Feature extraction method based on 2T0C DRAM-1T1R RRAM fusion operator
Efficient feature extraction of point cloud data is achieved through the 2T0C DRAM-1T1R RRAM fusion operator, which solves the problem of point cloud accelerators frequently accessing external memory, improves processing speed and energy efficiency, reduces system overhead, and is suitable for a variety of point cloud neural network algorithms.
Patent Information
- Application Number
- CN202510729372.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-16
AI Technical Summary
When performing downsampling operations, existing point cloud accelerators need to frequently access the original point cloud data in external memory, resulting in low cache hit rate and large data access bandwidth occupancy. In addition, existing in-memory computing technology needs to convert data to the analog domain for calculation, introducing additional data conversion overhead.
A 2T0C DRAM-1T1R RRAM fusion operator is used. The 2T0C DRAM stores the sign, exponent, and mantissa bits of the point cloud data, and the 1T1R RRAM stores the weights. This allows point cloud data to be converted from the floating-point digital domain to the analog domain. Matrix-vector multiplication is performed in the analog domain, and feature extraction is performed using a digital-to-analog converter and a shift adder.
It eliminates the need for expensive data conversion overhead, significantly improves processing speed and energy efficiency, provides efficient hardware acceleration infrastructure, reduces system overhead, and is suitable for a variety of point cloud neural network algorithms.
Smart Images

Figure CN120655931A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of novel in-memory computing technology and three-dimensional point cloud recognition, and specifically relates to a feature extraction method based on a 2T0CDRAM-1T1R RRAM fusion operator. Background Art
[0002] As an important means of three-dimensional environmental perception, three-dimensional point cloud recognition technology plays an important role in many fields such as autonomous driving, robot navigation, virtual reality, augmented reality, industrial inspection and smart cities, and has irreplaceable application value. Compared with traditional two-dimensional images, three-dimensional point cloud data directly obtains spatial coordinate information through lidar or depth camera, which can accurately characterize the geometric structure and spatial relationship of the target. In particular, the recognition robustness under complex lighting conditions and dynamic scenes is significantly better than that of visual solutions. At present, a variety of neural networks have been developed in the field for efficient three-dimensional point cloud recognition. Among them, point-based point cloud neural networks (such as PointNet, PointNet++, etc.) have become the mainstream in the current research field due to their direct and efficient processing of disordered point clouds and excellent network performance, omitting additional processing steps such as voxelization.
[0003] Existing point cloud accelerators frequently access raw point cloud data stored in external memory when performing downsampling operations. Because point cloud data is inherently a sparse, unstructured set of 3D coordinates, it lacks good spatial or cache locality, resulting in low cache hit rates and high data access bandwidth usage.
[0004] Currently, proposals have been made to accelerate matrix-vector multiplication using analog in-memory computing techniques (e.g., those based on resistive random access memory) to address the high computational overhead of feature extraction. However, this requires converting the original point cloud from its original floating-point data type to the analog domain for computation, introducing additional data conversion overhead. Therefore, utilizing in-memory computing techniques to optimize the design of efficient operators for the key steps of point cloud recognition, such as downsampling and feature calculation, and to achieve matching of data structure types, is of great research significance. Summary of the Invention
[0005] The present invention provides a feature extraction method based on a 2T0C DRAM-1T1R RRAM fusion operator, which can significantly improve processing speed and energy efficiency, and provide a universal and efficient hardware acceleration basic configuration for various point cloud neural network algorithms.
[0006] To achieve the above objectives, the present invention provides the following technical solutions:
[0007] A feature extraction method based on a 2T0C DRAM-1T1R RRAM fusion operator comprises the following steps:
[0008] 1) The 2T0C DRAM array is used to store input point cloud data. Each 2T0C DRAM cell consists of a write transistor and a read transistor. The WWL and WBL of the 2T0C DRAM array are used for data writing, and the RWL and RBL of the 2T0C DRAM array are used for data reading and calculation. The drain of the write transistor is connected to the gate of the read transistor to form an information storage node SN. The 1T1R RRAM array is used to store the first-layer MLP weights. Each 1T1R RRAM cell consists of a transistor and a resistive memory device connected in series. The word line WL of the 1T1R RRAM array is parallel to the source line SL, and the bit line BL is perpendicular to the word line WL. The read word line RWL of the 2T0C DRAM is converted to a 0 / 1 voltage by the sense amplifier SA and then directly connected to the gate of a column of transistors in the 1T1R RRAM array;
[0009] 2) A read voltage is applied to the RBL of the 2T0C DRAM array to read a specific bit in the mantissa of a point cloud data. Within the fusion operator, the mantissa bit information is output from the 2T0C DRAM array through the RWL of the 2T0C DRAM array. After being stabilized by the sense amplifier SA, it is converted into a stable voltage and connected to the WL of the 1T1R RRAM array. At the same time, a bit-by-bit inference voltage is input to the BL of the 1T1R RRAM array. The relationship between the inference voltage and the storage weight of the row is as follows:
[0010]
[0011] Where n is the number of rows of resistive memory devices in the 1T1R RRAM array, V BLn is the input voltage of the nth row, V inf Represents the benchmark inference voltage, which is the weighted current value obtained by multiplying a certain bit in the mantissa of the point cloud data by the weight on the SL of the 1T1R RRAM array;
[0012] 3) Reading the mantissa bits of the point cloud data row by row, repeating step 2) to obtain a weighted current value obtained by multiplying all mantissa bits in the point cloud data by the weight; after quantization by a digital-to-analog converter, extracting the exponent bit information and sign bit information of the point cloud data, performing shift and addition operations, obtaining the output of the first layer MLP, and entering the next layer MLP calculation;
[0013] 4) For the second and subsequent MLP layers, the entire process is implemented in the analog domain. The weight storage is represented by the resistance value of the RRAM in the 1T1R RRAM unit. By applying an analog inference voltage to the BL of the 1T1R RRAM, the analog current result of the multiplication and accumulation operation is obtained on the SL of the 1T1R RRAM. The result is then fed into the next MLP layer for calculation. When all MLP layers complete the calculation, all point cloud features are extracted.
[0014] Furthermore, in step 1), the sign bit, exponent bit, and mantissa bit of the point cloud data are stored separately, wherein the mantissa bit of the same point cloud data is stored separately in a column of the 2T0C DRAM array, and all the mantissa bit information constitutes the mantissa bit storage array, and the sign bit and exponent bit of the same point cloud data are stored together in a column of 2T0C DRAM cells in the 2T0C DRAM array.
[0015] Furthermore, in step 1), the pre-trained point cloud neural network weights are quantized to be mapped to the RRAM device.
[0016] Furthermore, in step 4), all point cloud features are pooled to aggregate features, and classified through a fully connected layer neural network to obtain the final recognition result.
[0017] The present invention has the following advantages compared with the prior art:
[0018] This invention constructs a point cloud feature extraction array based on a 1T1R RRAM array, accelerating matrix-vector multiplication computations in situ. A 2T0CDRAM-1T1R RRAM fusion operator is used to convert point cloud data from the floating-point digital domain to the analog domain. This is then accelerated using in-memory computation in the analog domain to extract all point cloud features. This invention eliminates the need for expensive data conversion overhead and organically combines two in-memory computation methods, further reducing system overhead. This provides a universal and efficient hardware acceleration infrastructure for a variety of point cloud neural network algorithms, making it an effective in-memory acceleration solution for point cloud recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of the 2T0CDRAM-1T1RRRAM fusion operator array structure and related peripheral circuits of a specific embodiment of the present invention. DETAILED DESCRIPTION
[0020] The present invention will be further described below with reference to the accompanying drawings through examples, which however do not limit the scope of the present invention in any way.
[0021] Figure 1 This is a schematic diagram of the principle of feature extraction based on the 2T0CDRAM-1T1RRRAM fusion operator of the present invention.
[0022] The 2T0C DRAM array of the present invention is used to store input point cloud data. Each 2T0C DRAM cell is composed of two transistors. The drain end of the write transistor is connected to the gate end of the read transistor to form an information storage node SN. The WWL and WBL of the 2T0C DRAM array are used for data writing, and the RWL and RBL of the 2T0C DRAM array are used for data reading and calculation. The present invention stores the sign bit, exponent bit, and mantissa bits of the point cloud data separately, wherein the mantissa bits of the same point cloud data are individually stored in a column (m, n...) of the 2T0C DRAM array. All mantissa bit information constitutes a mantissa bit storage array. The sign bit (s) and exponent bit (e) of the same point cloud data are stored together in a column of 2T0C DRAM cells in the 2T0C DRAM array.
[0023] The 1T1R RRAM array of the present invention is used to store the weights of the first layer of MLP in the feature calculation process. The conductance of the RRAM is positively correlated with the weights. Each 1T1R RRAM cell in the 1T1R RRAM array is composed of a transistor and a resistive memory device connected in series. The word lines WL of the 1T1R RRAM array are parallel to the source lines SL, and the bit lines BL are perpendicular to the word lines WL. The weights of the pre-trained point cloud neural network are mapped to the RRAM device through weight quantization. In the 1T1R RRAM cell, an inference voltage is applied to the bit line BL to obtain a current value on the source line SL that is the product of the inference voltage and the RRAM conductance. In the 1T1R RRAM array, the currents of multiple cells on the SL are summed to complete a multiplication and accumulation operation. The cell is connected to the peripheral circuit of the 1T1R RRAM array, and the peripheral circuit includes a multiplexer, an analog-to-digital converter, and a shift adder.
[0024] The present invention performs a feature extraction method based on a 2T0C DRAM-1T1R RRAM fusion operator, and the steps include:
[0025] 1) After the read word line RWL of the 2T0C DRAM is converted to a 0 / 1 voltage by the sense amplifier SA, it is directly connected to the gate of the transistors in a column of the 1T1RRRAM array. The number of rows in this column is determined by the quantization accuracy required by the first-layer MLP. Higher quantization accuracy requires more rows and devices. The devices in this column are all set to binary quantization to ensure higher calculation accuracy. The number of rows in this column is consistent with the quantization accuracy of the first-layer MLP. If it consists of 8 rows of devices, 8-bit quantization can be achieved.
[0026] 2) When performing a fusion operator calculation on the 2T0C DRAM-1T1R RRAM array, a read voltage is applied to the RBL of the 2T0C DRAM array to read a specific bit in the mantissa of a point cloud data point. Within the fusion operator, this bit information is output from the 2T0C DRAM array via the RWL and input to the sense amplifier SA through a connection. After being stabilized by the sense amplifier SA, it is converted into a stable voltage output and continuously connected to the WL of the 1T1R RRAM array. At the same time, a bit-by-bit inference voltage is input to the BL of the 1T1R RRAM array. The magnitude of the inference voltage is related to the position of the storage weight of the row, as shown in the following relationship:
[0027]
[0028] Where n is the row number of the device, corresponding to the weight position. At this time, the weighted current value obtained on the SL of the 1T1R RRAM array is the product of a certain bit in the mantissa of the point cloud data and the weight.
[0029] In the specific embodiment of the present invention, feature extraction of point M is performed as an example. First, a read voltage is applied to RBL to obtain the mantissa of the floating-point coordinate of point M in a certain direction (for example, the X direction):
[0030]
[0031] Taking m0 as an example, after being converted into a stable 0 / 1 voltage signal by the SA sense amplifier, it is input to the gate of the transistor in the 1T1R RRAM array that stores the first-layer MLP weight column. Taking a weight W as an example, the weight is expressed as follows after binary quantization:
[0032] W=W0·2 0 +W1·2 -1 +W2·2 -2 +…+W n 2 -n
[0033] Where (n+1) represents the number of quantized bits for the weights. If n=7, the weights are quantized to 8 bits, representing 0 / 1. At the same time, a bitwise inference voltage is input to the BL of the 1T1R RRAM weight array. The magnitude of the inference voltage is related to the bit position of the weight stored in that row. The specific relationship is as follows:
[0034]
[0035] Where n is the number of rows of resistive memory devices in the 1T1R RRAM array, V BLn is the input voltage of the nth row, V infIndicates the reference inference voltage. At this time, the weighted current value obtained on the SL is the product of a certain digit in the point cloud data (for example, m0) and the weight:
[0036]
[0037] 3) Read the mantissa bits of the point cloud data row by row, repeat the above operation to obtain the current value of all the mantissa bits of the point multiplied by the weight, quantize it through the digital-to-analog converter in the peripheral circuit, and perform shift and addition operations on the exponent bit information and sign bit information of the point to obtain the final first-layer MLP output and enter the next-layer MLP calculation.
[0038] 4) For the second and subsequent MLP layers, the entire process is implemented in the analog domain. Weight storage is represented by the resistance value of the RRAM in the 1T1RRRAM unit, leveraging the multi-valued nature of RRAM to improve computational efficiency. By applying an analog inference voltage to the BL, the analog current result of the multiplication-accumulation operation can be obtained on the SL. This current is quantized through multiplexers and digital-to-analog converters in the peripheral circuits and enters the next MLP layer for calculation.
[0039] When all MLP layers complete the calculation, all the extracted point cloud features are pooled and aggregated, and then classified through the fully connected layer neural network to obtain the final recognition result.
[0040] The above embodiments are only some preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any changes that adopt the design principles of the present invention and are made through non-creative work on this basis should fall within the scope of protection of the present invention.
Claims
1. A feature extraction method based on a 2T0C DRAM-1T1R RRAM fusion operator, comprising the following steps: 1) The 2T0C DRAM array is used to store input point cloud data. Each 2T0C DRAM cell consists of a write transistor and a read transistor. The WWL and WBL of the 2T0C DRAM array are used for data writing, and the RWL and RBL of the 2T0C DRAM array are used for data reading and calculation. The drain of the write transistor is connected to the gate of the read transistor to form an information storage node SN. The 1T1R RRAM array is used to store the first-layer MLP weights. Each 1T1R RRAM cell consists of a transistor and a resistive memory device connected in series. The word line WL of the 1T1R RRAM array is parallel to the source line SL, and the bit line BL is perpendicular to the word line WL. The read word line RWL of the 2T0C DRAM is converted to a 0 / 1 voltage by the sense amplifier SA and then directly connected to the gate of a column of transistors in the 1T1R RRAM array; 2) A read voltage is applied to the RBL of the 2T0C DRAM array to read a specific bit in the mantissa of a point cloud data. Within the fusion operator, the mantissa bit information is output from the 2T0C DRAM array through the RWL of the 2T0C DRAM array. After being stabilized by the sense amplifier SA, it is converted into a stable voltage and connected to the WL of the 1T1R RRAM array. At the same time, a bit-by-bit inference voltage is input to the BL of the 1T1R RRAM array. The relationship between the inference voltage and the storage weight of the row is as follows: Where n is the number of rows of resistive memory devices in the 1T1R RRAM array, V BLn is the input voltage of the nth row, V inf Represents the benchmark inference voltage, which is the weighted current value obtained by multiplying a certain bit in the mantissa of the point cloud data by the weight on the SL of the 1T1R RRAM array; 3) Reading the mantissa bits of the point cloud data row by row, repeating step 2) to obtain a weighted current value obtained by multiplying all mantissa bits in the point cloud data by the weight; after quantization by a digital-to-analog converter, extracting the exponent bit information and sign bit information of the point cloud data, performing shift and addition operations, obtaining the output of the first layer MLP, and entering the next layer MLP calculation; 4) For the second and subsequent MLP layers, the entire process is implemented in the analog domain. The weight storage is represented by the resistance value of the RRAM in the 1T1R RRAM unit. By applying an analog inference voltage to the BL of the 1T1R RRAM, the analog current result of the multiplication and accumulation operation is obtained on the SL of the 1T1R RRAM. The result is then fed into the next MLP layer for calculation. When all MLP layers complete the calculation, all point cloud features are extracted.
2. The feature extraction method based on the 2T0C DRAM-1T1R RRAM fusion operator according to claim 1, wherein: In step 1), the sign bit, exponent bit, and mantissa bit of the point cloud data are stored separately, wherein the mantissa bit of the same point cloud data is stored separately in a column of the 2T0C DRAM array. All the mantissa bit information constitutes the mantissa bit storage array, and the sign bit and exponent bit of the same point cloud data are stored together in a column of 2T0C DRAM cells in the 2T0C DRAM array.
3. The feature extraction method based on the 2T0C DRAM-1T1R RRAM fusion operator according to claim 1, wherein: In step 1), the pre-trained point cloud neural network weights are quantized to be mapped to the RRAM device.
4. The feature extraction method based on the 2T0C DRAM-1T1R RRAM fusion operator according to claim 1, wherein: In step 4), all point cloud features are pooled and aggregated, and classified through a fully connected layer neural network to obtain the final recognition result.