Video processing system and method based on three-dimensional resistive random access memory array

By using a three-dimensional resistive memory array for video processing in edge-end devices, high-energy-efficient and high-frame-rate video processing effects are achieved, and the function of long-term data storage is provided, which solves the problem of low performance and energy efficiency in the prior art.

CN120201225APending Publication Date: 2025-06-24SEMICON TECH INNOVATION CENT(BEIJING) CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510468550.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing edge-side video processing hardware is difficult to achieve high energy efficiency, high frame rate and long-term data storage at the same time. Especially when processing high-resolution video and high frame rate video, the performance and energy efficiency are low, making it difficult to meet the real-time video processing needs at the edge-side.

Method used

A video processing system based on a three-dimensional resistive memory array is adopted. The video acquisition module is used to split the video into multiple frame images, and the three-dimensional resistive memory array is used for parallel processing to realize video compression and feature extraction. At the same time, multi-frame compressed images are fused in the time domain and long-term data storage is used using a storage module.

Benefits of technology

It realizes an energy-efficient and high frame rate edge video processing system, which can effectively improve the energy efficiency and frame rate of edge devices when performing video processing tasks, and also has the ability to store long-term data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201225A_ABST
    Figure CN120201225A_ABST
Patent Text Reader

Abstract

The invention provides a video processing system and method based on a three-dimensional resistive random access memory array, and relates to the field of semiconductor devices and integrated circuits. The system comprises a video acquisition module used for acquiring a to-be-processed video, splitting the to-be-processed video into continuous multiple frames of images, and converting each frame of image in the multiple frames of images into a voltage signal; the three-dimensional resistive random access memory array is used for carrying out parallel processing on the voltage signals of the multiple frames of images to obtain a compressed image of each frame of image, the three-dimensional resistive random access memory array comprises a plurality of two-dimensional resistive random access memory arrays arranged in the same direction, and each two-dimensional resistive random access memory array is used for processing one frame of image; the time domain fusion module is used for fusing at least one frame of compressed image in the multiple frames of compressed images in a time domain; and the storage module is used for storing the fused compressed image. According to the invention, based on the in-memory computing architecture, the energy efficiency and the frame rate when the edge end equipment executes the video processing task can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semiconductor devices and integrated circuits, and particularly to a video processing system and method based on a three-dimensional resistive random access memory (RRAM) array. Background Art

[0002] With the rapid development of artificial intelligence and Internet of Things technologies, video processing plays an increasingly important role in edge intelligent systems. Edge devices such as smartphones, smart cameras, and Internet of Things devices need to process high-resolution videos while meeting the requirements of real-time performance and energy efficiency. In image processing tasks, multiply-accumulate operations are the most common computational forms. In traditional solutions, a large number of multipliers and adders are required, which leads to large area and power consumption overheads, and the system performance is often limited by the bandwidth between the memory and the computing unit and thus difficult to further improve. In video processing tasks, the dimension of time is added, which poses higher requirements on the computing power and energy efficiency of the video processing system, especially for application scenarios with high frame rate video processing requirements at the edge.

[0003] Currently, edge devices mainly use hardware solutions such as system-on-chip (SoC), digital signal processor (DSP), or application-specific integrated circuit (ASIC) for video compression. However, these solutions face many challenges in practical applications. The SoC integrates multiple processing units and provides a comprehensive solution. However, its performance and energy efficiency are relatively low, making it difficult to meet the growing video processing requirements at the edge, especially in real-time high frame rate processing tasks. The DSP is designed specifically for signal processing tasks and also faces the problem of relatively low performance and energy efficiency in high frame rate video processing tasks. The ASIC is a chip customized for specific tasks and shows high energy efficiency in video compression. However, ASIC chips based on traditional CMOS processes and von Neumann architectures still have energy efficiency bottlenecks and have significant room for improvement in edge applications that pursue extreme energy efficiency.

[0004] The in-memory computing architecture is considered to be one of the effective solutions to further improve the system energy efficiency. Among them, resistive random access memory (RRAM) is considered to be one of the most promising candidates for implementing in-memory computing systems due to its low programming energy consumption, fast read / write speed, high degree of integration, and compatibility with CMOS processes. Researchers have also conducted extensive research on edge image processing systems based on this. However, in existing solutions, the use of RRAM arrays to implement image processing is limited to two-dimensional images and is difficult to effectively process three-dimensional videos containing time-domain information. Summary of the Invention

[0005] In view of the problems that existing edge video processing hardware is difficult to achieve high energy efficiency, high frame rate, and long-term data storage simultaneously, the present invention provides a video processing system and method based on a three-dimensional RRAM array.

[0006] On the one hand, the present invention provides a video processing system based on a three-dimensional resistive random access memory (RRAM) array, including: a video acquisition module, configured to acquire a video to be processed, split the video to be processed into a series of consecutive frames of images, and convert each frame of image in the series of frames of images into a voltage signal; a three-dimensional RRAM array, configured to perform parallel processing on the voltage signals of the series of frames of images to obtain a compressed image of each frame of image, wherein the three-dimensional RRAM array includes a plurality of two-dimensional RRAM arrays arranged in the same direction, and each two-dimensional RRAM array is configured to process one frame of image; a time-domain fusion module, configured to fuse at least one compressed image in the series of compressed images in the time domain; and a storage module, configured to store the fused compressed image.

[0007] On the other hand, the present invention provides a video processing method based on a three-dimensional RRAM array, including: acquiring a video to be processed, splitting the video to be processed into a series of consecutive frames of images, and converting each frame of image in the series of frames of images into a voltage signal; using a three-dimensional RRAM array to perform parallel processing on the voltage signals of the series of frames of images to obtain a compressed image of each frame of image, wherein the three-dimensional RRAM array includes a plurality of two-dimensional RRAM arrays arranged in the same direction, and each two-dimensional RRAM array is configured to process one frame of image; fusing at least one compressed image in the series of compressed images in the time domain; and using a storage module to store the fused compressed image.

[0008] The video processing system and method based on the three-dimensional RRAM array provided by the present invention have at least the following beneficial effects:

[0009] Based on the in-memory computing principle, the present invention develops an energy-efficient and high-frame-rate edge-side video processing system using a three-dimensional RRAM array. The video processing includes but is not limited to video compression, feature extraction, etc.

[0010] The present invention can effectively improve the energy efficiency and frame rate of edge-side devices when performing video processing tasks, and can also store the processed data locally for a long time. Thanks to the in-memory computing architecture based on resistive random access memory, the embodiments of the present invention have extremely high energy efficiency when performing image processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0012] Figure 1 Schematically shows an architecture diagram of a video processing system based on a three-dimensional RRAM array according to an embodiment of the present invention;

[0013] Figure 2A And Figure 2BSchematically shown are a circuit topology diagram and a structural diagram of a three-dimensional resistive random access memory (RRAM) array according to an embodiment of the present invention;

[0014] Figure 3A and Figure 3B Schematically shown are a flowchart and a schematic diagram of a video processing method based on a three-dimensional RRAM array according to an embodiment of the present invention;

[0015] Figure 4 Schematically shown is a flowchart of video decompression according to an embodiment of the present invention. Detailed Embodiments

[0016] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention fall within the scope of protection of the present invention.

[0017] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0019] It should be noted that the video processing system and method based on a three-dimensional RRAM array provided in the embodiments of the present invention can be applied to edge devices, and video processing can specifically be video compression, feature extraction, etc.

[0020] The following takes the video compression task processed by this system as an example to elaborate on the content of the present invention in detail.

[0021] Figure 1 Schematically shown is an architecture diagram of a video processing system based on a three-dimensional RRAM array according to an embodiment of the present invention.

[0022] As Figure 1 shown, the video processing system based on a three-dimensional RRAM array in this embodiment includes a video acquisition module, a three-dimensional RRAM array, a time-domain fusion module, and a storage module.

[0023] Among them, the video acquisition module is used to acquire the video to be processed, split the video to be processed into a series of consecutive frames of images, and convert each frame of image in the series of frames of images into a voltage signal. The three-dimensional resistive random access memory (RRAM) array is used to perform parallel processing on the voltage signals of the series of frames of images to obtain the compressed image of each frame of image. Among them, the three-dimensional RRAM array includes a plurality of two-dimensional RRAM arrays arranged in the same direction, and each two-dimensional RRAM array is used to process one frame of image. The time-domain fusion module is used to fuse at least one of the series of compressed images in the time domain. The storage module is used to store the fused compressed image.

[0024] Please continue to refer to Figure 1 , in some embodiments, the storage module is another three-dimensional RRAM array. The compressed video data fused with time-domain information can be stored in another three-dimensional RRAM array. Due to its high density and non-volatile characteristics, the data can be stored for a long time and interact with external storage when needed.

[0025] Furthermore, the video processing system based on the three-dimensional RRAM array may further include a decompression network, which is used to: when restoring the video to be processed, read the fused compressed image from another three-dimensional RRAM array; decompress the fused compressed image to reconstruct the video to be processed.

[0026] Through the embodiments of the present invention, the video to be processed will be sequentially processed by the above-mentioned modules, so as to complete the compression, time information fusion, compressed data storage, and decompression of the video to be processed.

[0027] Figure 2A and Figure 2B respectively schematically show the circuit topology diagram and structure diagram of the three-dimensional RRAM array according to the embodiments of the present invention.

[0028] As Figure 2A and Figure 2B shown, in this embodiment, the three-dimensional RRAM array includes:

[0029] a plurality of bit-line metal layers arranged along the length direction;

[0030] a plurality of selection-line layers arranged along the width direction;

[0031] a stacked structure arranged above the plurality of selection-line layers, where the stacked structure includes insulating layers and word-line metal layers stacked alternately along the height direction;

[0032] a plurality of columnar electrodes, and each columnar electrode vertically passes through a bit-line metal layer, a selection-line layer, and the stacked structure along the height direction respectively;

[0033] A resistive switching dielectric layer is formed between each columnar electrode and the stacked structure, the corresponding bit line metal layer, and the corresponding selection line layer respectively.

[0034] Among them, each columnar electrode serves as a gated transistor to connect the stacked structure to the corresponding bit line metal layer, and the selection of any word line metal layer in the stacked structure is controlled by the selection line layer.

[0035] In this embodiment, in the three-dimensional resistive switching memory array, each two-dimensional resistive switching memory array stores the convolution kernel weights for processing one frame of an image among multiple frames of images, and the convolution kernel weights are stored in the form of conductance.

[0036] Based on this, please continue to refer to Figure 2A , in the three-dimensional resistive switching memory array, a two-dimensional resistive switching memory array is in the xy plane. In each two-dimensional resistive switching memory array, the word lines and the selection lines are parallel to the x direction, the bit lines are parallel to the y direction, and the convolution kernel weights for compressing and processing (including but not limited to edge extraction, contrast enhancement, etc.) one frame of an image among multiple frames of images are stored in the form of conductance in this two-dimensional resistive switching memory array.

[0037] Please continue to refer to Figure 2B , in the three-dimensional resistive switching memory array, the bottom is a layer of vertical transistors, which serve as gated transistors to connect the resistive switching devices above to the bit lines, and the selection of any word line metal layer in the stacked structure is controlled by the selection line layer. The gated transistors are connected to the columnar electrodes along the z direction, and these columnar electrodes will serve as one electrode of the resistive switching devices, and the columnar electrodes pass through the word line metal layers and the insulating layers that are sequentially overlapped in the xy plane.

[0038] Through the above structure, the three-dimensional resistive switching memory array is divided into multiple slices along the z direction, that is, two-dimensional resistive switching memory arrays. Each two-dimensional resistive switching memory array includes n gated transistors, corresponding to n columns of resistive switching memories, where n is a positive integer, and each two-dimensional resistive switching memory array processes one frame of an image. The present invention adopts such a three-dimensional integrated resistive switching memory array, which can obtain good resistive switching performance and can further improve the integration density of the resistive switching memory array, so that the integration density is less than 4F 2 / bit, where F is the feature size.

[0039] In this embodiment, when the three-dimensional resistive switching memory array is used to perform convolution operations, the three-dimensional resistive switching memory array is used for:

[0040] Clamping each bit line metal layer at a fixed voltage, and applying the voltage signal of each frame of an image among multiple frames of images to the word line metal layer of the corresponding two-dimensional resistive switching memory array;

[0041] For each two-dimensional resistive random access memory (RRAM) array, obtain the voltage difference across the two-dimensional RRAM array according to the voltage signal, multiply the voltage difference by the convolution kernel weights stored in the two-dimensional RRAM array to obtain the current flowing through the two-dimensional RRAM array.

[0042] On multiple bit-line metal layers, accumulate the currents of each two-dimensional RRAM array to obtain the convolution operation result.

[0043] It can be understood that for each two-dimensional RRAM array, according to Ohm's law \(I = V\times G\), where \(I\) is the current flowing through the two-dimensional RRAM array, \(V\) is the voltage difference across the two-dimensional RRAM array, and \(G\) is the conductance of the two-dimensional RRAM array. Thus, any one frame of the multiple frames of images can be multiplied by the corresponding weight data first, and then the currents of each two-dimensional RRAM array are converged and accumulated on the bit-line metal layer, and thus the complete multiply-accumulate / convolution operation is completed.

[0044] Next, the two-dimensional RRAM array is periodically extended along the \(z\) direction, thus forming a three-dimensional RRAM array. Different two-dimensional RRAM arrays along the \(z\) direction perform parallel processing on different frames of images of the video to be processed, greatly improving the processing efficiency.

[0045] In some embodiments, the convolution kernel weights stored in each two-dimensional RRAM array are obtained through the following method:

[0046] Train a pre-constructed video compression network to obtain the trained weight data for subsequent participation in the operation;

[0047] Map the trained weight data to different two-dimensional RRAM arrays of the three-dimensional RRAM array.

[0048] Furthermore, the weight data is mapped to the three-dimensional RRAM array through the following method: store the weight data for processing one frame of image in the \(xy\) plane, and store the weight data of different frames of images in the \(z\) direction.

[0049] Correspondingly, for the aforementioned decompression network, the training and weight mapping methods for the video compression network can also be used for corresponding operations, and the specific details of the present invention will not be elaborated here.

[0050] In this embodiment, the time-domain fusion module fuses at least one of the multiple frames of compressed images in the time domain, including: inputting any two frames of the multiple frames of compressed images into the time-domain fusion module and outputting the difference value between the two frames of compressed images.

[0051] In this way, the differential weights can be stored for the processing slices of different frames in the 3D resistive random access memory (RRAM) array in-situ to directly complete the temporal difference operation of the video. In other embodiments, additional non-linear modules can be used to perform non-linear operations or 3D convolution operations on any two compressed images in the multi-frame compressed images, so as to fuse the temporal information of several consecutive compressed images into one frame of image, further compressing and reducing the data volume.

[0052] In the embodiments of the present invention, after the multi-frame images of the video to be processed are processed in parallel on different slices of the 3D RRAM array to obtain the compressed images of each frame, the compressed images of different frames are fused in the time domain to further improve the efficiency of system operation and data transmission, so as to realize an edge-side video processing system with high energy efficiency and high frame rate.

[0053] It should also be noted that the time-domain fusion module in the embodiments of the present invention can be a dedicated circuit module or an operation method for the 3D RRAM array. For example, if two slices in the 3D RRAM array are opened simultaneously, the output current of the 3D RRAM array is naturally the difference value between two input images, thus realizing the time-domain fusion of inter-frame difference.

[0054] From the above description, it can be seen that in the above embodiments of the present invention, the video to be processed is split frame by frame, and then each frame of image is converted into a voltage signal as the input of different slices of the 3D RRAM array. In the 3D RRAM array, the in-memory computing method is used to process each frame of video image, and then the temporal information of different frames of images is fused in the time domain. The compressed video data fused with the time-domain information is stored in another 3D RRAM array. The above embodiments of the present invention can effectively improve the energy efficiency and frame rate when the edge-side device executes video processing tasks, and at the same time, the processed data can be stored locally for a long time. Thanks to the in-memory computing architecture based on resistive random access memory, the embodiments of the present invention have extremely high energy efficiency when executing image processing tasks.

[0055] Taking the most basic convolution operator in the field of image processing as an example, since the convolution operation involves a large number of multiply-accumulate operations, in a traditional CMOS (Complementary Metal Oxide Semiconductor) circuit, a large number of multipliers and adders are required to complete it, which not only occupies a large area and consumes a large amount of power, but also the system performance is often difficult to be further improved due to the bandwidth limitation between the memory and the computing unit.

[0056] However, in the above embodiments of the present invention, the multiply-accumulate operation is completed in-situ in the 3D RRAM array in an analog form without complex calculation and data transmission modules, and has extremely high energy efficiency and computing parallelism.

[0057] Further, in the above embodiments of the present invention, the convolution kernel weights that need to participate in the operation are mapped frame by frame to different slices (i.e., two-dimensional resistive memory arrays) of the three-dimensional resistive memory array, so that the information of different frames in the video stream can be processed in parallel, further improving the system computing parallelism.

[0058] Based on the system disclosed in the above embodiments, the present invention also provides a video processing method based on a three-dimensional resistive memory array, which will be described in detail below.

[0059] Figure 3A and Figure 3B respectively schematically show a flowchart and a schematic diagram of a video processing method based on a three-dimensional resistive memory array according to an embodiment of the present invention.

[0060] As Figure 3A and Figure 3B shown, according to the video processing method based on a three-dimensional resistive memory array of this embodiment, it may include operation S310 to operation S340.

[0061] Operation S310, obtain a video to be processed, split the video to be processed into a plurality of consecutive frame images, and convert each frame image in the plurality of frame images into a voltage signal.

[0062] Split the video to be processed frame by frame, and then convert each frame image into a voltage signal as the input of different slices of the three-dimensional resistive memory array.

[0063] Operation S320, use the three-dimensional resistive memory array to perform parallel processing on the voltage signals of the plurality of frame images to obtain a compressed image of each frame image, wherein the three-dimensional resistive memory array includes a plurality of two-dimensional resistive memory arrays arranged in the same direction, and each two-dimensional resistive memory array is used to process one frame image.

[0064] Taking the convolution operation as an example, the voltage signal converted from one frame image can be multiplied by the convolution kernel weights stored in the form of conductance in a two-dimensional resistive memory array, and the accumulation is completed in the form of current on the bit line metal layer to implement the convolution operation and complete feature extraction and image compression.

[0065] Operation S330, fuse at least one compressed image among the plurality of compressed images in the time domain.

[0066] For example, operations such as time-domain difference or three-dimensional convolution can be performed to fuse the information of several consecutive frames of video into one frame image, further compressing and reducing the data volume.

[0067] Specifically, the three-dimensional resistive random access memory can be utilized in-situ to store differential weights for the processing slices of different frames to directly complete the temporal difference operation of the video, perform non-linear operations using an additional non-linear module, etc., and fuse at least one compressed image in the multi-frame compressed image in the time domain.

[0068] Operation S340, use the storage module to store the fused compressed image.

[0069] For example, the storage module is another three-dimensional resistive random access memory array. The compressed video data fused with the time domain information can be stored in another three-dimensional resistive random access memory array. Due to its high density and non-volatile characteristics, the data can be stored for a long time and interact with external storage when needed.

[0070] Figure 4 Schematically shows a flowchart of video decompression according to an embodiment of the present invention.

[0071] In some embodiments, the video processing method may further include operations S350 to S360 after the above operation S340.

[0072] Operation S350, in the case of restoring the video to be processed, read the fused compressed image from another three-dimensional resistive random access memory array.

[0073] Operation S360, use the decompression network to decompress the fused compressed image to reconstruct the video to be processed.

[0074] For example, when it is necessary to restore the original video to be processed, read the fused compressed image from another three-dimensional resistive random access memory array, and then input it into the decompression network to reconstruct and output the decompressed video.

[0075] It should be noted that the embodiments of the system part are similar to the embodiments of the method part, and the achieved technical effects are also similar. For the specific details of the embodiments of the two parts, please refer to each other and will not be elaborated here.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0077] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the words "a" or "an" preceding an element do not exclude the existence of a plurality of such elements.

[0078] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0079] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A video processing system based on a three-dimensional resistive memory array, characterized in that: include: A video acquisition module, used for acquiring a video to be processed, splitting the video to be processed into multiple continuous frames of images, and converting each frame of the multiple frames of images into a voltage signal; A three-dimensional resistive memory array, used for processing the voltage signals of the multiple frames of images in parallel to obtain a compressed image of each frame of image, wherein the three-dimensional resistive memory array includes a plurality of two-dimensional resistive memory arrays arranged in the same direction, and each two-dimensional resistive memory array is used for processing one frame of image; A time domain fusion module, used for fusing at least one compressed image frame of the multiple compressed images in the time domain; The storage module is used to store the fused compressed image.

2. The system according to claim 1, characterized in that The storage module is another three-dimensional resistive memory array.

3. The system according to claim 2, characterized in that The system also includes a decompression network for: When the video to be processed is restored, the fused compressed image is read from the other three-dimensional resistive random access memory array; The fused compressed image is decompressed to reconstruct the video to be processed.

4. The system according to claim 1, characterized in that The three-dimensional resistive memory array comprises: A plurality of bit line metal layers arranged along the length direction; A plurality of selection line layers arranged along the width direction; A stacked structure disposed above the plurality of selection line layers, wherein the stacked structure comprises insulating layers and word line metal layers alternately stacked in a height direction; A plurality of columnar electrodes, each columnar electrode vertically passes through one of the bit line metal layers, one of the selection line layers and the stacked structure along a height direction; A resistive dielectric layer is formed between each columnar electrode and the stacked structure, the corresponding bit line metal layer, and the corresponding selection line layer; Each of the columnar electrodes serves as a gate transistor to connect the stacked structure to a corresponding bit line metal layer, and the selection line layer controls the gate of any word line metal layer in the stacked structure.

5. The system according to claim 4, characterized in that In the three-dimensional resistive memory array, each two-dimensional resistive memory array stores a convolution kernel weight for processing one frame of the multiple frames of images, and the convolution kernel weight is stored in the form of conductance.

6. The system according to claim 5, characterized in that When the three-dimensional resistive memory array is used to perform a convolution operation, the three-dimensional resistive memory array is used to: Clamp each bit line metal layer at a fixed voltage, and apply the voltage signal of each frame of the multiple frames of images to a corresponding word line metal layer of a two-dimensional resistive memory array; For each two-dimensional resistive memory array, a voltage difference between two ends of the two-dimensional resistive memory array is obtained according to the voltage signal, and the voltage difference is multiplied by the convolution kernel weight stored in the two-dimensional resistive memory array to obtain a current flowing through the two-dimensional resistive memory array; On the multiple bit line metal layers, the current of each two-dimensional resistive memory array is accumulated to obtain a convolution operation result.

7. The system according to claim 5, characterized in that The convolution kernel weights stored in each two-dimensional resistive memory array are obtained in the following manner: Train the pre-built video compression network to obtain trained weight data for subsequent operations; The trained weight data is mapped to different two-dimensional resistive memory arrays in the three-dimensional resistive memory array.

8. The system according to claim 7, characterized in that The time domain fusion module fuses at least one compressed image frame of the multiple compressed images in the time domain, including: Any two compressed image frames among the multiple compressed image frames are input into the time domain fusion module, and the difference value between the two compressed image frames is output.

9. A video processing method based on a three-dimensional resistive memory array, characterized in that: include: Acquire a video to be processed, split the video to be processed into multiple continuous frames of images, and convert each frame of the multiple frames of images into a voltage signal; Using a three-dimensional resistive memory array to process the voltage signals of the multiple frames of images in parallel to obtain a compressed image of each frame of image, wherein the three-dimensional resistive memory array includes a plurality of two-dimensional resistive memory arrays arranged in the same direction, and each two-dimensional resistive memory array is used to process one frame of image; Fusing at least one compressed image frame among the multiple compressed images in the time domain; Use the storage module to store the fused compressed image.

10. The method according to claim 9, characterized in that The storage module is another three-dimensional resistive memory array; the method further includes: When the video to be processed is restored, the fused compressed image is read from the other three-dimensional resistive random access memory array; The fused compressed image is decompressed using a decompression network to reconstruct the video to be processed.