Image processing data analysis platform based on memristor storage and calculation integrated architecture
By adopting dynamic weight synchronization network and parallel sharded convolution accelerator in the image processing data analysis platform, combining the memristor array calculation module and dynamic resource allocation mechanism, the efficiency and reliability problems of the memristor image processing system in the prior art are solved, and efficient and low-energy image processing is achieved.
Patent Information
- Application Number
- CN202510390967.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Existing memristor-based image processing systems face problems such as static weight configurations that are difficult to adapt to image feature changes, a single array cannot meet the needs of high-resolution image processing, and high sensitivity to simulation computing, resulting in insufficient generalization capabilities, low processing efficiency and insufficient reliability.
The image processing data analysis platform based on the XC7Z020 FPGA control module is adopted, combined with the memristor array calculation module and dynamic resource allocation mechanism, and the dynamic weight synchronization network and parallel sharded convolution accelerator are used to realize high-speed processing of image data and low-energy operation.
It significantly improves the speed, energy efficiency and reliability of image processing, improves recognition accuracy and throughput, meets the real-time processing requirements of high-resolution images, and achieves low error recognition rates in industrial scenarios.
Smart Images

Figure CN120029969A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data recognition technology, and in particular to an image processing data analysis platform based on a memristor storage and computing integrated architecture. Background Art
[0002] As the global manufacturing industry accelerates its transformation towards intelligence and automation, computer vision technology, as the core means of information perception and decision-making in industrial scenarios, has been widely used in intelligent detection, precise sorting, robot navigation and other fields. However, the current mainstream image processing systems are mostly based on traditional CPU or GPU architectures. The von Neumann architecture with separated computing and storage leads to frequent data migration, resulting in a serious "memory wall" bottleneck. Taking real-time defect detection on industrial production lines as an example, traditional systems need to transmit massive image data from sensors to computing units, and return control instructions after processing by multi-layer convolutional neural networks. This process not only consumes a lot of time, but also usually has a single frame processing delay of more than 10ms. In addition, the energy consumption generated by data handling accounts for more than 60% of the total power consumption. In addition, complex environmental factors such as lighting changes and mechanical vibrations in industrial scenes lead to a significant increase in image noise. Traditional fixed-weight convolution models are prone to misidentification due to feature matching deviations, with a typical misjudgment rate of >5%, which seriously restricts production line efficiency and quality control level. Existing solutions such as the use of dedicated ASIC chips can improve computing speed, but their hardware solidification characteristics are difficult to adapt to dynamically changing industrial needs, and the development cost is high, making it difficult to deploy on a large scale.
[0003] In this context, the memristor storage-computing integrated architecture shows breakthrough potential. As a new type of device with non-volatile resistive memory characteristics, memristors can directly characterize weight parameters through conductance values, realizing the parallel processing mode of "storage is computing" at the physical level. In theory, memristor arrays can complete data storage and matrix multiplication and addition operations in the same hardware, eliminating the data handling overhead in traditional architectures, thereby significantly reducing latency and energy consumption. However, existing memristor-based image processing systems still face three key challenges: First, static weight configuration is difficult to adapt to the dynamic changes of image features in industrial scenarios, resulting in insufficient model generalization ability; second, the physical scale of a single memristor array is limited and cannot meet the parallel processing requirements of high-resolution images: 4K industrial cameras; third, the inherent noise sensitivity of analog computing may cause result drift, making it difficult to meet industrial-grade reliability requirements, such as the misrecognition rate of <0.1%. Existing technologies attempt to improve by stacking multiple memristor arrays or introducing redundant check modules, but these solutions often lead to a surge in system complexity and fail to fundamentally solve the problem of balancing computing efficiency and accuracy. Summary of the invention
[0004] 1. Technical issues to be resolved
[0005] In view of the shortcomings of the prior art, the present invention provides an image processing data analysis platform based on a memristor storage and computing integrated architecture, which solves the problem of how to achieve high-speed processing, low-energy operation and high-reliability identification of image data in industrial scenarios through a memristor storage and computing integrated architecture.
[0006] (II) Technical solution
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: an image processing data analysis platform based on a memristor storage and computing integrated architecture, comprising:
[0008] The XC7Z020 FPGA control module is connected to the OV5640 camera image acquisition module, the memristor array computing module, and the Internet of Things module respectively; it should be further explained that the platform of the present invention takes the XC7Z020 FPGA control module as the core, and realizes the complete process of image acquisition, processing, calculation and mechanical control through the collaboration of hardware interface and software logic.
[0009] The OV5640 camera image acquisition module is used to acquire raw image data and transmit it to the XC7Z020FPGA control module; it should be further explained that the OV5640 camera image acquisition module is connected to the PS end of the FPGA through the I2C bus to complete the camera initialization configuration, including resolution and frame rate settings, and transmits the raw image data to the DDR3 cache area of the FPGA through the DCMI interface.
[0010] The memristor array computing module includes a D / A converter, a memristor array, a transimpedance amplifier and an A / D converter, and is used to receive the image data processed by the XC7Z020 FPGA control module, complete the convolution operation through parallel analog calculation, and return the result to the XC7Z020 FPGA control module; it should be further explained that the memristor array computing module is composed of a multi-channel D / A converter, a memristor array, a transimpedance amplifier and an A / D converter. The PL end of the FPGA distributes the pre-processed image data to the D / A converter, converts it into an analog voltage signal and loads it to the bit line (BLx) of the memristor array. After completing the convolution operation, the current signal of SLx is converted into a voltage signal through the transimpedance amplifier, and then transmitted back to the FPGA by the A / D converter.
[0011] The Internet of Things module receives instructions from the XC7Z020 FPGA control module via the CAN bus and executes predefined actions. It should be further explained that the Internet of Things module receives instructions from the FPGA via the CAN bus and drives the servo motor and the pneumatic gripper to perform sorting, positioning or obstacle avoidance operations.
[0012] The platform also includes a dynamic resource allocation mechanism Dynamic Resource Allocation Module, DRAM, which integrates a real-time task scheduling unit through the XC7Z020 FPGA control module to dynamically allocate computing tasks to the FPGA logic unit or the memristor array computing module according to the complexity of the input image; The real-time task scheduling unit generates task priorities by analyzing the spatial frequency, edge density and grayscale distribution of the image, and allocates the convolution operation of the high-complexity area to the memristor array computing module for parallel analog computing, while the low-complexity processing tasks are retained in the FPGA logic unit for execution; Among them, the weights of the memristor array are pre-written in the form of conductance and loaded to the array nodes through voltage signals to achieve convolution operations of storage and computing. It should be further explained that the weights of the memristor array are pre-written in the form of conductance, and the dynamic resource allocation mechanism (DRAM) is used to dynamically schedule computing resources according to the complexity of the image to achieve efficient processing of storage and computing.
[0013] Preferably, the XC7Z020 FPGA control module is used to segment, compress and grayscale the original image to generate preprocessed data; and transmit the preprocessed data to the memristor array computing module, and control the Internet of Things module to perform actions. It should be further explained that the XC7Z020 FPGA control module segments, compresses and grayscales the original image captured by the OV5640 camera. Image segmentation adopts a dynamic blocking strategy to divide the image into sub-areas of different sizes according to texture complexity, including: high-texture areas use 64×64 tiles, low-texture areas use 256×256 tiles, and the edges of the tiles are processed by bilinear interpolation. The compression process uses differential pulse code modulation DPCM and Huffman coding, with a compression ratio of 2:1, and the hardware logic unit achieves a throughput of 1.5Gbps. Grayscale processing uses the brightness formula recommended by ITU: Y=0.299R+0.587G+0.114B; RGB images are converted into 8-bit grayscale data, and the FPGA's DSP Slice completes multiplication and addition operations in parallel, with a single frame processing delay of 1.08ms. The pre-processed data is transmitted to the memristor array computing module via the LVDS interface. The FPGA also generates a robot arm action command based on the recognition result, which is encapsulated into a data frame containing coordinates and speed parameters through the CAN bus protocol, and drives the servo motor to complete the operation.
[0014] Preferably, the memristor array computing module includes a dynamic weight synchronization network DWS-Net, that is: the XC7Z020 FPGA control module is embedded with a lightweight feature extraction unit, and the spatial frequency and texture distribution of the input image are analyzed in real time to generate a feature vector; based on the feature vector, the conductance values of different nodes in the memristor array are dynamically adjusted through the D / A converter to achieve adaptive allocation of convolution kernel weights; the dynamic weight synchronization network optimizes weight matching through a closed-loop feedback mechanism to reduce redundant computing energy consumption. It should be further explained that the dynamic weight synchronization network DWS-Net dynamically adjusts the weight allocation of the memristor array through real-time image feature feedback. The PS end of the FPGA is embedded with a lightweight feature extraction unit, and the high-frequency energy proportion is calculated using FFT and the Sobel operator is used to count the edge density to generate a 24-bit feature vector. According to the feature vector, the D / A converter generates a voltage pulse of a specific amplitude, wherein +3V is used for SET operation and -3V is used for RESET operation, and the conductance value of the memristor node is dynamically adjusted. Including: when the high-frequency energy proportion is >60%, the conductance value of the edge detection kernel is enhanced, and the weight of the smoothing filter kernel is reduced in low-contrast scenes. The closed-loop feedback mechanism combines the dual-modal verification results. If the confidence level is lower than the threshold, the conductance value is corrected. The weight adjustment realizes the weighted summation of the current signal through Kirchhoff's current law. The single adjustment delay is <0.5ms, and the false recognition rate is reduced.
[0015] Preferably, the memristor array computing module integrates a parallel slice convolution accelerator PSCA, that is: the XC7Z020 FPGA control module divides the input image into N×N independent sub-regions, which are regarded as slices, and each slice is assigned to an independent sub-array of the memristor array; each sub-array completes the local convolution operation through parallel analog calculation, and synchronously outputs the current signal through the multiplexing technology of the transimpedance amplifier; the A / D converter converts the analog signals of the multi-channel sub-arrays into digital signals, and the XC7Z020 FPGA control module performs global feature fusion. It should be further explained that the parallel slice convolution accelerator PSCA divides the high-resolution image into N×N independent sub-regions, including 4×4 and 8×8 slices, and each slice is assigned to an independent sub-array of the memristor array. FPGA distributes the slice data to the multi-channel D / A converter through the AXI-Stream interface, converts it into a 0-1V analog voltage and loads it to the BLx line. Each sub-array completes the analog convolution operation in parallel. The current signal output by SLx is converted into a voltage signal through the transimpedance amplifier gain, and then switched to the A / D converter for synchronous sampling through the 32-way analog switch ADG732. The FPGA aligns the slicing results by timestamp and uses weighted average: high-complexity slicing weight 0.7 or maximum value fusion algorithm to generate a global feature map. In slicing mode, the 4K image processing speed is compressed from 12ms to 2.4ms, and the throughput is increased by 5 times.
[0016] Preferably, the XC7Z020 FPGA control module deploys a dual-modal verification mechanism DMVM: constructs a memristor array analog convolution operation path for the first mode and an FPGA digital logic verification path for the second mode; compares the calculation results of the two modes through a confidence assessment module, and triggers secondary calculation or alarm if the difference exceeds a preset threshold; the dual-modal verification mechanism controls the misrecognition rate below 0.05% in a noise interference environment. It should be further explained that the dual-modal verification mechanism DMVM constructs a dual path of analog calculation and digital verification. The analog path completes the convolution operation by the memristor array, and the digital path performs fixed-point multiplication and addition operations through the FPGA logic unit. The confidence assessment module calculates the Euclidean distance of the results of the two paths, and triggers secondary calculation or alarm if the difference exceeds a dynamic threshold, wherein the initial value is 0.1, and is adaptively adjusted according to historical data. Including: Under noise interference SNR=15dB, the misrecognition rate of the analog path is 0.12%, which is reduced to 0.03% after digital verification. The FPGA outputs high and low level alarm signals through the GPIO pins, drives the LED or buzzer to indicate abnormalities, and records the fault type to the Block RAM for subsequent diagnosis.
[0017] Preferably, the weight adjustment process of the dynamic weight synchronization network DWS-Net includes: generating a voltage pulse signal through the D / A converter to adjust the initial conductance value of the memristor array; dynamically correcting the conductance distribution according to the real-time image characteristics to match the spatial frequency and grayscale gradient of the input data; the weight adjustment process realizes the weighted summation of the current signal through Kirchhoff's current law to complete the convolution operation. It should be further explained that the weight adjustment process of DWS-Net includes voltage pulse regulation and closed-loop calibration. In the initialization stage, the D / A converter writes the conductance value into the memristor array through ±3V pulses, and immediately reads the SLx current after writing to infer the actual conductance value, triggering a rewrite when the error exceeds 5%. In real-time processing, the FPGA dynamically corrects the conductance distribution according to the spatial frequency and grayscale gradient.
[0018] Preferably, the slice processing process of the parallel slice convolution accelerator PSCA includes: dividing the high-resolution image into 4×4 or 8×8 slices, each slice corresponds to an independent memristor subarray; the input voltage of each subarray is generated by the D / A converter based on the slice grayscale value, and the current signal is synchronously output through the transimpedance amplifier; the A / D converter synchronously samples the multi-channel subarray signals to ensure the timing consistency of the global convolution result. It should be further explained that the slice processing process divides the high-resolution image into 4×4 or 8×8 slices, and each slice is bound to an independent subarray. The D / A converter generates a 0-1V voltage signal based on the slice grayscale value, loads it to the BLx line, and the transimpedance amplifier synchronously outputs the multi-channel subarray current signal. The A / D converter samples at a rate of 100MSPS, and the FPGA integrates the results through timestamp matching and weighted fusion algorithm.
[0019] Preferably, the circuit structure of the memristor array computing module comprises: a multi-channel D / A converter and a transimpedance amplifier, respectively connected to the bit line BLx and the source line SLx of the memristor array; The D / A converter regulates the conductance value of the memristor through a voltage pulse signal, and the transimpedance amplifier converts a tiny current signal into a voltage signal; The A / D converter converts the analog signal into a digital signal after filtering by a high-speed operational amplifier and transmits the digital signal back to the XC7Z020 FPGA control module.
[0020] It should be further explained that the memristor array circuit includes multiple D / A converters, transimpedance amplifiers and A / D converters. The D / A converter AD9744 receives the FPGA digital signal through the LVDS interface, generates an analog voltage and loads it to the BLx line to adjust the memristor conductance value. The transimpedance amplifier OPA657 converts the μA current of SLx into a mV voltage, which is sampled by the A / D converter AD9434 as a 12-bit digital signal after the second-order low-pass filter LMH6521 eliminates the noise and then transmitted back to the FPGA. The power management chip LM27762 provides ±5V voltage for the D / A converter to ensure the stability of SET / RESET operation. The multi-channel analog switch ADG732 switches the signal path according to the slice index, and the timing deviation is <1ns.
[0021] Preferably, the process of the Internet of Things module receiving instructions through the CAN bus includes: the XC7Z020 FPGA control module generates an action instruction according to the full connection result of the convolutional neural network; the instruction is transmitted to the robot drive circuit through the CAN_H and CAN_L signals to perform sorting, positioning or obstacle avoidance operations. It should be further explained that the Internet of Things module receives the instruction frame generated by the FPGA through the CAN bus, and the standard extended frame ID is: 0x18FF0000. The FPGA encapsulates the full connection result of the convolutional neural network into an 8-byte data field containing action type, coordinates, and speed parameters, and attaches a CRC-16 check code. The CAN transceiver SN65HVD230 converts the TTL signal into a differential level CAN_H / CAN_L and transmits it to the robot drive circuit. The servo motor converts the coordinates into joint angles according to the inverse kinematics model, and the PWM signal controls the rotation with an accuracy of ±0.1°. In the obstacle avoidance scenario, the ToF sensor VL53L1X detects obstacles in real time, the FPGA plans the path through the A* algorithm, the response time is <10ms, and the gripper pressure sensor feedback data forms a closed-loop control.
[0022] Preferably, the dynamic resource allocation mechanism optimizes the platform energy consumption through a load balancing algorithm, including: when it is detected that the complexity of multiple consecutive frames of images is lower than a preset threshold, the power supply of some memristor subarrays is turned off to reduce power consumption; in a high-load scenario, the redundant subarrays of the parallel sliced convolution accelerator PSCA are enabled for hyperthreading acceleration. It should be further explained that the dynamic resource allocation mechanism DRAM optimizes energy consumption through a load balancing algorithm. The FPGA calculates the image complexity score in real time, including spatial frequency, edge density, and grayscale entropy weighting. When the load is low, that is, when the score is <40, 50% of the subarrays are turned off, and the power consumption is reduced from 5.2W to 3.1W; when the load is high, that is, when the score is >80, the redundant subarrays are enabled, and the hyperthreading accelerates the throughput to 480 frames / second. The power management chip TPS22965 controls the power supply of the subarray, and the sleep power consumption is only 5μW. The closed-loop regulation dynamically adjusts the voltage according to the real-time power consumption MAX44284 monitoring. For example, when the power consumption exceeds 5.5W, the voltage is reduced to 0.8V to ensure that the energy efficiency ratio reaches 96 GOPS / W. The historical learning model optimizes the shard merging strategy and compensates for the loss of computing accuracy to <0.5%.
[0023] (III) Beneficial effects The present invention provides an image processing data analysis platform based on a memristor storage and computing integrated architecture. It has the following beneficial effects: (I) This image processing data analysis platform based on the memristor storage and computing integrated architecture has significantly improved the speed, energy efficiency and reliability of image processing through the deep integration of the memristor storage and computing integrated architecture and dynamic optimization technology. First, the memristor array implements parallel convolution operations in an analog manner, eliminating the data handling bottleneck of the traditional von Neumann architecture. The processing speed has been improved compared to the traditional GPU solution, and the system energy consumption has been reduced. The dynamic weight synchronization network and parallel sharding convolution accelerator further optimize the computing efficiency, adjust the weight distribution and sharding processing strategy through real-time feature feedback, so that the recognition accuracy in complex industrial scenarios is improved, and the throughput has also been broken through to meet the real-time processing requirements of high-resolution images.
[0024] (II) The image processing data analysis platform based on the memristor storage and computing integrated architecture has a dual-modal verification mechanism and a dynamic resource allocation mechanism that enhances the robustness and adaptability of the system. DMVM reduces the misrecognition rate and improves industrial reliability through cross-verification of analog and digital paths; DRAM dynamically starts and stops sub-arrays and redundant resources according to the load, reduces power consumption at low loads, and improves throughput through hyperthreading acceleration at high loads, taking into account both energy efficiency and flexibility. The overall solution shows superior performance in scenarios such as semiconductor testing and intelligent sorting, which is conducive to promoting the intelligent upgrading of the manufacturing industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of the overall framework of the present invention; Figure 2 Schematic diagram of the internal flow of the memristor of the present invention. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] See also Figure 1 and Figure 2 The present invention provides a technical solution: an image processing data analysis platform based on a memristor storage and computing integrated architecture, comprising: The XC7Z020 FPGA control module is connected to the OV5640 camera image acquisition module, the memristor array computing module, and the Internet of Things module; The OV5640 camera image acquisition module is used to collect raw image data and transmit it to the XC7Z020 FPGA control module; the OV5640 camera image acquisition module is connected to the PS end of the FPGA through the I2C bus for camera initialization configuration, including resolution and frame rate settings. The raw image data collected by the camera is transmitted to the PL end of the FPGA through the D0-D7 data line of the DCMI to achieve high-speed data reception. The process is: first, image data is collected, and the FPGA configures the working mode of the OV5640 camera through the I2C bus: 1080P@30fps, and triggers image acquisition. The original image data is transmitted to the DDR3 buffer area of the PS side of the FPGA in YUV format through the DCMI interface to complete frame synchronization and data alignment; then image preprocessing is performed. The PS side of the FPGA converts the YUV data into an 8-bit grayscale image through the weighted average method in the RGB to grayscale algorithm to reduce the complexity of subsequent calculations. Subsequently, the median filtering algorithm, that is, a 3×3 window is used to remove salt and pepper noise and improve image quality. Then, according to the preset ROI area, the image is divided into multiple sub-areas, that is, 512×512 pixel blocks, to facilitate subsequent parallel processing.
[0028] The memristor array computing module includes a D / A converter, a memristor array, a transimpedance amplifier and an A / D converter, which is used to receive the image data processed by the XC7Z020 FPGA control module, complete the convolution operation through parallel analog calculation, and return the result to the XC7Z020 FPGA control module; the memristor array computing module is connected to the D / A converter and the A / D converter through the GPIO pin of the PL end of the FPGA. Among them, the D / A converter receives the digital signal processed by the FPGA, converts it into an analog voltage signal and loads it to the bit line BLx of the memristor array to complete the weight writing and input data loading; the transimpedance amplifier TIA converts the tiny current signal of the array source line SLx into a voltage signal, which is then transmitted back to the FPGA by the A / D converter after filtering.
[0029] The convolution kernel weights are determined through offline training and pre-written into the memristor array in the form of conductance values. The specific process is as follows: FPGA generates a voltage pulse signal of a specific amplitude through the D / A converter: AD9744, including +3V for SET operation and -3V for RESET operation, which is applied to the BLx and SLx cross nodes of the memristor array to adjust its conductance value to the target range, that is, 10μS to 100μS; after writing is completed, the accuracy of the conductance value is verified by reading the current signal to ensure that the weight is consistent with the training model.
[0030] Parallel analog convolution operation is performed, including: the pre-processed image grayscale value (0-255) is converted into an analog voltage signal (0-1V) through a D / A converter and loaded into the BLx line of the memristor array; according to Kirchhoff's current law, the output current of each SLx line is the sum of the product of the BLx voltage and the corresponding conductance value, that is, , realizing the analog calculation of convolution operation. The current signal of SLx is converted into a voltage signal by the transimpedance amplifier OPA657, with the gain set to 1000 times, and then sampled into a 12-bit digital signal by the high-speed A / D converter AD9434 and transmitted to the PL end of FPGA.
[0031] The IoT module receives instructions from the XC7Z020 FPGA control module through the CAN bus and executes predefined actions; that is, the IoT module is connected to the PS end of the FPGA through the CAN bus, receives the recognition result instructions, including the sorting coordinates and action types, and drives the robotic arm to perform the operation; including: the PL end of the FPGA performs pooling and full connection operations on the convolution results returned by the memristor array to generate the final recognition result, wherein the pooling is the maximum pooling, the window is 2×2, and the full connection is based on the pre-trained weight matrix, and the final recognition result includes the defect category and coordinate information; The recognition result is encapsulated as a command frame through the CAN bus protocol and sent to the robot drive circuit by the FPGA's CAN controller. The standard frame ID is 0x18FF0000, and the CAN controller is integrated on the PS side. The robot performs sorting, positioning or obstacle avoidance operations according to the instructions, with a response delay of less than 5ms. The sorting operation includes the action of the pneumatic gripper, the positioning operation includes the rotation angle of the servo motor, and the obstacle avoidance operation is a path planning algorithm.
[0032] The platform also includes a dynamic resource allocation mechanism Dynamic Resource Allocation Module, DRAM, which integrates a real-time task scheduling unit through the XC7Z020 FPGA control module to dynamically allocate computing tasks to the FPGA logic unit or the memristor array computing module according to the complexity of the input image; the real-time task scheduling unit generates task priorities by analyzing the spatial frequency, edge density and grayscale distribution of the image, and allocates the convolution operation of the high-complexity area to the memristor array computing module for parallel simulation calculation, while the low-complexity processing tasks are retained in the FPGA logic unit for execution; The DRAM implementation of the dynamic resource allocation mechanism includes: task scheduling unit design and dynamic power management; wherein the task scheduling unit design process includes: The PS side of the FPGA embeds a lightweight feature extraction algorithm to calculate the spatial frequency and edge density of the image in real time and generate a task priority score ranging from 0 to 100. The spatial frequency is calculated by Fourier transform frequency domain energy ratio, and the edge density is calculated by Sobel operator gradient amplitude statistics. Task allocation strategy: High complexity area, i.e. score > 70: assigned to the memristor array computing module for parallel analog convolution, taking advantage of its high speed and low power consumption; Low-complexity tasks, i.e., scores ≤ 70: Digital convolution is performed by the FPGA logic unit PL to avoid occupying high-performance resources.
[0033] Dynamic power consumption management includes low-load mode and high-load mode. In low-load mode, when the complexity scores of 10 consecutive frames of images are all lower than 50, the FPGA turns off the power supply of some memristor subarrays through the power management chip TPS65218, leaving only 1 / 4 of the array working, and the system power consumption is reduced to 3.8W; in high-load mode, when a sudden high-complexity task is detected, that is, when the score is >90, the redundant subarray is enabled, that is, it is switched through the multiplexing switch to process the fragmented data in parallel, and the single-frame processing delay is compressed to 0.8ms.
[0034] The weights of the memristor array are pre-written in the form of conductance and loaded to the array nodes through voltage signals to achieve storage-computation integrated convolution operations.
[0035] The XC7Z020 FPGA control module is used to segment, compress and grayscale the original image to generate pre-processed data; transmit the pre-processed data to the memristor array computing module and control the robotic arm to perform actions.
[0036] It should be further explained that in the specific implementation process, image segmentation includes: the PL side of the XC7Z020 FPGA, that is, the Programmable Logic configures the image segmentation hardware acceleration module, and divides the original image collected by the OV5640 camera, such as 1920×1080 resolution, into multiple sub-regions through the predefined ROI, that is, the Region of Interest algorithm. The block size is dynamically adjusted according to the complexity of the image content. For example, in the industrial defect detection scenario, high-texture areas, that is, solder joints and edges, use small blocks of 64×64 pixels, and low-texture background areas use large blocks of 256×256 pixels to balance computing resources and accuracy requirements. The block edge pixels are processed by the bilinear interpolation algorithm to avoid image distortion caused by segmentation. The parallel pipeline design of the FPGA allows multiple blocks to be processed simultaneously, and the single-frame segmentation delay is less than 0.2ms, with a resolution of 1080P.
[0037] Image compression uses differential pulse code modulation DPCM technology to predictively encode the grayscale image data. The specific steps are as follows: the current pixel value is predicted based on the left neighbor pixel to generate a residual signal. The residual is mapped to an 8-bit signed integer range (-128~127) to reduce the amount of data. The data is further compressed through Huffman coding, with a compression ratio of up to 2:1 (measured average). The PL side of the FPGA integrates dedicated compression logic units, such as LUT-based encoders, with a compression throughput of 1.5Gbps to ensure real-time performance.
[0038] Grayscale processing includes: RGB to grayscale formula: using the brightness formula recommended by the International Telecommunication Union ITU: Y=0.299R+0.587G+0.114B; where R, G, and B are the three-channel values of the original pixel, and Y is the grayscale value. Parallel multiplication and addition operations are implemented through the FPGA's DSP Slice, and the single pixel processing cycle is 1 clock cycle, that is, at a 100MHz main frequency, the grayscale conversion of a single frame of 1080P image takes 1.08ms. After grayscale conversion, real-time median filtering and a 3×3 window are embedded to eliminate sensor noise and improve the accuracy of subsequent convolution operations.
[0039] The preprocessed image data is transmitted to the D / A converter AD9744 of the memristor array computing module through the FPGA's PL-side GPIO pins, such as the LVDS differential pair of Bank34. The 8-bit grayscale value (0-255) is converted to a 14-bit digital signal (0-16383) through linear scaling to match the input accuracy of the D / A converter. The D / A converter converts the digital signal into an analog voltage signal (0-1V) and loads it to the bit line (BLx) of the memristor array. The FPGA generates a synchronous clock signal (CLK+ / -) to control the sampling timing of the D / A converter, ensuring strict synchronization between data loading and convolution operations.
[0040] During the system initialization phase, the trained convolution kernel weights are written into the memristor array in the form of conductance values through the D / A converter. +3V pulses are used to increase conductance (SET operation), and -3V pulses are used to reduce conductance (RESET operation). A charge pump chip (LM27762) provides a stable voltage source. In real-time processing, the conductance value is dynamically corrected according to image features in combination with the dynamic weight synchronization network DWS-Net described below to optimize convolution kernel matching.
[0041] Control the robot to perform actions: The PL side of the FPGA performs maximum pooling (window 2×2) and full connection operations on the convolution results returned by the memristor array. The pre-trained fully connected layer weights are stored in the Block RAM of the FPGA, and matrix multiplication is completed through the parallel multiplier (DSP48E1 Slice) to output classification probability or coordinate information. The recognition results (such as defect category, target coordinates) are encapsulated as 32-bit data packets (including check bits) and transmitted to the PS side (Processing System) of the FPGA.
[0042] The robot control command transmission adopts the standard CAN 2.0B extended frame format, the frame ID is 0x18FF0000, and the data field includes the action type (1 byte), coordinates X / Y (2 bytes each), and speed parameters (1 byte). The PS side of the FPGA integrates a CAN controller (IP Core), supports a 1Mbps communication rate, and the command transmission delay is less than 0.1ms.
[0043] The IoT module parses the CAN command and drives the servo motor to rotate to the target angle (accuracy ±0.1°) through the PWM signal. In complex scenarios, such as multi-target sorting, the FPGA presets the A* algorithm to generate the optimal path to avoid robot arm collision. A force sensor and a visual feedback module are installed at the end of the robot arm to send back the execution status (such as clamping force, position deviation) to the FPGA in real time. If an abnormality is detected (such as clamping failure), the FPGA triggers the recalculation process and updates the robot arm instructions to ensure operational reliability.
[0044] This implementation process fully demonstrates the advantages of the memristor-based memory and computing integrated architecture in terms of speed, energy consumption, and reliability through hardware-algorithm co-optimization, dynamic resource allocation, and closed-loop control mechanisms.
[0045] The memristor array computing module includes a Dynamic Weight Synchronization Network (DWS-Net). The DWS-Net includes a lightweight feature extraction unit embedded in the XC7Z020 FPGA control module, which analyzes the spatial frequency and texture distribution of the input image in real time to generate a feature vector. Based on the feature vector, the conductance values of different nodes in the memristor array are dynamically adjusted through a D / A converter to achieve adaptive allocation of convolutional kernel weights. The DWS-Net optimizes weight matching through a closed-loop feedback mechanism to reduce redundant computing energy consumption.
[0046] It should be further noted that in the specific implementation process, the architecture design of the DWS-Net aims to dynamically adjust the weight allocation of the memristor array through real-time image feature feedback to address the problem of insufficient generalization ability caused by traditional fixed weights. Its implementation process includes the following key steps: Design and implementation of the lightweight feature extraction unit: The input image is decomposed in the frequency domain using the Fast Fourier Transform (FFT) to calculate the proportion of high-frequency energy, such as the component ratio exceeding 50Hz in the frequency domain, to characterize the complexity of image details.
[0047] Formula: ; where F(u,v) is the two-dimensional Fourier transform result of the image.
[0048] Texture features are extracted based on the Gray-Level Co-Occurrence Matrix (GLCM), including parameters such as contrast, energy, and correlation, and a texture complexity score is comprehensively generated.
[0049] Contrast calculation: ; where P(i,j) is the joint probability of gray levels i and j in the gray-level co-occurrence matrix.
[0050] The PL (Programmable Logic) side of the FPGA integrates an FFT IP core (such as Xilinx FFT v9.0), which supports real-time transformation of 512×512 pixel blocks (with a delay < 0.1ms). The gray-level co-occurrence matrix is stored in the Block RAM of the FPGA, and parameters such as contrast and energy are calculated in parallel using DSPSlice. The single-frame processing time is 0.3ms. Parameters such as the proportion of high-frequency energy (1D), texture contrast (1D), and energy (1D) are normalized to 8-bit integers and combined into a 3D feature vector for subsequent weight adjustment.
[0051] Dynamic adjustment mechanism of the memristor array conductance value: The convolutional kernel weights are determined through offline training, and each weight value wij Corresponding to the memristor cross point BL i , SL j The target conductivity value G ij , the mapping formula is: ; where k is the proportionality coefficient, which is determined by the conductance range of the memristor, G min is the minimum conductance value to prevent the conductance from returning to zero. Including: Step 1: Feature vector analysis: The PS end of the FPGA receives the feature vector and generates a conductance adjustment instruction according to the preset rules. Including: If the high-frequency energy ratio is >60%, it is determined that the image is rich in details, and the weight of the edge detection convolution kernel needs to be enhanced, that is, the corresponding conductance value is increased. If the texture contrast is <30, it is determined that the image is smooth, and the weight of the high-frequency filter convolution kernel needs to be reduced, that is, the corresponding conductance value is reduced. Step 2: Voltage pulse generation: A voltage pulse signal of a specific amplitude is generated through a D / A converter (AD9744), including SET operation and RESET operation; the SET operation includes a +3V pulse (duration 10ns) to increase the conductance value. The RESET operation includes a -3V pulse (duration 10ns) to reduce the conductance value. The pulse amplitude and width are dynamically adjusted according to the conductance adjustment amplitude. For example, when the conductance needs to be increased by 20%, the pulse amplitude is increased to +3.5V. Step 3: Conductance value writing verification: Immediately after writing, read the SL line current through the transimpedance amplifier (TIA), infer the actual conductance value, and compare it with the target value. If the error exceeds 5%, trigger a second write or alarm to ensure the accuracy of the weight. Closed-loop feedback mechanism optimizes weight matching: Through the confidence score output by the dual-modal verification mechanism DMVM, such as the difference in Euclidean distance. Real-time acquisition of the power consumption of the memristor array (through the current sensor MAX44284) to calculate the energy consumption per unit frame. If the confidence score is lower than the threshold, such as <90%, it is determined that the current weight matching is insufficient, triggering the feature extraction unit to re-analyze the image and generate a new conductance adjustment instruction. If the energy consumption per unit frame exceeds the preset value, such as >1mJ, it is determined that the weight is redundant and the conductance value of the non-critical convolution kernel is automatically reduced. Optimize the conductance adjustment rule based on historical data (such as feature vectors of 10 consecutive frames), for example, predict the optimal weight distribution through a linear regression model. Run the lightweight feature extraction algorithm (FFT, GLCM calculation), occupying 20% of the ARM Cortex-A9 core resources. Conductance adjustment control logic (occupies 10% of LUT resources). D / A converter interface and transimpedance amplifier signal processing (occupies 15% of GPIO pins). From feature extraction to completion of conductance adjustment, the whole process delay is <0.5ms, meeting the 30fps real-time processing requirements. Dynamic weight adjustment reduces redundant calculations by 40%, and the overall system power consumption is reduced to 3.8W (compared to 4.5W for static weight scheme). In the industrial defect detection scenario, the false recognition rate is reduced from 0.1% to 0.04%.
[0052] The memristor array computing module integrates a parallel sliced convolution accelerator (PSCA), namely: the XC7Z020 FPGA control module divides the input image into N×N independent sub-areas, which are regarded as slices, and each slice is assigned to an independent sub-array of the memristor array; each sub-array completes the local convolution operation through parallel analog computing, and synchronously outputs the current signal through the multiplexing technology of the transimpedance amplifier; the A / D converter converts the analog signals of the multiple sub-arrays into digital signals, and the XC7Z020 FPGA control module performs global feature fusion.
[0053] It should be further explained that in the specific implementation process, the parallel sliced convolution accelerator (PSCA) aims to break through the physical scale limitation of a single array by dividing the high-resolution image into independent sub-regions (slices) and utilizing the parallel computing capability of the memristor sub-array to significantly improve the processing speed. The implementation process includes the following key steps: S1. Image slicing strategy and hardware implementation: The slicing size is dynamically determined according to the resolution and complexity of the input image. This includes: For high-resolution images, such as 3840×2160 pixel images captured by 4K industrial cameras, 8×8 slicing is used, with each slicing being 480×270 pixels, to ensure efficient use of subarray computing resources. For high-texture areas (such as solder joints of electronic components), 4×4 slicing is used, with each slicing being 240×135 pixels, to improve the accuracy of local feature extraction.
[0054] An overlapping tile strategy (Overlap=10%) is used to avoid edge information loss, including: adjacent tiles overlap 24 pixels in the horizontal and vertical directions (480×270 tiles), and smooth transition is achieved through interpolation algorithm (bicubic interpolation).
[0055] The FPGA slicing logic includes: the PL (Programmable Logic) of the XC7Z020 FPGA integrates an image slicing hardware acceleration module, which is implemented through pixel address mapping and parallel data distribution steps; pixel address mapping maps the pixel coordinates of the original image to the slicing coordinate system and generates a slicing index table (stored in Block RAM). Parallel data distribution transmits the slicing data to the D / A converter interface of multiple memristor subarrays simultaneously through the FPGA's AXI-Stream interface.
[0056] Using a pipeline architecture, the single-frame slice processing delay is less than 0.3ms (4K resolution).
[0057] S2. Parallel computing process of memristor subarrays: The pixel grayscale value (0-255) of each slice is converted into an analog voltage signal (0-1V) through linear scaling, and loaded to the bit line BLx of the corresponding subarray by the D / A converter (AD9744); a negative voltage (-1V) is generated by the charge pump chip (LM27762) to represent the negative weight value of the convolution kernel. The memristor conductance value of each subarray is pre-written according to the slice task type (such as edge detection, texture classification) to ensure the functional independence of the subarray.
[0058] According to Kirchhoff's current law, the output current of the source line SLx of each sub-array is: ; Among them, V BLi is the bit line voltage, G i is the conductance value of the corresponding memristor node.
[0059] The SLx line of each sub-array is connected to an independent transimpedance amplifier (TIA, such as OPA657), and the gain is set to 1000 times to convert the μA-level current signal into a mV-level voltage signal. The sub-array signal currently being sampled is selected through a multi-channel analog switch (such as ADG732) to ensure the timing synchronization of multiple outputs. The high-speed A / D converter (AD9434) synchronously samples the voltage signals of multiple sub-arrays with 12-bit accuracy, and the sampling rate is set to 100MSPS to ensure that the data is distortion-free. The PL end of the FPGA matches the convolution results of each slice through the timestamp to eliminate the time offset of the slice processing. The weighted average algorithm is used to fuse the slice results, and the weight is dynamically determined by the slice complexity, including a high-texture slice weight of 0.7 and a low-texture slice weight of 0.3.
[0060] S3. Hardware circuit design and performance optimization: Metal wires are used to isolate BLx and SLx of different sub-arrays to avoid signal crosstalk. Each sub-array contains 128×128 memristor cross nodes. Each sub-array is powered independently. The power switch (such as TPS22965) is controlled by FPGA to turn off idle sub-arrays when the load is low, reducing power consumption by 30%. ADG732 analog switches are used to support dynamic switching of 32 SLx signals with a switching delay of less than 10ns. A high-speed op amp (LMH6521) is embedded at the output of the transimpedance amplifier to form a low-pass filter (cutoff frequency 50MHz) to eliminate high-frequency noise.
[0061] The XC7Z020 FPGA control module deploys a dual-modal verification mechanism DMVM: constructs a memristor array analog convolution operation path for the first mode and an FPGA digital logic verification path for the second mode; compares the calculation results of the two modes through a confidence assessment module. If the difference exceeds a preset threshold, a secondary calculation or an alarm is triggered; the dual-modal verification mechanism controls the misrecognition rate to below 0.05% in a noisy environment.
[0062] It should be further explained that in the specific implementation process, the dual-modal verification mechanism (DMVM) solves the potential noise interference and precision drift problems in memristor analog calculation by constructing cross-validation of heterogeneous computing paths (analog and digital), ensuring high-reliability identification in industrial scenarios. Its implementation process includes the following steps: T1. Memristor array simulates convolution operation path: The trained convolution kernel weights are pre-written into the memristor array in the form of conductance values. Edge detection convolution kernel , the corresponding memristor node conductance value is mapped proportionally (negative weight is represented by negative voltage). The preprocessed image grayscale value is converted into an analog voltage signal (0-1V) by a D / A converter (AD9744) and loaded into the bit line (BLx) of the memristor array. Based on Kirchhoff's law, the output current of the source line (SLx) is the sum of the product of the bit line voltage and the conductance value, realizing the analog domain convolution operation. The transimpedance amplifier (TIA) converts the μA-level current signal into a mV-level voltage signal, which is sampled into a 12-bit digital result by the A / D converter (AD9434) and transmitted to the PL end of the FPGA. Combined with the parallel slice convolution accelerator (PSCA), the image is divided into sub-regions for parallel processing, and the single-frame analog convolution delay is compressed to 0.8ms (1080P resolution). A low-pass filter (LMH6521, cut-off frequency 50MHz) is embedded before the A / D conversion to eliminate high-frequency noise interference and ensure signal stability.
[0063] T2.FPGA digital logic verification path: The PL side of the FPGA integrates a digital convolution IP core, supports multiple convolution kernel sizes such as 3×3 and 5×5, and implements parallel multiplication and addition operations through the hardware description language (HDL). After the original image data is grayed, it is distributed to the analog path and the digital path at the same time to ensure input consistency. 16-bit fixed-point numbers are used to represent weights and intermediate results to balance accuracy and resource usage (occupying 15% of DSP Slice resources).
[0064] Among them, the digital convolution algorithm process is: traverse the image with a 3×3 window, and perform multiplication and addition operations in each window: ; A zero-filling strategy is used to process image edges to avoid information loss. Through the FPGA pipeline architecture, the single-frame digital convolution delay is 1.2ms (1080P resolution).
[0065] T3. Design and implementation of confidence assessment module: The output results of the analog path and the digital path are matched by timestamps to ensure the comparison of convolution results in the same image area. The Euclidean distance is used to quantify the difference between the two modal results: ; Among them, S iis the analog path output, D i is the digital path output, n is the result dimension. According to the requirements of industrial scenarios, the initial threshold is set to 0.1 (normalized distance). The distance mean μ and standard deviation σ of 100 consecutive frames are counted, and the threshold is dynamically adjusted to μ+2σ. If the environmental noise is detected to be enhanced, such as the current sensor monitoring the power supply fluctuation exceeding 5%, the threshold is temporarily reduced to 0.08; if the distance exceeds the threshold, the FPGA control module starts the following process: t1. Path switching: reallocate the current slice data to another group of memristor sub-arrays or digital logic units for review; t2. Result arbitration: If the review result is consistent with the original result, it is determined to be environmental noise interference and the threshold is updated; if it is inconsistent, it is determined to be a hardware failure and an alarm is triggered.
[0066] Output high and low level signals through GPIO pins, such as 3.3V high level indicates serious error, driving external LED or buzzer to prompt the operator.
[0067] The weight adjustment process of the dynamic weight synchronization network DWS-Net includes: generating a voltage pulse signal through a D / A converter to adjust the initial conductance value of the memristor array; dynamically correcting the conductance distribution according to the real-time image characteristics to match the spatial frequency and grayscale gradient of the input data; the weight adjustment process realizes the weighted summation of the current signal through Kirchhoff's current law to complete the convolution operation.
[0068] It should be further explained that in the specific implementation process, by combining real-time image features with the dynamic correction of the memristor conductance value, the matching accuracy of the convolution kernel is optimized, thereby improving the accuracy and energy efficiency of image recognition. The specific implementation process is as follows: E1. Writing and verification of initial conductance values: The convolution kernel weights are determined by offline training (such as fine-tuning the VGG16 model on an industrial defect dataset). Each weight value w ij Corresponding to the memristor cross point (BL i , SL j) The target conductivity value G ij , the mapping formula is: ; Where k is the proportionality coefficient, which is determined by the conductance range of the memristor, G min is the minimum conductance value to prevent the conductance from returning to zero.
[0069] Among them, the conductance writing process generates a +3V voltage pulse (duration 10ns) through a D / A converter (AD9744), which is applied to the target memristor node to increase its conductance value; a -3V voltage pulse (duration 10ns) is generated to reduce the conductance value; the pulse amplitude is dynamically adjusted according to the target conductance deviation. For example, when the conductance needs to be increased by 20%, the pulse amplitude is increased to +3.5V.
[0070] Conductance value verification and calibration includes: reading the SL line current I through a transimpedance amplifier (TIA) immediately after writing SLx , inversely calculate the actual conductivity value ; If the actual conductivity value deviates from the target value by more than 5%, a secondary write or alarm is triggered, that is, a high-level signal is output through the GPIO pin of the FPGA. All conductivity adjustment records are stored in the Block RAM of the FPGA for subsequent dynamic correction of historical data analysis.
[0071] E2. Real-time image feature analysis and dynamic conductivity correction: The input image is fast Fourier transformed through the FPGA FFT IP core to calculate the high-frequency energy ratio (the ratio of components above 50Hz). The formula is: ; The Sobel operator is used to calculate the image gradient amplitude, and the percentage of pixels with gradient amplitude greater than the threshold is counted to characterize the edge density. The high-frequency energy percentage (8 bits), edge density (8 bits), and average grayscale value (8 bits) are combined into a 24-bit feature vector and transmitted to the conductivity adjustment control unit; in high-frequency dominant scenarios, that is, the high-frequency energy percentage > 60%, the conductivity value of the edge detection convolution kernel is enhanced (such as increasing by 15%), and the conductivity value of the smoothing filter kernel is suppressed (reducing by 10%). In low-contrast scenarios, that is, the edge density < 20%, the weight of the high-frequency filter kernel is reduced, and the weight of the texture enhancement kernel is increased. Combined with the results of dual-modal verification, if the simulated path confidence is lower than the threshold, the conductivity value is triggered to be readjusted, and the historical feature vectors and conductivity adjustment records are analyzed through a linear regression model to predict the optimal weight distribution.
[0072] E3. Convolution operation based on Kirchhoff's current law: The grayscale value of the preprocessed image is mapped into a voltage signal V through a D / A converter. BLi , loaded to the bit line BLi of the memristor array.
[0073] Current calculation: According to Kirchhoff's current law, the output current of the source line SLx is: Among them, G iThe dynamically adjusted conductance value directly corresponds to the convolution kernel weight; the TIA (gain 1000 times) converts the μA-level current signal into an mV-level voltage signal, which is filtered by a high-speed op amp (LMH6521) and input into the A / D converter; the A / D converter (AD9434) samples the voltage signal with 12-bit accuracy, generates a digital convolution result, and transmits it to the FPGA for pooling and full-connection operations.
[0074] The slice processing process of the parallel slice convolution accelerator PSCA includes: dividing the high-resolution image into 4×4 or 8×8 slices, each slice corresponds to an independent memristor subarray; the input voltage of each subarray is generated by the D / A converter based on the slice grayscale value, and the current signal is synchronously output through the transimpedance amplifier; the A / D converter synchronously samples the multi-channel subarray signals to ensure the timing consistency of the global convolution result.
[0075] It should be further explained that in the specific implementation process, the parallel calculation of high-resolution images is achieved through slice processing, which includes slice segmentation, sub-array voltage generation, transimpedance amplifier synchronous output and global feature fusion. The implementation process is as follows: Dynamic sharding rule settings include: 4×4 Slicing Mode: Suitable for highly complex images, such as micron-level defect detection, the image is divided into 4×4 grids, each slice size is , for example, each slice of a 4K image is 960×540 pixels; 8×8 slicing mode: suitable for low-complexity scenes, such as logistics sorting with a smooth background. The slicing size is , for example, each slice of a 4K image is 480×270 pixels; Adjacent tiles overlap by 10% in the horizontal and vertical directions, such as 960×540 tiles overlap by 96 pixels, and the edges are smoothed by the bicubic interpolation algorithm to avoid information loss; The FPGA slicing logic is that the PL side of the FPGA receives the original image data through the AXI-Stream interface, generates a slicing index table according to the slicing rules, and stores it in the Block RAM; the slicing data is simultaneously transmitted to the D / A converter interface of multiple memristor sub-arrays through multiple DMA channels, supports 64-channel slicing parallel processing, adopts a pipeline architecture, and the single-frame slicing processing delay is <0.4ms, that is, 4K resolution.
[0076] Subarray input voltage generation and loading include: the 8-bit grayscale value (0-255) of each pixel in the slice is converted into an analog voltage through the formula: ; Among them, V maxis the maximum output voltage of the D / A converter, for example, 1V. A negative voltage (-1V) is generated by a charge pump chip (LM27762) to represent the conductance node corresponding to the negative weight of the convolution kernel.
[0077] The D / A converter configuration uses a 14-bit high-precision D / A converter (AD9744), which supports a 0-1V output range and a conversion rate of 100MSPS; the FPGA generates a synchronous clock signal (CLK+ / CLK-) to control the sampling timing of all sub-array D / A converters to ensure strict synchronization of slice data loading; the D / A converter of each sub-array is independently controlled through the FPGA's GPIO pins, supporting dynamic allocation of slices to idle sub-arrays.
[0078] The process of synchronously outputting the current signal of the transimpedance amplifier includes: using a low-noise operational amplifier (OPA657) with a gain set to 1000 times to convert the μA-level current signal into a mV-level voltage signal, and selecting R according to the current range, such as 1μA-100μA. f =1MΩ, output voltage V out = I SL· R f ; A second-order low-pass filter (cutoff frequency 50MHz) is embedded at the output of the TIA to eliminate high-frequency noise; a 32-channel analog switch (ADG732) is used to switch the SLx signals of the subarray to the A / D converter in the order of the slice index. The FPGA adds a timestamp (based on the system clock count) to the SLx signal of each slice to ensure data alignment during global fusion.
[0079] For A / D converter synchronous sampling and global feature fusion: a 12-bit high-speed A / D converter (AD9434) is used with a sampling rate of 100MSPS to support multi-channel synchronous sampling; the PL end of the FPGA generates a synchronous sampling clock (SYNC_CLK) to drive all A / D converters to start sampling at the same time, with a deviation of <1ns; the FPGA reorganizes the convolution results of each slice according to the original spatial position through the timestamp to restore the complete feature map; the high-complexity slice weight is set to 0.7, and the low-complexity slice weight is 0.3, the formula is: ; For edge detection tasks, the maximum response value of the corresponding position of each slice is selected. The formula is: ; The BLx and SLx of each sub-array are routed through independent metal layers to avoid signal crosstalk. The sub-array power supply is independently controlled through load switches (TPS22965), and 50% of the sub-arrays are turned off when the load is low, reducing power consumption by 35%. The D / A and A / D converters use LVDS differential interfaces to enhance the ability to resist common-mode noise. An independent ground loop is designed for each sub-array to reduce the impact of ground bounce noise.
[0080] The circuit structure of the memristor array computing module includes: a multi-channel D / A converter and a transimpedance amplifier, which are respectively connected to the bit line BLx and the source line SLx of the memristor array; The D / A converter controls the conductance of the memristor through a voltage pulse signal, and the transimpedance amplifier converts the tiny current signal into a voltage signal; The A / D converter converts the analog signal into a digital signal after filtering through a high-speed operational amplifier and transmits it back to the XC7Z020FPGA control module.
[0081] It should be further explained that in the specific implementation process, the circuit structure design of the memristor array computing module includes the coordinated work of multiple D / A converters, memristor arrays, transimpedance amplifiers and A / D converters to achieve storage and computing integrated analog convolution operations. The implementation process is as follows: D / A converter selection and configuration: The 14-bit high-precision D / A converter AD9744 is used, which supports 0-1V analog voltage output and a conversion rate of 100MSPS to meet the high-speed writing requirements of the memristor array. The bit line (BLx) of each memristor subarray is connected to an independent D / A converter channel to support parallel data loading. For example, a 128×128 memristor array requires 128 D / A converters. The D / A converter is controlled by FPGA to generate ±3V voltage pulses (duration 10ns), which are used to increase (SET) or decrease (RESET) the conductance value of the memristor. The preprocessed image grayscale value (0-255) is converted to 0-1V analog voltage through linear scaling and loaded to the BLx line to represent the input data of the convolution operation.
[0082] The PL side of the FPGA generates a synchronous clock signal (CLK+ / CLK-) to drive all D / A converters to sample synchronously, ensuring the timing consistency of multiple voltage pulses. The digital input of the D / A converter is connected through the Bank34 LVDS differential pair of the FPGA (such as B34_L4_P / N to B34_L10_P / N), supporting high-speed parallel data transmission.
[0083] The physical structure and conductance control of the memristor array include: the memristor array is composed of BLx (bit line) and SLx (source line) intersections, and each intersection integrates a memristor unit, whose conductance value represents the convolution kernel weight. The conductance value is adjusted to a preset range (such as 10μS-100μS) by applying SET / RESET pulses through the D / A converter. According to the real-time image features, that is, the dynamic weight synchronization network, the conductance value of a specific node is dynamically corrected.
[0084] When the input voltage is applied to the BLx line, the output current of the SLx line is the sum of the products of each BLx voltage and the corresponding conductance value, that is: ; This current signal directly implements the convolution operation in the analog domain.
[0085] The current-to-voltage conversion of the transimpedance amplifier (TIA) includes: using the low-noise transimpedance amplifier OPA657, with a gain-bandwidth product of 1.6GHz and an input bias current of 1pA, which is suitable for μA-level current signal amplification. Set the feedback resistor R according to the current range (1μA-100μA) f =1MΩ, output voltage V out = I SLx· R f ; Connect a 0.1μF ceramic capacitor to the power pin of OPA657 to suppress high-frequency noise; Among them, the SLx signal line uses differential routing to reduce electromagnetic interference.
[0086] The 32-channel analog switch ADG732 is used to switch the SLx signal to the A / D converter according to the slice index, with a switching delay of <10ns. The FPGA adds a timestamp to each SLx signal to ensure data alignment during global feature fusion.
[0087] A / D conversion and high-speed op amp filtering include: A / D converter configuration, using 12-bit high-speed A / D converter AD9434, sampling rate 100MSPS, supporting differential input, dynamic range 70dB. Connect high-speed op amp LMH6521 to the front end of the A / D converter, configured as a second-order low-pass filter (cut-off frequency 50MHz) to eliminate high-frequency noise. The 0-1V signal output by the TIA is adjusted to the input range of the AD9434 (0-2V) through a resistor divider network. The differential output of the AD9434 (D0+ / - to D11+ / -) is connected to the PL terminal through the Bank35 LVDS pins (B35_L2_P / N to B35_L14_P / N) of the FPGA. The FPGA generates a synchronous sampling signal (SYNC) to ensure strict synchronization between the A / D converter and the D / A converter, with a sampling deviation of <1ns.
[0088] To this end, the system collaborative workflow includes: Data loading and convolution operation: Step 1: FPGA distributes the pre-processed image slice data to multiple D / A converters to generate corresponding BLx voltage signals; Step 2: The memristor array completes the analog convolution calculation based on the conductance value, and the SLx output current signal is converted into a voltage signal through the TIA; Step 3: The A / D converter samples the filtered analog signal into a 12-bit digital result and sends it back to the FPGA for global fusion and post-processing.
[0089] Closed-loop verification and calibration: Regularly read the SLx current to infer the actual conductance value, compare it with the target value, and trigger a second write when the deviation exceeds 5%; through the noise power spectral density analysis of AD9434, dynamically adjust the filter parameters to ensure signal integrity.
[0090] The process of the IoT module receiving instructions through the CAN bus includes: the XC7Z020 FPGA control module generates action instructions based on the fully connected results of the convolutional neural network; the instructions are transmitted to the robotic arm drive circuit through CAN_H and CAN_L signals to perform sorting, positioning or obstacle avoidance operations.
[0091] It should be further explained that in the specific implementation process, the complete process of the IoT module receiving instructions through the CAN bus includes FPGA generating action instructions, CAN protocol encapsulation, data transmission and robot arm execution operations. Its implementation process covers hardware circuit design, communication protocol implementation and closed-loop control mechanism, as follows: The process of FPGA generating action instructions: The PL side of the FPGA integrates the pre-trained fully connected layer weight matrix (stored in Block RAM), completes matrix multiplication through the parallel multiplier (DSP48E1 Slice), and outputs classification probability or coordinate information. Including: For defect detection tasks, the output is the defect category, such as "false solder joint" and "short circuit" and its coordinates (x, y) in the image, and the classification result and coordinate information are encoded into a 32-bit data packet in the format of: Data packet = [category code (4 bits) | coordinate X (12 bits) | coordinate Y (12 bits) | check bit (4 bits)]; Among them, the check mechanism uses the CRC-4 check algorithm to ensure the integrity of data transmission.
[0092] For command priority scheduling, the PS side of the FPGA maintains the command queue and dynamically adjusts the priority according to the urgency of the task. Including: High priority: Robotic arm obstacle avoidance command (response delay <2ms). Low priority: Sorting or positioning command (response delay <5ms). Interrupt processing: If an emergency event (such as collision risk) is detected, the high priority command is immediately triggered to be inserted into the queue head.
[0093] For the implementation of the CAN bus communication protocol, in the hardware circuit design, the CAN controller is configured as follows: the PS side of the XC7Z020FPGA integrates the CAN controller IP core (such as Xilinx AXI CAN), supports the CAN 2.0B protocol, and the communication rate is 1Mbps; for the CAN transceiver in the physical layer interface: the TI SN65HVD230 chip is used to convert the TTL signal of the FPGA into a differential signal (CAN_H / CAN_L). Connect 120Ω terminal resistors at both ends of the CAN bus to suppress signal reflection. The CAN controller of the FPGA is connected to the robot drive circuit through the MIO pin (B1_MIO41: CAN_H, B1_MIO42: CAN_L). Among them, a frame of data is sent every 5ms to avoid bus congestion; if the ACK response from the robot is not received, the FPGA automatically resends the command after 1ms, and retries up to 3 times.
[0094] Robotic arm drive and action execution, after the robot arm drive circuit parses the CAN command, it generates a PWM signal through a microcontroller such as STM32F407 to control the servo motor to rotate to the target angle. The coordinates (x, y) are converted into joint angles (θ) through the inverse kinematics model. 1 ,θ 2 ). Use photoelectric encoders (such as OMRON E6B2-CWZ6C) to feedback the actual angle and achieve closed-loop control (accuracy ±0.1°). Synchronously control multiple servo motors through the EtherCAT bus to ensure smooth movement of the end effector of the robot arm.
[0095] Sorting and positioning operations: According to the weight of the target object (matched by the preset parameter library), the air pump pressure (0-0.6MPa) is adjusted. The surface of the gripper is covered with high friction coefficient silicone. The pressure sensor Honeywell FSS1500NGT monitors the gripping force in real time and triggers an alarm when it exceeds the limit. The micro camera OV9281 is integrated at the end of the robotic arm, and the real-time image sent back by the FPGA is used for secondary calibration of the position, with a positioning error of <0.05mm.
[0096] For obstacle avoidance path planning, the ToF sensor (such as VL53L1X) is used to measure the distance of surrounding obstacles in real time (detection range 0-4m, accuracy ±1cm). The path planning algorithm uses the A* algorithm to preset the grid map in the FPGA and generate the optimal path based on the obstacle position. If a sudden obstacle is detected, the path is replanned through the RRT* (Rapidly Extended Random Tree) algorithm with a response time of <10ms. Mechanical limit switches are set at the joints of the robot arm to forcibly limit the range of motion and prevent hardware damage.
[0097] Closed-loop feedback and exception handling are performed, that is, the robot drive circuit collects the following parameters in real time and transmits them back to the FPGA through the CAN bus: joint angle, end position, clamping force, power supply voltage, and temperature.
[0098] The data frame format is return frame = [status code (4 bits) | joint angle (24 bits) | clamping force (12 bits) | CRC-8]; if it is detected that the clamping force exceeds the limit, that is, >50N, or the temperature is too high, that is, >80℃, the robot arm immediately stops moving and sends an emergency alarm frame (ID: 0x18FF0001); after receiving the alarm, the FPGA triggers the system self-check process, resets the faulty module and restores the initial state.
[0099] The dynamic resource allocation mechanism optimizes the platform energy consumption through a load balancing algorithm, including: when it is detected that the complexity of multiple consecutive frames of images is lower than a preset threshold, the power supply of some memristor subarrays is turned off to reduce power consumption; in high-load scenarios, the redundant subarrays of the parallel sliced convolution accelerator PSCA are enabled for hyper-threading acceleration.
[0100] It should be further explained that in the specific implementation process, the enabled state of the memristor subarray is dynamically adjusted through the load balancing algorithm to optimize the system energy consumption, including shutting down some subarrays to reduce power consumption when the load is low, and enabling redundant subarrays for hyperthreading acceleration to improve throughput when the load is high. The implementation process covers load monitoring, resource scheduling and hardware control, as follows: The PS (Processing System) of the FPGA is embedded with a lightweight algorithm to calculate the key features of the input image in real time: the proportion of high-frequency energy is calculated through fast Fourier transform (FFT), and the percentage of energy of components above 50Hz in the total energy.
[0101] The Sobel operator is used to count the proportion of pixels whose gradient amplitude exceeds the threshold, and the entropy value of the image grayscale histogram is calculated to represent the information complexity; the above features are normalized to a complexity score (Score) of 0-100, and the formula is: Score = 0.4×high-frequency energy ratio+0.3×edge density+0.3×grayscale entropy; among them, the load state judgment includes: if the average score of 10 consecutive frames of images is <40, it is judged as low load, if the score of a single frame is>80 or the average score of 3 consecutive frames is>70, it is judged as high load.
[0102] During the energy consumption optimization in low-load mode, for the sub-array shutdown strategy, the slice mapping relationship is: according to the slice processing mode, such as 4×4 or 8×8 slices, each slice is bound to a specific sub-array group, wherein the dynamic shutdown process includes: FPGA sends a control signal to the power management chip TI TPS22965 through the GPIO pin to turn off the power supply of the idle sub-array, leaving only the core sub-array, such as 1 / 4 of the total array working, and the remaining sub-arrays enter a sleep state.
[0103] In the power management circuit, the power supply of each sub-array is controlled by an independent load switch, which supports nanosecond start and stop. The static power consumption of the dormant sub-array is reduced from 50mW to 5μW, a reduction of 99%.
[0104] During task redistribution and computational compensation, slice merging is performed, including: merging the slices originally assigned to the shut-down subarray into the adjacent active subarray, maintaining processing power by increasing the slice size (e.g., from 8×8 to 4×4), increasing the input voltage of the active subarray proportionally (e.g., 1.2 times) to compensate for the increased computational effort after slice merging, and weighting the results through a global fusion algorithm to avoid precision loss.
[0105] For hyperthreading acceleration in high-load mode, during the redundant sub-array enabling mechanism, the system reserves 20% of the memristor sub-array as redundant resources, which are in a dormant state under normal circumstances.
[0106] The dynamic activation process includes: after the FPGA detects a high load, it wakes up the redundant sub-array through the power management chip; the redundant sub-array is connected to the computing network and processes the same slice data in parallel with the original slice-bound sub-array (hyper-threading mode).
[0107] Hyperthreading acceleration implementation includes: copying the same slice data to multiple sub-arrays, extracting multi-dimensional features in parallel through different convolution kernel weights (such as edge detection kernel and texture enhancement kernel), selecting the maximum value of each sub-array output as the final feature, enhancing the edge detection effect, dynamically allocating weights according to the historical accuracy of the sub-array, improving the reliability of the fusion result, coordinating the calculation timing of redundant sub-arrays through the global clock signal of FPGA (such as 200MHz), avoiding data conflicts, and allocating high-complexity slices to redundant sub-arrays first to ensure the processing speed of critical tasks.
[0108] During the load balancing algorithm and energy consumption control, the algorithm logic is designed: input real-time complexity score (Score), historical energy consumption data, and sub-array status; output sub-array start and stop instructions, voltage adjustment parameters, and task allocation strategy.
[0109] The core steps include: first calculate the current frame score and update the sliding window mean (window size = 10 frames); if the mean is <40 and lasts for 5 frames, trigger the low load mode and shut down 50% of the subarrays; if the single frame score is >80, trigger the high load mode and enable all redundant subarrays. Dynamically adjust the subarray voltage according to the real-time power consumption (collected by the current sensor MAX44284) to ensure that the total power consumption is ≤5W.
[0110] If the real-time power consumption exceeds the threshold (such as 5.5W), the voltage of non-critical sub-arrays is automatically reduced (such as from 1V to 0.8V). If the power consumption is lower than the target (such as 4W), the voltage is gradually increased in exchange for higher computing accuracy. The relationship between historical load and energy consumption is analyzed through machine learning models (such as linear regression) to predict the optimal resource allocation strategy.
[0111] This data analysis platform significantly improves the speed, energy efficiency and reliability of image processing through the deep integration of memristor storage and computing architecture and dynamic optimization technology. First, the memristor array implements parallel convolution operations in an analog manner, eliminating the data handling bottleneck of the traditional von Neumann architecture. The processing speed is improved compared to the traditional GPU solution, and the system energy consumption is reduced. The dynamic weight synchronization network and parallel sharding convolution accelerator further optimize the computing efficiency, adjust the weight distribution and sharding processing strategy through real-time feature feedback, so that the recognition accuracy in complex industrial scenarios is improved, and the throughput is also broken through to meet the real-time processing requirements of high-resolution images.
[0112] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0113] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image processing data analysis platform based on a memristor storage and computing integrated architecture, characterized in that: include: The XC7Z020 FPGA control module is connected to the OV5640 camera image acquisition module, the memristor array computing module, and the Internet of Things module; The OV5640 camera image acquisition module is used to collect raw image data and transmit it to the XC7Z020 FPGA control module; The memristor array computing module includes a D / A converter, a memristor array, a transimpedance amplifier and an A / D converter, and is used to receive the image data processed by the XC7Z020 FPGA control module, complete the convolution operation through parallel analog calculation, and return the result to the XC7Z020 FPGA control module; The IoT module receives instructions from the XC7Z020 FPGA control module via the CAN bus and executes predefined actions; The platform also includes a dynamic resource allocation mechanism, which integrates a real-time task scheduling unit through the XC7Z020 FPGA control module to dynamically allocate computing tasks to the FPGA logic unit or the memristor array computing module according to the complexity of the input image; The real-time task scheduling unit generates task priorities by analyzing the spatial frequency, edge density and grayscale distribution of the image, and allocates the convolution operation of the high-complexity area to the memristor array computing module for parallel analog computing, while the low-complexity processing tasks are retained in the FPGA logic unit for execution; The weights of the memristor array are pre-written in the form of conductance and loaded to the array nodes through voltage signals to achieve storage-computation integrated convolution operations.
2. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 1, characterized in that: The XC7Z020 FPGA control module is used to segment, compress and grayscale the original image to generate pre-processed data; The pre-processed data is transmitted to the memristor array computing module, and the robotic arm is controlled to perform actions.
3. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 2, characterized in that: The memristor array computing module includes a dynamic weight synchronization network DWS-Net, that is: the XC7Z020FPGA control module is embedded with a lightweight feature extraction unit, which analyzes the spatial frequency and texture distribution of the input image in real time to generate a feature vector; based on the feature vector, the conductance values of different nodes in the memristor array are adjusted through the D / A converter.
4. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 3, characterized in that: The memristor array computing module integrates a parallel sliced convolution accelerator (PSCA), that is: the XC7Z020 FPGA control module divides the input image into N×N independent sub-areas, which are regarded as slices, and each slice is allocated to an independent sub-array of the memristor array; each sub-array completes the local convolution operation through parallel analog calculation, and synchronously outputs the current signal through the multiplexing technology of the transimpedance amplifier; the A / D converter converts the analog signals of the multiple sub-arrays into digital signals, and the XC7Z020 FPGA control module performs global feature fusion.
5. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 4, characterized in that: The XC7Z020 FPGA control module deploys a dual-modal verification mechanism DMVM: constructs a memristor array simulation convolution operation path for the first mode and an FPGA digital logic verification path for the second mode; compares the calculation results of the two modes through a confidence assessment module, and triggers secondary calculation or an alarm if the difference exceeds a preset threshold.
6. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 5, characterized in that: The weight adjustment process of the dynamic weight synchronization network DWS-Net includes: generating a voltage pulse signal through the D / A converter to adjust the initial conductance value of the memristor array; dynamically correcting the conductance distribution according to the real-time image characteristics to match the spatial frequency and grayscale gradient of the input data; the weight adjustment process realizes the weighted summation of the current signal through Kirchhoff's current law to complete the convolution operation.
7. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 6, characterized in that: The slice processing process of the parallel slice convolution accelerator PSCA includes: dividing the high-resolution image into 4×4 or 8×8 slices, each slice corresponds to an independent memristor subarray; the input voltage of each subarray is generated by the D / A converter based on the slice grayscale value, and the current signal is synchronously output through the transimpedance amplifier; the A / D converter synchronously samples the multi-channel subarray signals.
8. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 7, characterized in that: The circuit structure of the memristor array computing module includes: a multi-channel D / A converter and a transimpedance amplifier, which are respectively connected to the bit line BLx and the source line SLx of the memristor array; The D / A converter regulates the conductance value of the memristor through a voltage pulse signal, and the transimpedance amplifier converts a tiny current signal into a voltage signal; The A / D converter converts the analog signal into a digital signal after filtering by a high-speed operational amplifier and transmits it back to the XC7Z020FPGA control module.
9. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 8, characterized in that: The process of the Internet of Things module receiving instructions through the CAN bus includes: the XC7Z020 FPGA control module generates action instructions according to the full connection results of the convolutional neural network; the instructions are transmitted to the robotic arm drive circuit through CAN_H and CAN_L signals to perform sorting, positioning or obstacle avoidance operations.
10. The image processing data analysis platform based on the memristor storage and computing integrated architecture according to claim 9, characterized in that: The dynamic resource allocation mechanism optimizes platform energy consumption through a load balancing algorithm, including: when it is detected that the complexity of multiple consecutive frames of images is lower than a preset threshold, turning off the power supply of some memristor subarrays to reduce power consumption; in high-load scenarios, enabling the redundant subarrays of the parallel sliced convolution accelerator PSCA for hyperthreading acceleration.
Citation Information
Patent Citations
Memristor-based low-power-consumption pulse convolutional neural network hardware architecture
CN112183739A
Binary pulse neural network dynamic image recognition system based on memristor cross array
CN118608921A
Data processing apparatus and method
US20200150971A1
Cited By
Manufacturing production line defect detection method and system based on big data intelligent algorithm
CN120991953A
Full-closed-loop control method and system for joint module of humanoid robot based on FPGA (Field Programmable Gate Array)
CN121200067A
Fixed-point matrix FFT calculation method based on memristor
CN121233885A
A Fixed-Point Matrix FFT Calculation Method Based on Memristors
CN121233885B
Multi-physical-quantity collaborative sensing brain synapse-like device and dynamic environment adaptive method thereof
CN121996076A