Space-time fusion detector double-reservoir network system based on multi-frame-in-one calculation

Through the spatiotemporal fusion detector dual-reservoir network system with multi-frame unified calculation, MoS2 photodetectors and memristor arrays are used to solve the problem of insufficient processing power of machine vision systems in dynamic motion recognition, and achieve efficient and accurate dynamic motion recognition.

CN120654751APending Publication Date: 2025-09-16SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510462480.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing machine vision systems lack processing power in dynamic motion recognition tasks, resulting in a sharp increase in data transmission volume, affecting system efficiency and accuracy.

Method used

A dual-reservoir network system of spatiotemporal fusion detectors based on multi-frame integration calculation is adopted, and MoS2 photodetector arrays and memristor cross arrays are used to realize the encoding and recognition of multi-frame information through nonlinear photoconductive properties and parallel processing.

Benefits of technology

It significantly improves the accuracy and efficiency of dynamic motion recognition, reduces redundant data flow, and reduces power consumption and latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654751A_ABST
    Figure CN120654751A_ABST
Patent Text Reader

Abstract

The invention discloses a time-space fusion detector double-reservoir network system based on multi-frame-in-one calculation, and relates to the technical field of bionic intelligent vision and brain-like calculation. The system comprises a photoelectric reservoir layer and an electrical readout layer, and the photoelectric reservoir layer comprises a retina-shaped MoS2 photoelectric detector array and is used for constructing the photoelectric reservoir layer by using the nonlinear continuous photoconduction characteristic of a two-dimensional MoS2 photoelectric detector. Detecting an optical signal and projecting the optical signal into an electrical readout layer with increased dimensions in a photocurrent form of a multi-frame-in-one pattern; the electrical readout layer includes a memristor cross array for processing an input signal in a parallel manner and generating an output result in real time. The system has the advantages of low power consumption, low delay, high recognition precision and the like, and is particularly suitable for complex application scenes such as infrared dynamic moving target recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bionic intelligent vision and brain-like computing technology, and in particular to a spatiotemporal fusion detector dual-reservoir network system based on multi-frame integration computing. Background Art

[0002] With the rapid advancement of machine vision technology, the amount of data accumulated and collected by detectors has reached an unprecedented scale. At the same time, as application scenarios continue to expand and become more complex, the requirements for the speed and computing power of machine vision systems to handle complex tasks are becoming increasingly stringent. This is especially true for demanding tasks such as dynamic motion recognition, which present unprecedented challenges for detectors.

[0003] In dynamic motion recognition tasks, since objects are constantly in motion, detectors must perform frequent frame-by-frame analysis to ensure accurate capture and identification of motion trajectories and dynamic changes. This frame-by-frame analysis paradigm leads to a dramatic increase in data transmission volume, placing extremely high demands on the detector's internal processing power and data transmission efficiency. To meet this challenge, improving the detector's internal processing power has become key to addressing energy consumption and latency issues in machine vision systems. As a core component of machine vision technology, improving the performance of the detector's internal vision system directly impacts the overall system's operational efficiency and accuracy.

[0004] However, current in-detector vision systems primarily focus on static image processing. While static image processing plays an important role in machine vision, simply capturing images is insufficient for tasks like dynamic motion recognition. Accurate motion recognition requires not only that the detector capture images but also that it record information about how these images change over time, enabling more accurate identification of motion trajectories and dynamic changes. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to provide a spatiotemporal fusion detector dual-reservoir network system capable of improving the accuracy of dynamic motion recognition.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a dual-reservoir network system of spatiotemporal fusion detectors based on multi-frame integration calculation, characterized in that it includes: a photoelectric reservoir layer and an electrical readout layer, the photoelectric reservoir layer includes a retinal MoS2 photodetector array, which is used to utilize the nonlinear continuous photoconductivity characteristics of the two-dimensional MoS2 photodetector to construct a photoelectron reservoir layer, detect the light signal and project it in the form of a photocurrent of a multi-frame integration pattern into the electrical readout layer with increased dimension; the electrical readout layer includes a memristor cross array, which is used to process the input signal in parallel and generate output results in real time.

[0007] A further technical solution is that the retinal-shaped MoS2 photodetector array includes several retinal-shaped MoS2 photodetection units, the retinal-shaped MoS2 photodetection units include two independent van der Waals heterojunction structures, the van der Waals heterojunction structure includes a sapphire substrate, the upper surface of the sapphire substrate is formed with a MoS2 detector, the upper surface of the MoS2 detector is formed with several cross-finger electrodes, and the above two independent MoS2 detectors are subjected to different O2 plasma treatments.

[0008] A further technical solution is that the memristor crossbar array includes a plurality of programmable memristor units, the conductivity of each memristor unit is adjusted by an external voltage, and the pre-trained weight is mapped to the conductivity of the memristor unit.

[0009] The beneficial effects of the above-described technical solution are as follows: The system described herein utilizes MoS2 photodetectors for in-sensor computing, integrating multiple frames into a single frame with persistent photoconductivity for dynamic motion recognition. The inherent photoconductivity of MoS2 photodetectors allows for the embedding of spatiotemporal information into a single frame, effectively reducing redundant data streams and simplifying dynamic visual tasks. The present invention utilizes a simple process to fabricate photodetectors with varying light responses and attenuation characteristics within the same pixel region, achieving the first implementation of a photoelectric multi-reservoir hardware neural network, significantly improving the recognition rate of similar networks.

[0010] The dual-terminal MoS2 photodetector design is not only highly scalable but also highly tolerant to heterogeneity in 2D material devices. This allows for the addition of more devices per pixel, enriching the reservoir state and further enhancing network training.

[0011] Through the above design and working mechanism, low-power and low-latency dynamic motion target recognition is achieved, and the recognition accuracy and efficiency of the system are significantly improved through multi-frame integrated spatiotemporal information encoding and multi-reservoir design. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0013] Figure 1 is a schematic structural diagram of the system according to an embodiment of the present invention;

[0014] Figure 2 2 is a schematic structural diagram of a MoS2 photodetector unit in a system according to an embodiment of the present invention;

[0015] Figure 31 is a schematic diagram of a top view of a MoS2 photodetector unit in a system according to an embodiment of the present invention;

[0016] Figure 4 The photoresponse of the Au / MoS2 / Au device with different O2 plasma exposure treatment times measured under illumination (λ=520nm, 1.97mW laser power) at a bias voltage of 0.2V provided by the present invention;

[0017] Figure 5 This is a diagram showing the experimental results of the persistent photocurrent effect for encoding time information obtained in Example 1;

[0018] Figure 6 Schematic diagram of the spatiotemporal multi-mask RC system based on MIP hardware implementation for dynamic motion recognition within the detector obtained in Example 2, and compared with traditional fully connected and single-mask reservoir networks;

[0019] Among them: 1. Retina-shaped MoS2 photodetector array; 2. Memristor cross array; 3. Retina-shaped MoS2 photodetection unit; 4. Sapphire substrate; 5. MoS2 detector; 6. Cross-finger electrode. DETAILED DESCRIPTION

[0020] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.

[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0022] like Figure 1 As shown, an embodiment of the present invention discloses a dual-reservoir network system of a spatiotemporal fusion detector based on multi-frame integration calculation, comprising: a photoelectric reservoir layer and an electrical readout layer, wherein the photoelectric reservoir layer comprises a retinal MoS2 photodetector array 1, which is used to construct a photoelectron reservoir layer by utilizing the nonlinear continuous photoconductivity characteristics of a two-dimensional MoS2 photodetector, detect light signals and project them in the form of photocurrents of a multi-frame integration pattern into an electrical readout layer with increased dimensionality; the electrical readout layer comprises a memristor cross array 2, which is used to process input signals in a parallel manner and generate output results in real time.

[0023] The retinal MoS2 photodetector array 1 includes several retinal MoS2 photodetection units 3. Preferably, the detector array size is 8×5 or 16×10; the retinal MoS2 photodetection unit 3 includes two independent van der Waals heterojunction structures, and the van der Waals heterojunction structure includes a sapphire substrate 4. The upper surface of the sapphire substrate 4 is formed with a MoS2 detector 5, and the upper surface of the MoS2 detector 5 is formed with several cross-finger electrodes 6. The above two independent MoS2 detectors 5 are subjected to different O2 plasma treatments.

[0024] like Figure 2 、 Figure 3 As shown, in the present invention, each pixel of the detector array (retinal-shaped MoS2 photoelectric detection unit 3) has two independent devices, and the independent device is a van der Waals heterojunction structure, which includes a sapphire substrate 1, a MoS2 detector 2 and a plurality of cross-finger electrodes 3 from bottom to top; the thickness of the sapphire substrate 1 is preferably 80μm to 120μm, more preferably 100μm; Figure 2 As shown, the left side of the MoS2 detector is a pure MoS2 detection device, and the right side is an O-doped device, and the two devices are independent of each other; the two independent devices have different photoelectric properties, especially the photocurrent decay time, which is used to define the nonlinear attenuation equation in subsequent network operations.

[0025] The device acts as multiple reservoirs in the reservoir layer with different input responses and decay characteristics, and obtains two sets of final states after nonlinear decay to remember all information of the past and current frames;

[0026] When identifying dynamic moving targets, the nonlinear persistent photoconductivity effect (PPC) of the MoS2 photodetector is used to encode the optical information of multiple frames into the photocurrent output of a single frame. The device acts as multiple reservoirs with different input responses and attenuation characteristics in the reservoir layer, and two sets of final states are obtained after nonlinear attenuation to remember all information of the past and current frames. The nonlinear encoding method not only retains the temporal information of the optical signal, but also further enriches the information dimension through the multi-reservoir design (two detectors in each pixel), providing richer feature extraction for the subsequent readout layer.

[0027] The output signal can be input to a readout layer via a transimpedance amplifier (TIA). The present invention provides a memristor crossbar array 2 as a readout layer. The memristor crossbar array can include multiple programmable memristor cells, and the conductivity of each memristor cell is adjusted by an external voltage. By mapping pre-trained weights to the conductivity of the memristor cells, the memristor crossbar array can process input signals in parallel and generate output results in real time. The memristor crossbar array linearly combines the output currents of multiple reservoirs according to Kirchhoff's law to generate the final recognition result.

[0028] The present invention also provides a method for preparing the MoS2 photodetector array described in the above technical solution, such as Figure 2-Figure 3 As shown, the following steps are included:

[0029] MoS2 / sapphire substrate grown by chemical vapor deposition;

[0030] The substrate is preferably prepared by chemical vapor deposition or physical vapor deposition. In the present invention, chemical vapor deposition is preferably the selected deposition method, and the deposition rate of the thermal evaporation deposition is preferably less than 0.5 angstroms / second for cadmium and less than 1 angstrom / second for gold, more preferably 0.1 to 0.5 angstroms / second. The present invention does not specifically limit the ion beam sputtering deposition, electron beam evaporation deposition, or thermal evaporation deposition process, and can be performed using processes familiar to those skilled in the art.

[0031] Using standard microfabrication processes, an Au / MoS2 / Au array with interdigitated electrodes is prepared on a MoS2 / sapphire substrate; the MoS2 film comprises four atomic layers; the metal electrodes are made of Cr / Au; the thickness of the Cr metal electrode is preferably 3 to 8 nm, and the thickness of the Au metal electrode is preferably 45 to 60 nm;

[0032] In the present invention, the metal electrode is preferably prepared by ion beam sputtering deposition, electron beam evaporation deposition, or thermal evaporation deposition. The selected evaporation deposition method is preferably thermal evaporation deposition, and the deposition rate of the thermal evaporation deposition is preferably less than 0.5 angstroms / second for cadmium and less than 1 angstrom / second for gold, more preferably 0.1 to 0.5 angstroms / second. The present invention does not specifically limit the ion beam sputtering deposition, electron beam evaporation deposition, or thermal evaporation deposition process; processes familiar to those skilled in the art can be used.

[0033] like Figure 2 、 Figure 3As shown, in the present invention, each detector unit includes two MoS2 detectors; the detector unit uses a 4×4 unit array to construct dual-reservoir hardware, with a total of 32 MoS2 detectors; in subsequent steps, the detectors will be subjected to different O2 plasma treatments to change their electrical and optoelectronic properties; a laser direct writing technique (MicroWriter ML3) is used to prepare a pattern on the prepared array to protect the right device 4 in each RC optoelectronic unit, and the left device 2 is sequentially subjected to O2 plasma treatment with specific parameters; after removing the photoresist, the right device 4 is subjected to O2 plasma treatment with other parameters using the same process;

[0034] Specifically, an annealing process is performed on the RC photoelectron array to obtain the MoS2 photodetector array; the annealing process temperature is preferably 100°C to 150°C, more preferably 120°C; the annealing time is preferably 3 to 8 minutes, more preferably 5 minutes.

[0035] The present invention also provides the O2 plasma treatment technology described in the above technical solution, such as Figure 4 As shown:

[0036] The exposure time of the device is 1 second. The O2 plasma exposure is as follows Figure 4 Indicated by long dash;

[0037] The device benefits from a short treatment time, which does not destroy the MoS2 lattice structure, and due to the carrier capture effect at the MoS2 / MoOx heterojunction interface, the photocurrent of the device changes by up to 10 times, while the decay time is slightly reduced, such as Figure 4 As shown in;

[0038] After the device has been processed for 2 seconds, the decay time slows down, such as Figure 4 As shown by the dashed lines in FIG, parallel MoS2 lattice fringes with uniform interlayer spacing extend only in the top layer; the uniform interlayer spacing of the device is preferably 0.5nm to 0.8nm, more preferably 0.6nm; the parallel MoS2 lattice fringes extend to a depth of 0.6nm to 1.0nm in the top layer; the device recovers when the exposure time is 3 seconds, as shown in FIG. Figure 4 As shown by the short dotted line in ; for this device, the current level continues to decrease with increasing plasma exposure time until it reaches the instrument noise level after 5 seconds of plasma treatment.

[0039] The present invention also provides a method for encoding optical information of multiple frames into a single frame of photocurrent output as described in the above technical solution:

[0040] enabling the one detector to receive light pulses of different sequences representing different information to generate different photocurrents;

[0041] The final state of the photocurrent is recorded by a semiconductor device analyzer;

[0042] irradiating the detector with three consecutive frames of light pulses at a fixed frequency to generate a photocurrent that decays over time;

[0043] The fixed frequency is preferably 8 Hz to 12 Hz, more preferably 10 Hz;

[0044] The greater the number of light pulses received by the detector, the higher the current level;

[0045] When the number of pulses remains unchanged, the different pulse sequences also determine the final current.

[0046] The present invention also provides for the use of the retinal-shaped MoS2 photodetector described in the above technical solution, or the retinal-shaped MoS2 photodetector prepared by the preparation method described in the above technical solution, in other biomimetic optoelectronic devices. The present invention does not particularly limit the specific implementation of such application, and the application can be carried out using processes familiar to those skilled in the art.

[0047] In order to further illustrate the present invention, the following detailed description of a spatiotemporal fusion detector dual-reservoir network (RC) system based on a multi-frame integration calculation mode provided by the present invention is given in combination with the accompanying drawings and embodiments, but they should not be understood as limiting the scope of protection of the present invention.

[0048] The structure of the system is described below with reference to specific embodiments and tests.

[0049] Example 1:

[0050] A detector receives different sequences of light pulses representing different information, generating different photocurrents. The final state of the photocurrent is recorded using a semiconductor device analyzer. Three consecutive frames of light pulses are irradiated onto the detector at a fixed frequency (10 Hz in this experiment), generating a photocurrent that decays over time. The present invention records the final state of the photocurrent after three frames and finds differences in the residual current. Generally, the greater the number of light pulses received, the higher the current level. When the number of pulses remains unchanged, the different pulse sequences also determine the final current. The results clearly demonstrate that only the last frame of data is needed to accurately identify and classify time-dependent input signals.

[0051] Example 2:

[0052] Three frames are played continuously, representing the movement of an object (in this experiment, a simulated car) at different speeds in eight possible directions. The MoS2 photodetector array receives time-sequential light pulses corresponding to the object's position at a fixed frequency. Each set of photodetectors representing a reservoir group has a unique response and attenuation to the same light sequence, thereby generating a dual-mode output current through the two reservoirs. Typically, due to the lack of temporal information, it is impossible to complete the task of identifying the direction of dynamic motion using only one frame. However, due to the persistent photoconductivity (PPC) effect of the device, the system of the present invention enables the final state to retain the spatiotemporal memory of multiple previous frames. The photocurrent representing the final state of all historical information is then input into the subsequent memristor crossbar array through the TIA.

[0053] Each pixel generates a light pulse based on the position of the real object and receives the light signal that is irradiated onto the photodetector array. After receiving and processing the light signal, the photodetector array converts this information into a photocurrent.

[0054] After generating this final state, the photodetector array outputs it to the memristor array. When the memristor crossbar array receives the input signal, it combines the states of each reservoir in parallel according to Kirchhoff's laws. The conductivity of the memristors can be finely tuned to ensure that each weight is accurately mapped to its corresponding memristor cell.

[0055] Test Example 1:

[0056] like Figure 5 As shown, a 3-bit optical pulse input from "000" to "111" has a pulse width and interval of 100ms and 400ms, respectively. The corresponding photoresponse characteristics are extracted from two Au / MoS2 / Au devices with different doping levels, one after 0s and the other after 2s of O2 plasma treatment.

[0057] Test Example 2:

[0058] like Figure 6 As shown, a car is assigned to a 4×4 mapping area. This car can move in any direction at a random speed. Furthermore, the car's route may not be strictly straight and may include small turns along the way, which is closer to reality. Each reservoir generates an information frame consisting of 16 node states.

[0059] The memristor array receives 32 outputs from the photodetector array. The final states of the two reservoirs exhibit distinct patterns, representing two dimensions extracted from the low-dimensional input space. These patterns are not only unique but also each represent two distinct dimensions extracted from the low-dimensional input space.

[0060] 6,000 training data were used to train the RC system within the detector, and 2,000 test data were used for verification. After 150 training cycles, 99.5% recognition accuracy was achieved through simulation, which is 4% higher than the traditional FC network and 3% higher than the single mask reservoir network.

[0061] Compared with the traditional fully connected network that calculates frame by frame, the multi-frame integrated multi-reservoir network system of the present invention improves the accuracy while reducing the readout layer size by 33%.

Claims

1. A dual-reservoir network system of spatiotemporal fusion detectors based on multi-frame unification calculation, characterized by include: A photoelectric reservoir layer and an electrical readout layer, wherein the photoelectric reservoir layer comprises a retina-shaped MoS2 photodetector array (1), which is used to construct a photoelectron reservoir layer by utilizing the nonlinear continuous photoconductivity characteristics of the two-dimensional MoS2 photodetector, detect light signals and project them into the electrical readout layer with increased dimension in the form of photocurrent of a multi-frame unified pattern; the electrical readout layer comprises a memristor cross array (2), which is used to process input signals in a parallel manner and generate output results in real time.

2. The dual-reservoir network system for spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 1 is characterized in that: The retinal MoS2 photodetector array (1) comprises a plurality of retinal MoS2 photodetection units (3), wherein the retinal MoS2 photodetection units (3) comprise two independent van der Waals heterojunction structures, wherein the van der Waals heterojunction structures comprise a sapphire substrate (4), wherein a MoS2 detector (5) is formed on the upper surface of the sapphire substrate (4), and wherein a plurality of interdigitated electrodes (6) are formed on the upper surface of the MoS2 detector (5), and wherein the two independent MoS2 detectors (5) are subjected to different O2 plasma treatments.

3. The dual-reservoir network system of spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 2 is characterized in that: The method for preparing the retinal MoS2 photodetector array comprises the following steps: growing a sapphire substrate (4) by chemical vapor deposition; Growing a MoS2 layer on the upper surface of a sapphire substrate (4) by chemical vapor deposition; fabricating an interdigitated electrode array on the MoS2 layer using a microfabrication process; Using laser direct writing technology to prepare patterns on the prepared array; The MoS2 detector on the left is treated with O2 plasma using the set parameters in sequence; After removing the photoresist from the MoS2 detector on the left, the MoS2 detector on the right was treated with O2 plasma with other set parameters using the same process; The device array is subjected to an annealing process to obtain a retina-shaped MoS2 photodetector array.

4. The dual-reservoir network system of spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 3 is characterized in that: The cross-finger electrodes are prepared by ion beam sputtering deposition, electron beam evaporation deposition or thermal evaporation deposition process.

5. The dual-reservoir network system of spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 3 is characterized in that: The thickness of the sapphire substrate is 80 μm to 120 μm.

6. The dual-reservoir network system for spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 3, characterized in that: When identifying dynamic moving targets, the nonlinear persistent photoconductivity (PPC) effect of the MoS2 photodetector is used to encode the optical information of multiple frames into the photocurrent output of a single frame. The MoS2 photodetector acts as multiple reservoirs with different input responses and attenuation characteristics in the reservoir layer, and two sets of final states are obtained after nonlinear attenuation to remember all information of the past and current frames.

7. The dual-reservoir network system for spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 3 is characterized in that: The annealing process temperature is 100° C. to 150° C., and the annealing time is preferably 3 to 8 minutes.

8. The dual-reservoir network system for spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 1 is characterized in that: The retina-shaped MoS2 photodetector array includes 8×5 or 16×10 retina-shaped MoS2 photodetection units.

9. The dual-reservoir network system for spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 1, characterized in that: The memristor crossbar array includes a plurality of programmable memristor units, the conductivity of each memristor unit is adjusted by an external voltage, and the pre-trained weight is mapped to the conductivity of the memristor unit.

10. The dual-reservoir network system of spatiotemporal fusion detectors based on multi-frame unification calculation according to claim 1 is characterized in that: The method for encoding optical information of multiple frames into a single frame of photocurrent output comprises the following steps: Setting the one two-dimensional MoS2 photodetector to receive light pulses of different sequences representing different information to generate different photocurrents; The final state of the photocurrent is recorded by a semiconductor device analyzer; irradiating the detector with three consecutive frames of light pulses at a fixed frequency to generate a photocurrent that decays over time; The greater the number of light pulses received by the detector, the higher the current level; when the number of pulses remains unchanged, different pulse sequences determine the final current.