Multimodal neuromorphic vision perception and processing system and method thereof
Patent Information
- Application Number
- CN202610531224.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-18
AI Technical Summary
这种专化设计导致系统在动态多变的真实环境中适应性不足,从而导致对场景中视觉信息的处理效率较低
[0015] This application proposes a multi-mode neuromorphic visual perception and processing system and method. When processing visual information, each multi-mode pixel unit in the system's multi-mode pixel unit array can adjust its corresponding unit architecture according to the input voltage of the voltage source. Different bias voltages can be provided to the multi-mode intelligent vision device based on the unit negative feedback loop composed of amplifiers and transistors, thus setting the multi-mode pixel units to different unit modes. Furthermore, the unit modes can be flexibly adjusted based on preset visual processing configurations and processing function requirements, thereby enabling the adjusted multi-mode pixel unit array to better adapt to multi-scene intelligent visual processing, improving the system's adaptability in dynamic and changing real-world environments, and increasing the processing efficiency of visual information in the scene.
Smart Images

Figure CN122596142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of neuromorphic semiconductor technology and bionic vision technology, and in particular to a multimodal neuromorphic vision perception and processing system and method thereof. Background Technology
[0002] Neuromorphic visual processing systems (also known as intelligent visual systems) refer to specialized hardware chips / systems that mimic the structure and working principles of biological nervous systems (such as visual receptors, afferent nerves, and central nervous systems) and are used to process visual information (such as light signals, images, and videos).
[0003] The intelligent vision systems employed in related technologies are often limited by specific processing paradigms, such as using event-driven processing methods only in moving scenes or frame-based processing methods only in static scenes. This specialized design results in insufficient adaptability of the system to dynamic and changing real-world environments, leading to low efficiency in processing visual information within the scene. Summary of the Invention
[0004] The main objective of this application is to propose a multimodal neuromorphic visual perception and processing system and method, which can flexibly adjust the neuromorphic visual processing system, improve the system's adaptability in dynamic and ever-changing real environments, and improve the processing efficiency of visual information in the scene.
[0005] To achieve the above objectives, a first aspect of the present application provides a multi-mode neuromorphic visual perception and processing system. The system includes a multi-mode pixel unit array composed of a row and column structure. Each multi-mode pixel unit in the multi-mode pixel unit array includes a multi-mode intelligent vision device and a multi-mode control circuit composed of a voltage source, an amplifier, and a transistor. The voltage source is used to provide input voltage to the multi-mode pixel unit in order to adjust the unit architecture of the multi-mode pixel unit; The amplifier and transistor in the multi-mode pixel unit constitute a unit negative feedback loop, and the unit negative feedback loop is used to provide different bias voltages for the multi-mode intelligent vision device, so that the multi-mode pixel unit is set to different unit modes. Based on different visual processing requirements and corresponding visual processing algorithms, multi-mode pixel units and their array configurations are reconstructed. The target unit mode corresponding to each multi-mode pixel unit is determined according to a preset visual processing configuration. Different visual processing configurations are used to perform different array functional partitioning, unit mode allocation, and data flow path configurations on the multi-mode pixel unit array, so as to achieve the corresponding visual information processing functions through the determined array functional partitioning, unit mode allocation, and data flow path configuration.
[0006] In some embodiments, the voltage source is used to provide an input voltage to the multi-mode pixel unit to adjust the cell architecture of the multi-mode pixel unit, including: The voltage source provides an enable signal to the amplifier in the multi-mode pixel unit, and the unit architecture of the multi-mode pixel unit is adjusted to the first circuit architecture. The first circuit architecture enables the multi-mode pixel unit array to form a 1R array architecture. Then, the multi-mode intelligent vision device generates a current signal, and the amplifier and transistor maintain the negative voltage of the device through negative feedback and convert the current into voltage output. By providing a disable signal to the amplifier in the multi-mode pixel unit through the voltage source, the unit architecture of the multi-mode pixel unit is adjusted to a second circuit architecture, which enables the multi-mode pixel unit array to form a 1T1R cross array architecture.
[0007] In some embodiments, the unit modes include photoresponsive neuron mode, photoactivated synapse mode, long-term memory synapse mode, and leakage current accumulation discharge neuron mode; The unit negative feedback loop is used to provide different bias voltages for the multi-mode intelligent vision device, so that the multi-mode pixel unit is set to different unit modes, including: The unit negative feedback loop provides a first bias voltage value to the multi-mode intelligent vision device, thereby setting the multi-mode pixel unit to the photoresponse neuron mode. The unit negative feedback loop provides a voltage in the first bias voltage range to the multi-mode intelligent vision device, thereby setting the multi-mode pixel unit to photoactivated synaptic mode. The unit negative feedback loop provides a second bias voltage range to the multi-mode intelligent vision device, thereby setting the multi-mode pixel unit to a long-term memory synaptic mode. The unit negative feedback loop provides the voltage of the third bias voltage range to the multi-mode intelligent vision device, so that the multi-mode pixel unit is set to the leakage current accumulation discharge neuron mode, wherein the first bias voltage range, the second bias voltage range and the third bias voltage range do not intersect each other.
[0008] In some embodiments, the visual processing configuration method includes a single-neural morphological perception configuration method, and the single-neural morphological perception configuration method includes an event-driven visual processing configuration algorithm and a frame-based visual processing algorithm. The step of determining the target unit mode corresponding to each multi-mode pixel unit according to a preset visual processing configuration includes: When the visual processing configuration method uses an event-driven visual processing configuration algorithm for visual perception, the target mode corresponding to all the multi-mode pixel units in the multi-mode pixel unit array is set to the light response neuron mode. When the preset visual processing configuration uses a frame-based visual processing algorithm for visual perception, the target mode corresponding to all the multi-mode pixel units in the multi-mode pixel unit array is set to the photoactivated synaptic mode.
[0009] In some embodiments, the visual processing configuration further includes a hybrid configuration based on neuromorphic perception and neuromorphic computing, and the hybrid configuration based on neuromorphic perception and neuromorphic computing includes a visual spiking artificial neural network algorithm, a visual spiking recurrent neural network algorithm, a visual artificial neural network algorithm, and a visual reservoir processing algorithm. The step of determining the target unit mode corresponding to each multi-mode pixel unit according to a preset visual processing configuration includes: When the visual processing configuration uses the visual impulse artificial neural network algorithm for visual processing, the multi-mode pixel unit array is divided into a photoresponse neuron array, a long-term memory synapse array, and a leakage current accumulation discharge neuron array, and the multi-mode pixel units contained in different arrays are adjusted in unit mode based on the corresponding array function. When the visual processing configuration uses the visual impulse recurrent neural network algorithm for visual processing, the multi-mode pixel unit array is divided into a photoresponse neuron array, a long-term memory synapse array, and a leakage current accumulation discharge neuron array, and the multi-mode pixel units contained in different arrays are adjusted in unit mode based on the corresponding array function. When the visual processing configuration uses a visual artificial neural network algorithm or a visual reservoir processing algorithm for visual processing, the multi-mode pixel unit array is divided into a photoactivated synaptic array and a long-term memory synaptic array, and the multi-mode pixel units contained in different arrays are adjusted in unit mode based on the corresponding array function.
[0010] To achieve the above objectives, a second aspect of this application proposes a neuromorphic visual image processing method, the method comprising: Acquire the visual images to be processed and the image recognition requirements data; Determine the target visual processing configuration that matches the visual image to be processed based on the image recognition requirement data; Based on the target visual processing configuration, the multi-mode pixel unit array in the multi-mode neuromorphic visual perception and processing system described in the first aspect is systematically adjusted to obtain the target neuromorphic visual processing system. The target neuromorphic visual processing system performs image processing on the visual image to be processed.
[0011] In some embodiments, the image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the target visual processing configuration method adopts an event-driven visual processing configuration algorithm for visual processing, the visual image to be processed is input into the target neuromorphic visual processing system, and the visual image to be processed is regarded as a light pulse signal image. The optical pulse signal image is filtered based on the device's illumination intensity threshold and cumulative time threshold to obtain the first target voltage pulse signal image.
[0012] In some embodiments, the image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the preset visual processing configuration uses a frame-based visual processing algorithm for visual processing, the visual image to be processed is input into a photoactivated synapse array of a multi-mode pixel unit array for image processing, resulting in a processed voltage non-pulse signal image.
[0013] In some embodiments, the image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the target visual processing configuration adopts the visual spiking artificial neural network algorithm for visual processing, the multi-mode pixel unit array in the target neuromorphic visual processing system includes a photoresponse neuron array, a long-term memory synapse array, and a leakage current accumulation discharge neuron array. The visual image to be processed is input into the optical response neuron array for image optical signal conversion to obtain a second voltage pulse signal image; The second voltage pulse signal diagram is input into the long-term memory synaptic array for weighted fusion to obtain the third voltage pulse signal diagram; The third voltage pulse signal is input into the leakage current accumulation discharge neuron array for unit accumulation to obtain the target voltage signal.
[0014] In some embodiments, the image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the target visual processing configuration uses a visual artificial neural network algorithm or a visual reservoir processing algorithm for visual processing, the multi-mode pixel unit array in the target neuromorphic visual processing system includes a photoactivated synapse array and a long-term memory synapse array. The visual image to be processed is input into the photoactivated synapse array for image enhancement or compression to obtain an enhanced or compressed voltage non-pulse signal image. The enhanced voltage non-pulse signal map is input into the long-term memory synaptic array for weighted fusion to obtain the third target voltage non-pulse signal map.
[0015] This application proposes a multi-mode neuromorphic visual perception and processing system and method. When processing visual information, each multi-mode pixel unit in the system's multi-mode pixel unit array can adjust its corresponding unit architecture according to the input voltage of the voltage source. Different bias voltages can be provided to the multi-mode intelligent vision device based on the unit negative feedback loop composed of amplifiers and transistors, thus setting the multi-mode pixel units to different unit modes. Furthermore, the unit modes can be flexibly adjusted based on preset visual processing configurations and processing function requirements, thereby enabling the adjusted multi-mode pixel unit array to better adapt to multi-scene intelligent visual processing, improving the system's adaptability in dynamic and changing real-world environments, and increasing the processing efficiency of visual information in the scene. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the structure of the multimodal neuromorphic visual perception and processing system provided in the embodiments of this application; Figure 2 This is a schematic diagram of a unit negative feedback loop provided in an embodiment of this application; Figure 3A This application provides schematic diagrams of the working modes of the multi-mode control circuit under various unit modes and schematic diagrams of the output modes of the multi-mode intelligent vision device under different unit modes. Figure 3B This is provided by the embodiments of this application. Figure 3A Input / output diagrams for different unit modes; Figure 4 This is a schematic diagram of a hardware architecture for visual information processing based on a hybrid NP / NC configuration provided in an embodiment of this application; Figure 5 This is a schematic diagram of the data stream processing flow under different configuration methods provided in the embodiments of this application; Figure 6 This is a schematic diagram of a neuromorphic visual image processing method provided in the application embodiment; Figure 7 yes Figure 6 The first flowchart of step S640; Figure 8 This is a schematic diagram of an event-driven visual processing configuration algorithm for visual processing provided in an embodiment of this application. Figure 9This is a schematic diagram of a frame-based visual processing configuration algorithm for visual processing provided in an embodiment of this application. Figure 10 yes Figure 6 The second flowchart of step S640; Figure 11 This is a schematic diagram of an algorithm for visual processing based on a visual spiking artificial neural network algorithm provided in an embodiment of this application; Figure 12 This is a schematic diagram of an algorithm for visual processing based on a visual impulse recurrent neural network algorithm provided in an embodiment of this application; Figure 13 This is a schematic diagram of an algorithm for visual processing based on a visual artificial neural network algorithm provided in an embodiment of this application; Figure 14 This is a schematic diagram of an algorithm for visual processing based on a visual reservoir processing algorithm provided in an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] Neuromorphic vision processing systems (also known as intelligent vision systems) refer to dedicated hardware chips / systems that mimic the structure and working principles of biological nervous systems (such as visual receptors, afferent nerves, and the central nervous system) to process visual information (such as light signals, images, and videos). An ideal intelligent vision system must simultaneously meet two key design goals: high flexibility to adapt to different scenarios and high power / area efficiency for specific applications, in order to achieve the next evolutionary breakthrough in intelligent vision technology. However, current technological solutions struggle to achieve these two conflicting goals, hindering the improvement of overall system performance. This contradiction is particularly pronounced when deployed in edge computing environments with limited hardware resources and strict power consumption constraints. Therefore, even under relatively unrestricted resource conditions, the system still needs to address the diverse, dynamic, and unpredictable challenges of visual events.
[0021] Furthermore, the intelligent vision systems employed in related technologies are often limited by specific processing paradigms, such as using event-driven visual processing methods only in moving scenes or frame-based visual processing methods only in static scenes. This specialized design leads to insufficient adaptability of the system in dynamic and ever-changing real-world environments, resulting in low efficiency in processing visual information within the scene. In recent years, cutting-edge reconfigurable CMOS intelligent vision systems have made some progress, enabling reconfigurability of system configurations in front-end neuromorphic perception (NP) and back-end neuromorphic computing (NC) vision processors, and supporting flexible algorithm deployment at the chip and module levels to address various application scenarios. Nevertheless, their overall performance still falls short of requirements, and the root cause of this problem lies in the deficiencies at the level of basic devices and architecture. Specifically, at the device unit level, the systems employed in related technologies mainly rely on heterogeneous integration and single-function devices, lacking cross-paradigm dynamic reconfigurability. At the system architecture level, the fixed "sensor-computing" structure also limits its flexibility and operational efficiency in various concrete scenarios. For example, scenarios oriented towards neuromorphic perception require high-quality intelligent visual acquisition, while scenarios oriented towards AI computing have higher requirements for high-precision neuromorphic computing processing. The current architecture is difficult to optimize and adapt to these different needs at the same time.
[0022] Therefore, there is an urgent need to propose a novel multimodal neuromorphic visual perception and processing system that can flexibly allocate device resources according to the dynamic needs of different scenarios, improve the system's adaptability in dynamic and ever-changing real environments, and improve the efficiency of processing visual information in the scenarios.
[0023] Based on this, embodiments of this application provide a multimodal neuromorphic visual perception and processing system and method, which can flexibly adjust the neuromorphic visual processing system, improve the system's adaptability in dynamic and changing real environments, and improve the processing efficiency of visual information in the scene.
[0024] Please see Figure 1 , Figure 1 This is a schematic diagram of a multimodal neuromorphic visual perception and processing system provided in an embodiment of this application. Figure 1 The multi-mode neuromorphic visual perception and processing system includes a multi-mode pixel unit array 110 composed of a row and column structure, and each multi-mode pixel unit in the multi-mode pixel unit array 110 includes a multi-mode intelligent vision device and a multi-mode control circuit composed of a voltage source, an amplifier, and a transistor.
[0025] The multi-mode pixel unit designed in this application embodiment is a reconfigurable multi-mode control 1A1T1R circuit architecture. This 1A1T1R architecture unit can add a low-power amplifier to the 1T1R architecture. The output terminal of the amplifier is connected to the gate of the NMOS transistor, and the non-inverting input terminal is connected to the drain of the NMOS transistor. A system composed of multiple multi-mode pixel units in a row-column structure is a reconfigurable system. That is, the visual unit (multi-mode pixel unit) array corresponding to the system can be divided into units in different ways to form a highly flexible architecture reconfiguration to achieve different visual information processing functions.
[0026] Each multi-mode pixel unit contains one device in a multi-mode pixel unit array (which can refer to a chip). Device resources indicate the number of devices on the chip. For example, if there are 100 devices on the chip, there are 100 multi-mode pixel units.
[0027] It should be noted that, in combination Figure 1 In this embodiment of the application, the 1A1T1R architecture units can be configured into a multi-mode pixel unit array using a homogeneous array, and a multiply-accumulate module that can define the working state (meaning working or not working) can be designed for each column of devices for calculation in NC mode.
[0028] Combination Figure 1 The voltage source is used to provide input voltage to the multi-mode pixel unit to adjust the unit architecture of the multi-mode pixel unit. The adjustment process of the unit architecture may specifically include: An enable signal is provided to the multi-mode pixel unit through a voltage source, adjusting the unit architecture of the multi-mode pixel unit to a first circuit architecture, which is similar to a 1R architecture. This first circuit architecture enables the multi-mode pixel unit array to form a 1R array architecture. A current signal is generated by the multi-mode intelligent vision device, and the amplifier and transistor maintain the negative terminal voltage of the device through negative feedback, converting the current into a voltage output. By providing a disable signal to the multi-mode pixel unit through a voltage source, the unit architecture of the multi-mode pixel unit is adjusted to a second circuit architecture. In the second circuit architecture unit, the amplifier does not work, and the device and transistor form a 1T1R unit. The second circuit architecture enables the multi-mode pixel unit array to form a 1T1R cross array architecture.
[0029] In this embodiment, the unit architecture of each multi-mode pixel unit can be switched by an enable signal. That is, by giving the amplifier an enable or disable signal, the unit module of the multi-mode pixel unit can be freely switched between the first circuit architecture (i.e., 1R architecture) and the second circuit architecture (i.e., 1T1R architecture).
[0030] Combination Figure 1 The amplifiers and transistors in the multi-mode pixel unit can form a unit negative feedback loop, which is used to provide different bias voltages for the multi-mode intelligent vision device, enabling the multi-mode pixel unit to be set to different unit modes. In other words, the embodiments of this application can form a clamping circuit through the unit negative feedback loop to provide different bias voltages for the multi-mode intelligent vision device, thereby realizing the switching of different unit modes.
[0031] For example, please refer to Figure 2 , Figure 2 This is a schematic diagram of a unit negative feedback loop provided in an embodiment of this application. As shown by the red arrow in the diagram, the unit negative feedback loop has a programmable bias voltage applied to the negative terminal of the amplifier. When the amplifier is enabled, the amplifier and the transistor together form the unit negative feedback loop. When the multi-mode intelligent vision device in this unit generates current, the voltage at the positive terminal of the amplifier (see the red dot in the diagram) increases, and then the amplifier outputs a voltage V. P At the same time, this voltage V P This increases the gate voltage of the NMOS and decreases its on-resistance, thus lowering the voltage at the positive terminal of the op-amp. This feedback circuit ensures that the voltage at the lower end of the device is always equal to the reference voltage provided by the negative terminal of the amplifier. At this time, the voltage applied to the negative terminal of the amplifier is precisely applied to the "multi-mode intelligent vision device" through the negative feedback clamping circuit, serving as its operating bias. Among them, SL... iThis is the source line, which can be used as a signal input line. 'i' represents the row index of the multi-mode pixel unit in the multi-mode pixel unit array. In the 1T1R architecture (without amplifier operation), the negative terminal of the device is directly connected to the drain of the transistor. During signal transmission, this line can serve as a data input line, inputting the processed voltage signal from the previous stage into this unit. BL k1 This is a bit line, which can be used as a signal output line or a power supply line. As a data output line, it sends the signal processed by this unit to the Multiply-and-accumulate driver (MAC driver) module located at the end of each column to convert the current signal into a voltage signal. k1 represents the column index of the multi-mode pixel unit in the multi-mode pixel unit array. In the 1T1R architecture, it is connected to the source of the transistor. SL i and BL k1 Together, in conjunction with the gate signal WL i This allows addressing and reading / writing of a specific cell located in the i-th row and k1-th column of the array. VP i,k1 This is the output voltage under the 1R architecture.
[0032] Please see Figure 3A , Figure 3A This is a schematic diagram illustrating a structure for constructing multiple unit modes according to an embodiment of this application. These unit modes can include a Light-response Neuron (LRN) mode, an Optical Interference Synaptic (OIS) mode, a Long-term Memory Synaptic (LMS) mode, and a Leaky-Integrate-and-Fire Neuron (LIFN) mode. The first two modes can be used for retinal-like neuromorphic perception, while the latter two can be used for brain-like neuromorphic computing. Therefore, the device used in this application can simultaneously meet both front-end intelligent sensing and back-end computing requirements. Further, as... Figure 3B The diagram shows the processing results of input and output data for each unit mode during application.
[0033] It should be noted that different modes employ different unit architectures. For example, the LMS mode uses a 1T1R architecture, where the source of the NMOS transistor is connected to the BL (Bit Lines) line and the gate is connected to the WL (Word Lines) line. The LRN, OIS, and LIFN modes all use a 1R architecture, and the bias voltage is used to switch between the four modes: LRN, OIS, LWS, and LIFN. Therefore, the unit architecture of this application can be freely switched between 1R and 1T1R architectures, allowing for individual control of the operating mode of each device while implementing the array, thus facilitating resource allocation at the subsequent algorithm level.
[0034] It should be noted that each column of the MAC driver module in the multi-mode pixel unit array consists of an NMOS transistor and an operational amplifier. When the multi-mode pixel unit operates in LMS mode, it performs a multiply-and-accumulate (MAC) operation on the input voltage signal to generate a current signal. When the multi-mode pixel unit is set to LMS mode, the MAC driver module operates, converting the current signal generated by the multi-mode pixel unit into a voltage signal for output. When the multi-mode pixel unit is designed in LRN, OIS, or LIFN mode and the MAC driver module is not used, a disable signal is given to the MAC driver module, putting it in an inactive state.
[0035] In some specific embodiments, the unit negative feedback loop is used to provide different bias voltages for the multi-mode intelligent vision device, so that the multi-mode pixel units are set to different unit modes, which may specifically include: The first bias voltage value is provided to the multi-mode intelligent vision device through the unit negative feedback loop, so that the multi-mode pixel unit is set to the photoresponse neuron mode. The unit negative feedback loop provides the voltage of the first bias voltage range to the multi-mode intelligent vision device, so that the multi-mode pixel unit is set to the photoactivated synaptic mode. The second bias voltage range is provided to the multi-mode intelligent vision device through the unit negative feedback loop, so that the multi-mode pixel unit is set to long-term memory synaptic mode. The third bias voltage range is provided to the multi-mode intelligent vision device through the unit negative feedback loop, so that the multi-mode pixel unit is set to the leakage current accumulation discharge neuron mode, wherein the first bias voltage range, the second bias voltage range and the third bias voltage range do not intersect each other.
[0036] In this application embodiment, the voltage across the multi-mode intelligent vision device can be 0V, setting the multi-mode pixel unit to a photoresponsive neuron mode. At 0V, the photoresponsive neuron can receive continuous light and generate electrical pulses. Alternatively, the voltage of the multi-mode intelligent vision device can be any voltage within the range of -0.1V to -0.04V (i.e., the first bias voltage range), setting the multi-mode pixel unit to a photoactivated synapse mode. Another embodiment allows the voltage to be any voltage within the range of 0.1V to 0.5V (i.e., the second bias voltage range), setting the multi-mode pixel unit to a long-term memory synapse mode. Finally, an embodiment allows the voltage to be greater than 1.3V (i.e., the third bias voltage range), setting the multi-mode pixel unit to a leakage current accumulation discharge neuron mode. These voltage ranges do not overlap, ensuring the determinism and reliability of mode switching.
[0037] The multi-mode intelligent vision device employs an indium tin oxide (ITO) / copper oxide (CO) / palladium electrode structure. In LRN mode, the device converts continuous light with intensity and accumulation time reaching a threshold into current pulses, exhibiting intensity-wavelength-dependent pulse emission characteristics. In OIS mode, when stimulated by a sequence of light pulses, the device current initially increases continuously, and then decays non-linearly after stimulation ceases. Higher light pulse intensity results in a larger current amplitude and a slower decay rate, while lower light pulse intensity results in a smaller current amplitude and a faster decay rate. In LMS mode, the device exhibits non-volatile resistive switching behavior; that is, it becomes low-resistance under a constant positive (0.5V, 50ms) voltage pulse and high-resistance under a negative (-0.5V, 50ms) voltage pulse, and the resistance remains unchanged after the voltage pulse stimulation stops. In LIFN mode, the device experiences leakage current accumulation and discharge behavior under continuous voltage pulse stimulation. When the received voltage pulse amplitude is in the range of 1.3 V-2.0 V, the output voltage decays exponentially back to the initial voltage over time. After receiving a continuous input voltage pulse train, the output current gradually increases with the sum of the inputs. Only when the amplitude of the voltage pulse received by the device reaches a threshold (V≥2.0V) will the device generate a current pulse exceeding the critical value, i.e., discharge behavior. Based on this, the first two modes can be used for retinal-like neuromorphic visual perception, and the latter two can be used for visual cortex-like neuromorphic computing. Therefore, the device used in this application can simultaneously meet the requirements of front-end intelligent perception and back-end brain-like computing.
[0038] In some embodiments, different multi-mode pixel unit array configuration methods can be set according to different visual processing needs. There are two types of configuration methods: high-resolution / high-quality intelligent visual perception (NP) configuration method (also known as NP only configuration method or single neuromorphic perception configuration method) or hybrid configuration method of neuromorphic perception NP / neuromorphic computing (NC) (also known as NP / NC hybrid configuration method). According to different configuration methods, the target unit mode of each multi-mode pixel unit can be determined. Different visual processing configuration methods are used to perform different array function partitioning, unit mode allocation and data flow path configuration of the multi-mode pixel unit array, so as to realize the corresponding visual information processing function through the determined array function partitioning, unit mode allocation and data flow path configuration.
[0039] Here, the target unit mode refers to the unit mode corresponding to each multi-mode pixel unit in the multi-mode pixel unit array, determined based on the vision processing configuration. By flexibly adjusting and dividing the unit mode corresponding to each multi-mode pixel unit in the multi-mode pixel unit array, different functions can be implemented on the same chip to achieve high precision, power, and area efficiency (because the entire area of the chip in this application can be used for photosensitive purposes, and the working modes of different areas of the unit array can be flexibly adjusted to improve the utilization of device resources, thereby achieving high area utilization). Array functional partitioning refers to allocating device resources in the multi-mode pixel unit array to different vision processing configurations according to different proportions for functional implementation. Data flow path configuration refers to the fact that different unit architectures may adopt different data flow path configurations.
[0040] The system's visual processing configuration includes a single neuromorphic perception (NP) configuration and a hybrid configuration based on neuromorphic perception NP and neuromorphic computation (NC), i.e., allocating device resources to NP and NC in different proportions. In the system constructed in this application embodiment, the flexibility to reconfigure different NP / NC sizes allows for maximum resource allocation within a unified array, achieving high accuracy, power, and area efficiency compared to vision systems where NP / NC cannot be reconfigured.
[0041] In some embodiments, the step of determining the target unit mode corresponding to each multi-mode pixel unit according to a preset visual processing configuration may specifically include: When the visual processing configuration method adopts the event-driven visual processing configuration algorithm for visual perception, the target unit mode corresponding to all multi-mode pixel units in the multi-mode pixel unit array is set to the light response neuron mode. When the preset visual processing configuration uses a frame-based visual processing algorithm for visual perception, the target unit mode corresponding to all multi-mode pixel units in the multi-mode pixel unit array is set to the light-activated synaptic mode.
[0042] Specifically, this NP-only configuration can include event-driven (pulse-based) visual processing configuration algorithms and frame-based (non-pulse-based) visual processing configuration algorithms. This single-neuron morphological perception configuration only considers multi-mode pixel units in LRN or OIS modes as sensory neurons, responsible for low-to-medium level imaging processing. Each sensory neuron can consist of an LRN or OIS mode multi-mode pixel unit, a bypassable capacitive sampler, and a driver. The bypass capacitive sampler supports reconfiguration of the pulsed or non-pulse NP data flow path. Specifically, the non-pulse OIS mode unit output is sampled at a configurable time, while the LRN mode unit output is directly relayed. The NP processing result can be read from the driver as the sensory neuron output. In this single-neuron morphological perception configuration, the system is primarily used as a high-performance intelligent vision sensor.
[0043] It should be noted that when using an event-driven vision processing configuration algorithm, the system sets the target mode of all multi-mode pixel units in the entire multi-mode pixel unit array to the photoresponse neuron (LRN) mode, forming a massively parallel event camera specifically designed to capture motion information from dynamic scenes. This configuration can be used for event-based high-resolution intelligent imaging. In color-mixing motion, light of different intensities can excite the unit array in LRN mode to generate pulses of different intensities. Therefore, the LRN mode unit array can intelligently capture and distinguish the continuous motion information L of the target vehicle (red and blue cars) within a 0-T period and output a high-resolution voltage pulse signal U. By filtering the voltage pulse signal, the continuous trajectory and shape contour of the target car can be extracted, while filtering out non-target objects (such as a green motorcycle) and static background.
[0044] It should be noted that when using frame-based visual processing algorithms, all units are set to photoactivated synaptic mode, utilizing their attenuation characteristics to suppress noise and enhance the contrast between the image subject and background at specific sampling moments. This approach sets all devices in the array to OIS (Optical Image Switching) mode. This method can also be used for frame-based high-resolution imaging and preprocessing, including image contrast enhancement and image compression. In OIS mode, the devices generate a slower but larger attenuation current for subject pixels than for noise pixels, resulting in a difference in the output voltage signal. During sampling, the current generated by background pixels is essentially zero, while the contrast between the subject and noise is more pronounced than in the input image, thus distinguishing subject and noise pixels and filtering background information.
[0045] In some embodiments, the step of determining the target unit mode corresponding to each multi-mode pixel unit according to a preset visual processing configuration may specifically include: When the visual processing configuration uses the Visual Spiking Artificial Neural Network (V-SANN) algorithm for visual processing, the multi-mode pixel unit array is divided into a photoresponse neuron array (LRN array), a long-term memory synapse array (LMS array), and a leakage current accumulation discharge neuron array (LIFN array), and the multi-mode pixel units contained in different arrays are adjusted according to the corresponding array functions. When the visual processing configuration uses the Visual Spiking Recurrent Neural Network (V-SRNN) algorithm for visual processing, the multi-mode pixel unit array is divided into a light response neuron array (LRN array), a long-term memory synapse array (LMS array), and a leakage current accumulation discharge neuron array (LIFN array), and the multi-mode pixel units contained in different arrays are adjusted according to the corresponding array functions. When the visual processing configuration uses the Visual Artificial Neural Network (V-ANN) algorithm or the Visual Reservoir Computing (V-RC) algorithm for visual processing, the multi-mode pixel unit array is divided into an Optically Activated Synaptic Array (OIS array) and a Long-Term Memory Synaptic Array (LMS array), and the multi-mode pixel units contained in different arrays are adjusted according to the corresponding array functions.
[0046] Specifically, this NP / NC hybrid configuration can include visual spiking artificial neural network algorithms, visual spiking recurrent neural network algorithms, visual artificial neural network algorithms, and visual reservoir processing algorithms. In other words, when the system is configured for NP / NC hybrid processing, the output generated by the NP part can be used as the input to the NC part. This means allocating a portion of the device resources to visual perception for low-to-mid-level processing and imaging, and then using the remaining portion for NC computation for higher-level processing. The NC module consists of LMS-mode synaptic unit arrays and LIFN-mode synaptic unit arrays for efficient MAC operations. For V-SRNN, additional recurrent paths have configurable delays between the synapse and neuron modules. Finally, the synaptic output is passed to the neuron module for decision-making. This module includes a dual-path pulse generator and a unified decision block to efficiently support both pulsed and non-pulsed signals. The dual-path pulse generator uses LIFN-mode units to output pulse sequences during pulsed processing, or uses a voltage-controlled oscillator (VCO) to convert non-pulsed voltage signals into pulse sequences in the non-pulsed case. Decisions are made by calculating the number of output peaks and finding the largest one using the unified decision block.
[0047] It should be noted that when using the V-SANN algorithm, the device resources are divided into three parts: an LRN array, an LMS array, and a LIFN array. The LRN array is used for event-driven visual processing, converting light signals into voltage pulse signals and filtering out noise and static backgrounds. The LMS array can serve as an adjustable synaptic layer of the artificial neural network, and the LIFN array can serve as the output layer of the artificial neural network. This approach can be used for the extraction, filtering, and recognition of event-based mixed-color motion trajectories. Specifically, the NP algorithm is first performed in the region of the LRN array to obtain the voltage pulse signal map of the target's motion trajectory. This obtained voltage pulse signal map is used as the input to the LMS array, where a MAC operation is performed. Next, the voltage pulses are passed to the LIFN array, where voltage pulse signals above a threshold are selected for final processing and output.
[0048] It should be noted that when using the V-SRNN algorithm, the LMS and LIFN arrays are further subdivided based on the V-SANN algorithm. This algorithm adds intermediate layers and back feedback to the V-SANN algorithm, enabling it to support more complex prediction functions. For example, this allows the system to intelligently identify the motion direction of the 11th step while simultaneously predicting the motion direction of the first 10 frames.
[0049] It should be noted that when using the V-ANN algorithm, this algorithm can allocate device resources into two parts: an OIS array and an LMS array. The OIS array is used for frame-based visual processing to obtain a contrast-enhanced voltage map. The obtained voltage map is used as input to the LMS array, where MAC operations are performed to obtain voltage signals of different amplitudes.
[0050] It should be noted that when using the V-RC algorithm, the resource allocation method is similar to that of the V-ANN algorithm, but it utilizes another image compression function of the OIS array. This algorithm can allocate device resources into two parts: the OIS array and the LMS array. The OIS array acts as a reservoir layer, compressing three consecutive light pulses from three adjacent pixels into a single analog conductance state. The compressed image is output as a voltage map. The resulting voltage map is used as input to the LMS array, where a MAC operation is performed to obtain voltage signals of different amplitudes as outputs.
[0051] It should be noted that the algorithms used in the embodiments of this application address the dynamic requirements between imaging-oriented and computation-oriented scenes, as well as the dynamic requirements between event-based motion and frame-based static scenes.
[0052] It should be noted that, in order to facilitate understanding of the specific implementation process of each unit mode, the LRN mode, OIS mode, LMS mode and LIFN mode will be explained in detail below.
[0053] For LRN mode, the devices can perform motion capture, extraction, and encoding filtering on the input image, and the specific process is expressed by the following formula:
[0054] Among them, U LRN This is a diagram of the high-resolution voltage pulse signal output by the array, where U is the high-resolution voltage pulse signal output by the unit, and k is the number of array devices (i.e., the number of multi-mode pixel units contained in the multi-mode pixel unit array). m It is the voltage pulse signal output by the m-th visual unit. This describes the behavior of a multi-mode pixel unit array in LRN mode. L represents the behavior at t0. t d The image contains continuous motion information within its period, therefore L m This represents the pixel acquired by the m-th device. This is a pulse code filtering algorithm, where m is the device index and P is the light intensity. m Let λ represent the intensity of the light collected by the m-th device, and λ be the wavelength of the light. m This represents the wavelength of the light collected by the m-th device.
[0055] It should be noted that when a portion of the cell array is set to LRN mode, the devices in that array exhibit Leaky-Integrate-and-Fire (LIF) pulse characteristics and thresholds based on illumination intensity and accumulation time. Only when the illumination intensity reaches a certain threshold and the accumulation time reaches the emission threshold can the LRN cell generate a sufficiently strong voltage pulse signal. Input t0 t d The LRN mode cell array, containing continuous motion information within a period of time, can output a high-resolution voltage pulse signal U based on illumination intensity and duration. LRN The filtering function of the LRN mode unit on the optical pulse signal can extract the continuous trajectory and shape contour of the target object, filtering out those with a duration shorter than the cumulative time threshold. (Also known as the emission threshold accumulation time) Non-target objects and light intensity below the light intensity threshold The static background. The voltage pulse signal generated by the vision unit can be directly used as the signal input for the next stage, or it can be read from the driver as the module's output. Based on this, the system can accurately extract and encode motion information in real time, effectively reflecting the spatial characteristics and temporal dynamics of the target object. The array's computing power can achieve a frame rate of 10fps with a latency of 0.1s.
[0056] In OIS mode, the device generates a slower but larger decay current for pixels with higher illumination intensity than for pixels with lower illumination intensity, resulting in a difference in the output voltage signal. During sampling (t= At this point, the current generated by pixels with low illumination intensity is essentially zero, making the contrast between different pixels more pronounced. The expression for the input image enhancement or compression process is shown below:
[0057] in, This is a diagram of the high-resolution voltage pulse signal output by the array. and This describes the optical contrast enhancement and data compression behavior of a multi-mode pixel unit array in OIS mode. An algorithm for enhancing device-level optical contrast in OIS mode. This is a device-level data compression algorithm, where k1 and k2 are the number of OIS devices used for contrast enhancement and data compression, respectively, and n c L represents the compression ratio. L is the image to be processed in the input array. Each device receives an optical signal consisting of eight different 3-bit optical flows from "000" to "111", which the device converts into eight different 1-bit conductance states, achieving an image compression ratio of 3.
[0058] For LMS mode, the MAC driver module in the circuit is also set to active. The device performs a MAC operation on the input voltage signal to generate a current signal, which is then converted into a voltage signal by the MAC driver module of each column. This voltage signal can be processed using a multiply-accumulate algorithm. The representation is as follows:
[0059] in, The output voltage pulse signal diagram is shown below. The device-level multiplication algorithm performed for each device operating in LMS mode, where i is the column index, j is the row index, r is the LMS device row number, and n... column This represents the number of device columns.
[0060] For LIFN mode, its output voltage pulse signal It is represented as follows:
[0061] in, The device-level threshold comparison algorithm executed for each device operating in LIFN mode. The voltage pulse output by the LMS array is also the input to the LIFN array, where k is the number of LIFN devices and m is the device index. The LIFN array can be located between two intermediate layers of a neural network. The output voltage pulse can be used as the input to the next layer's calculation or as the output array. The output voltage pulse is directly used for decision-making, and the decision can be made based on the magnitude and number of output pulse peaks.
[0062] For example, please refer to Figure 4 , Figure 4 This is a schematic diagram of a hardware architecture for visual information processing based on a hybrid NP / NC configuration, as provided in an embodiment of this application. The visual processing array is composed of 1A1T1R multi-mode pixel units operating in LRN or OIS mode. The specific circuit structure of the multi-mode neuromorphic visual perception and processing system can be flexibly adjusted by combining the previously mentioned algorithms.
[0063] In some embodiments, different visual processing configurations are used to perform different array functional partitioning, cell mode allocation, and data flow path configurations on the multi-mode pixel unit array, so as to realize the corresponding visual information processing functions through the determined array functional partitioning, cell mode allocation, and data flow path configuration. Here, array functional partitioning refers to the process of determining the cell mode for each multi-mode pixel unit in the multi-mode pixel unit array for both the NP-only configuration and the NP / NC hybrid configuration.
[0064] For example, please refer to Figure 5 , Figure 5 This is a schematic diagram of the data stream processing flow under different configuration methods provided in the embodiments of this application. Wherein, the Figure 5 The data flow processing under different unit modes and algorithm architectures is marked with arrows of different colors.
[0065] The multi-mode neuromorphic visual perception and processing system provided in this application achieves a synergistic improvement in hardware flexibility, energy efficiency, and area efficiency through a fully reconfigurable design from the bottom-level circuit units to the top-level system architecture. Its technical effects are systemic: at the circuit unit level, the innovative 1A1T1R structure and its negative feedback control mechanism enable individual devices to switch precisely and stably between four biomimetic operating modes based on the applied bias voltage, and support dynamic switching of the array between 1R and 1T1R basic architectures, laying the physical foundation for multi-functional integration. At the system architecture level, through preset visual processing configuration methods, the massive unit array is dynamically partitioned, mode-assigned, and data flow-configured, allowing the same physical chip to be reconfigured in real time into dedicated processing systems with vastly different functions, such as ultra-low-power neuromorphic event cameras, high-performance spiking neural network accelerators, or visual perception and computing processors. This allows for multi-functionality, perfectly adapting to diverse visual tasks ranging from high-speed motion detection to complex static recognition. At the information processing level, this method achieves end-to-end non-pulse / pulse domain processing from optical signal perception to intelligent decision-making. Utilizing the physical characteristics of multi-mode intelligent vision devices, it directly performs multiply-accumulate calculations within memory, avoiding energy-intensive data transfer and frequent analog-to-digital conversions. This results in energy efficiency and area efficiency several orders of magnitude higher than existing technologies when performing complex vision tasks. Ultimately, this invention provides a comprehensive intelligent vision hardware solution that combines versatility and specialized efficiency for fields such as edge computing, autonomous driving, and robotics, which have extreme requirements for real-time performance, power consumption, and integration.
[0066] However, there is a lack of corresponding systems to allocate resources according to the dynamic needs of different scenarios for such multifunctional devices. In order to allocate resources for multi-mode pixel unit arrays, an intelligent vision system that can flexibly set the device's operating mode needs to be designed at the circuit, architecture, and algorithm levels.
[0067] Compared with the prior art, the present invention has the following beneficial effects: (1) It realizes a hybrid processing model centered on neuromorphic perception (NP) and neuromorphic perception / neuromorphic computing (NP / NC). This adaptability enables the system provided in this application embodiment to handle various environmental scenarios, which is not possible for devices based on single-function devices.
[0068] (2) In the neuromorphic perception-based processing, the computational energy efficiency of event-driven visual processing in the system provided in the embodiments of this application can reach 9.1 TOPS / W, and the computational energy efficiency of frame-based visual processing can reach 52.6 TOPS / W, which is 86-107 times higher than the currently reported advanced level of CMOS.
[0069] (3) In the neuromorphic computing-based processing, the system provided in the embodiments of this application shows that the computational efficiency of the visual spiking artificial neural network is 2.3 TOPS / W, the computational efficiency of the visual spiking recurrent neural network is 2.7 TOPS / W, the computational efficiency of the visual artificial neural network is 76.5 TOPS / W, and the computational efficiency of the visual reservoir processing is 76.0 TOPS / W.
[0070] (4) The system provided in this application embodiment can uniquely perform hybrid neuromorphic perception / neuromorphic computing processing, wherein the computational energy efficiency of the visual spiking artificial neural network is 2.3 TOPS / W, the computational energy efficiency of the visual spiking recurrent neural network is 2.7 TOPS / W, the computational energy efficiency of the visual artificial neural network is 75.5 TOPS / W, and the computational energy efficiency of the visual reservoir processing is 74.7 TOPS / W, which is 28-922 times higher than the energy efficiency of the most advanced CMOS reported to date. The area efficiency is 2.29×10 4 Up to 7.53×10 5 OPS / F 2 This is 2986 to 3900 times faster than currently reported advanced CMOS neuromorphic vision systems. In summary, the design of this application's embodiments achieves high processing efficiency, with area utilization improved by three orders of magnitude compared to existing reconfigurable vision chips. These results provide further enhancements to the development of neuromorphic vision systems, enabling the expansion of system capabilities to support more complex models and a wider range of visual scenarios.
[0071] Please refer to Figure 6 , Figure 6 This is a schematic diagram of a neuromorphic visual image processing method provided in an embodiment of this application. This neuromorphic visual image processing method can be applied to the multimodal neuromorphic visual perception and processing system described in the above embodiment. Specifically, this neuromorphic visual image processing method may include, but is not limited to, steps S610 to S640.
[0072] Step S610: Obtain the visual image to be processed and the image recognition requirement data; Step S620: Determine the target visual processing configuration method that matches the visual image to be processed based on the image recognition requirement data; Step S630: According to the target visual processing configuration, the multi-mode pixel unit array in the multi-mode neuromorphic visual perception and processing system is systematically adjusted to obtain the target neuromorphic visual processing system. Step S640: Perform image processing on the visual image to be processed according to the target neuromorphic visual processing system.
[0073] In steps S610 and S620 of some embodiments, this application embodiment can receive real-time light signals from an optical lens or read digital image data from a memory to obtain a visual image to be processed. These visual images to be processed can be used to process specific real-world targets, such as detecting moving vehicles or recognizing traffic signs. Furthermore, this application embodiment can determine a target visual processing configuration that matches the visual image to be processed based on image recognition requirement data. For example, for the requirement of detecting moving vehicles, a visual spiking artificial neural network algorithm is selected as the target configuration.
[0074] In step S630 of some embodiments, the embodiments of this application may perform system adjustments on multiple multi-mode pixel units in a multi-mode neuromorphic visual perception and processing system according to the target visual processing configuration. This means that the system controller loads the instruction set corresponding to the configuration, sends corresponding voltage control signals and mode switching signals to all units in the array, and performs architecture switching, mode allocation, and data stream connection operations, thereby transforming a general-purpose hardware platform into the target neuromorphic visual processing system, that is, instantiating it into a dedicated processing engine optimized for the current task.
[0075] In step S640 of some embodiments, the embodiments of this application can perform image processing on the visual image to be processed according to the target neuromorphic visual processing system. This means inputting the image data into the reconstructed hardware system and using its pre-configured specific sensor-computer integrated pathway to complete the entire process from raw data to advanced information.
[0076] Please refer to Figure 7 , Figure 7 This is a schematic diagram of step S640 provided in an embodiment of this application. Step S640 may specifically include, but is not limited to, steps S710 to S720.
[0077] Step S710: When the target visual processing configuration method adopts the event-driven visual processing configuration algorithm for visual processing, the visual image to be processed is input into the target neuromorphic visual processing system, and the visual image to be processed is regarded as a light pulse signal image. Step S720: Filter the light pulse signal diagram according to the device's light intensity threshold and cumulative time threshold to obtain the first target voltage pulse signal diagram.
[0078] When the target visual processing configuration uses an event-driven visual processing configuration algorithm, the entire multi-mode pixel unit array is in full LRN mode. Image processing at this time can be image perception (high-resolution intelligent imaging) processing. At the start of processing, this embodiment can first input the visual image to be processed into the target neuromorphic visual processing system, that is, project the continuous light signal of the scene onto the photosensitive area of the array. The array works in parallel and outputs a first target voltage pulse signal map. This map is not a traditional pixel array, but a sparse set composed of a large number of asynchronous pulse events. Each pulse marks a brightness change exceeding the basic sensitivity at a specific location at a particular time.
[0079] Exceeding the baseline sensitivity means that the input light pulse reaches the device's illumination intensity threshold and cumulative time threshold. Only pulses generated by pixels with sufficiently strong light intensity (e.g., representing reflected light from a moving vehicle) and a sufficiently long duration (excluding brief flash noise) will be recognized and retained by the system. Through this dual threshold filtering, environmental background noise and transient interference can be effectively eliminated, ultimately yielding a first target voltage pulse signal map that clearly and purely depicts the trajectory outline of the moving target.
[0080] For example, such as Figure 8 The diagram illustrates an event-driven visual processing configuration algorithm provided in this application for visual processing. Specifically, when the array element count is 13456, because the LRN mode device requires a certain light intensity threshold and a cumulative time threshold for emission, the device can only extract the input intensity (…). > ) and cumulative time ( > When the input light intensity is above the threshold for a moving target vehicle, a voltage pulse is generated. Conversely, when the input light intensity is below the threshold for a static dark background (where the input light intensity is below the threshold), a voltage pulse is generated. > but < ) and green locomotives with accumulated time below the launch threshold ( > but < No voltage pulses are generated, thus effectively filtering them out. Based on this, the system can accurately and in real-time extract and encode the motion information of the red and blue cars, effectively reflecting the spatial characteristics and temporal dynamics of the target objects. Therefore, the input visual image to be processed undergoes motion capture, extraction, encoding, and filtering, and the specific process is expressed by the following formula:
[0081] In some embodiments, when the preset visual processing configuration uses a frame-based visual processing algorithm for visual processing, the embodiments of this application can input the visual image to be processed into a light-activated synaptic array for image enhancement or compression, resulting in an enhanced or compressed voltage non-pulse signal image. For example... Figure 9 The diagram illustrates a frame-based visual processing configuration algorithm provided in this application for visual processing. When the target visual processing configuration uses the frame-based visual processing configuration algorithm, the device generates a slower but larger attenuation current for the subject pixel than for the noise pixel, thus making the contrast between the subject and the noise more pronounced. During sampling ( When the background pixels have essentially reached zero, the current generated by them is essentially zero, thus filtering out background information. Based on this, the specific process of enhancing or compressing the input visual image at this point is expressed by the following formula:
[0082] Please refer to Figure 10 , Figure 10 This is a schematic diagram of step S640 provided in an embodiment of this application. Step S640 may specifically include, but is not limited to, steps S1010 to S1040.
[0083] Step S1010: When the target visual processing configuration adopts the visual spur artificial neural network algorithm for visual processing, the multi-mode pixel unit array in the target neuromorphic visual processing system includes a light response neuron array, a long-term memory synapse array, and a leakage current accumulation discharge neuron array. Step S1020: Input the visual image to be processed into the optical response neuron array for image optical signal conversion to obtain the second voltage pulse signal image; Step S1030: Input the second voltage pulse signal diagram into the long-term memory synaptic array for weighted fusion to obtain the third voltage pulse signal diagram; Step S1040: Input the third voltage pulse signal graph into the leakage current accumulation discharge neuron array for unit accumulation to obtain the target voltage signal.
[0084] In this case, when the target visual processing configuration uses the visual spiking artificial neural network algorithm for visual processing, the image processing process is equivalent to the image recognition process (low-order, mid-order, and high-order). For example... Figure 11The diagram illustrates a visual processing algorithm based on a visual spiking artificial neural network (LSN) provided in this application. When the target visual processing configuration employs the LRN algorithm, the NP algorithm is first performed in the LRN array region to extract, encode, and filter the continuous color mixing motion information (L, 0-T) of the target vehicle with a size of 20×20, thus obtaining the target vehicle's motion over the time interval t0-t0. 20 Second voltage pulse signal diagram of motion trajectory The specific implementation process is shown in the following formula:
[0085] Furthermore, the obtained second voltage pulse signal can be input into the LMS array for multiply-accumulate operations. The algorithm for this operation is derived from the following formula, where Device-level multiplication algorithm executed for each device operating in LMS mode.
[0086]
[0087] Furthermore, the resulting third voltage pulse signal diagram is passed to the LIFN array, where voltage pulse signals above a threshold are selected for final processing and output. The processing procedure of the LIFN array algorithm can be found in the following formula. The threshold comparison algorithm executed for each device operating in LIFN mode ultimately outputs the following voltage pulse signal:
[0088] It should be noted that the LIFN array, which serves as the output neuron, has a total of 12 neurons, corresponding to 12 different categories. The final result is the pulse rate identified by the LIFN array. The trained system achieved a 91.7% recognition accuracy across 12 mixed-color vehicle motion categories.
[0089] In some embodiments, when using the V-SRNN algorithm, an LRN array of size 21×21 can be set, and the LMS array and LIFN array can be further subdivided. For example, Figure 12The diagram illustrates an algorithm for visual processing based on a visual impulse recurrent neural network (LRN) provided in this embodiment. This embodiment divides the LMS array into three parts: 32×8 fully connected (FC) synapses, 441×32 recurrent forward connection (RFC) synapses, and 32×32 backward recurrent connection (BRC) synapses. The LIFN array is divided into two parts, with 8 neurons serving as output neurons and 32 as recurrent neurons. The computational structure of the LRN array in this case is shown in the following formula:
[0090] Wherein, each input motion image (t0 to t) 10 The sequence comprises 10 consecutive motion steps, each occupying one cycle. The voltage pulse diagram of the motion trajectory output by the LRN array for the first 10 frames is as follows: The obtained voltage pulse diagram The calculation process for the three LMS arrays, which are used as inputs, is as follows:
[0091] Where U', R, and U''' represent the MAC output voltage pulses of the RFC, BRC, and FC synapses, respectively. Step 11 is... arrive The predicted direction of motion is determined by arrive The period has the highest pulse count The output neuron determines the motion trajectory. The motion trajectory dataset contains 1640 different motion images with 8 possible motion directions. 1312 pulse images were used for training, and the remaining 328 pulse images were used as the test set. Testing showed that this configuration achieved a prediction accuracy of 92.5%.
[0092] In some embodiments, when the V-ANN algorithm is used, step S640 may further include the following steps: When the target visual processing configuration adopts a visual artificial neural network algorithm for visual processing, the multi-mode pixel unit array in the target neuromorphic visual processing system includes a photoactivated synaptic array and a long-term memory synaptic array. The visual image to be processed is input into a light-activated synaptic array for image enhancement, resulting in an enhanced voltage non-pulse signal image. The enhanced voltage non-pulse signal map is input into a long-term memory synaptic array for MAC calculation to obtain the third target voltage non-pulse signal map.
[0093] The configuration method of the OIS array can refer to frame-based visual processing. Inputting a traffic sign image L yields a contrast-enhanced voltage non-pulse signal image. This traffic sign image dataset contains 8 classes of 28×28 noisy traffic sign images, such as warning electric, bumpy, hazard, low temperature freezing, slippery road surface, harmful, CCTV operation, and highly flammable. Image noise includes perspective distortion (application probability of 50%) and salt-and-pepper noise (amount=0.5). The obtained voltage non-pulse signal image is used as input to an LMS array, and MAC operation is performed on the LMS array to obtain voltage signals of different amplitudes. This is to determine the non-pulse signal diagram of the third target voltage. Based on this, the specific calculation process is as follows:
[0094] For example, such as Figure 13 The diagram illustrates an algorithm for visual processing based on a visual artificial neural network algorithm provided in this application embodiment. Based on experimental measurement results of OIS-patterned devices, a simulated OIS-patterned cell array behavior dataset (28×28 voltage maps) was generated. Period-to-period and cell-to-cell variations were extracted and modeled as Gaussian noise, which was added during the simulation to simulate real device level variations. The OIS-patterned cell array behavior dataset was used to train the NC part in this configuration. In the OIS-patterned cell array behavior dataset, 192 samples were used for training, 39 for validation, and 48 for testing. Training used the Adam optimizer (learning rate 0.001, β1=0.9, β2=0.999, ε=1×10⁻⁶). -8 Training was performed using CrossEntropyLoss. A maximum of 50 epochs were used, with early stopping based on validation loss and a patience period of 15 epochs. The final model was selected from the historical models that achieved the highest validation accuracy before any overfitting was observed. Training accuracy and training loss were monitored simultaneously to avoid underfitting. The voltage signal with the largest amplitude in the final output represents the final classification category with the highest probability. This system can identify 8 types of noisy traffic sign images with an accuracy of 97%, which is 18% higher than the accuracy of traditional processing systems without contrast enhancement.
[0095] In some embodiments, when the V-RC algorithm is used, another image compression function of the OIS array is utilized. This configuration allocates device resources to two parts: the OIS array (729 = 27×27) and the LMS array (729×6). The OIS array, as a reservoir layer, can compress three consecutive light pulses of three adjacent pixels into a single analog conductance state.
[0096] For example, such as Figure 14 The diagram shown is a schematic of a visual processing algorithm based on a visual reservoir processing algorithm provided in this application embodiment. The dynamic traffic image L (27×27×3) is compressed to 27×27×1, and the compressed image is represented as a voltage map. The output is in the form of [database name missing]. The traffic image dataset used here includes six categories of noisy traffic images: buses, bicycles, cars, trucks, signs, and pedestrians. Image noise includes perspective distortion, random rotation (angle ∈ [-15°, 15°]), and salt-and-pepper noise (amount = 0.2), with each type of noise having a 50% probability of application. Furthermore, the obtained voltage map can be used as input to an LMS array, where a MAC operation is performed to obtain voltage signals of different amplitudes. Implement the output layer. Based on this, the specific calculation process is as follows:
[0097] During the simulation, multimodal inter-pixel noise and periodic noise during compression were combined into Gaussian noise extracted from small-scale measurements. Of the entire dataset, 480 images were used for training, 115 for validation, and 120 for testing. Training used the Adam optimizer (β1=0.9, β2=0.999, ε=1×10⁻⁶). -8 The learning rate is 0.001, and the CrossEntropyLoss is also considered. The model is trained for a maximum of 30 epochs. Validation loss is monitored in each epoch, with a patience period of 15 epochs. Training accuracy and training loss are monitored simultaneously to avoid underfitting. After training, this configuration can classify six classes of noisy traffic images with a classification accuracy of 95.8%. Therefore, this embodiment of the application achieves NP-centric and NP / NC hybrid processing. This adaptability enables the system to handle various environmental scenarios, which is unattainable for previous devices based on single-function components.
[0098] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0099] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0101] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0102] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0103] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0104] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0105] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0106] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0107] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A multimodal neuromorphic visual perception and processing system, characterized in that, The system includes a multi-mode pixel unit array composed of a row and column structure. Each multi-mode pixel unit in the multi-mode pixel unit array includes a multi-mode intelligent vision device and a multi-mode control circuit composed of a voltage source, an amplifier, and a transistor. The voltage source is used to provide input voltage to the multi-mode pixel unit in order to adjust the unit architecture of the multi-mode pixel unit; The amplifier and transistor in the multi-mode pixel unit constitute a unit negative feedback loop, and the unit negative feedback loop is used to provide different bias voltages for the multi-mode intelligent vision device, so that the multi-mode pixel unit is set to different unit modes. The target unit mode corresponding to each multi-mode pixel unit is determined according to the preset visual processing configuration method. Different visual processing configuration methods are used to perform different array function partitioning, unit mode allocation and data flow path configuration on the multi-mode pixel unit array, so as to realize the corresponding visual information processing function through the determined array function partitioning, unit mode allocation and data flow path configuration.
2. The system according to claim 1, characterized in that, The voltage source is used to provide an input voltage to the multi-mode pixel unit to adjust the unit architecture of the multi-mode pixel unit, including: The voltage source provides an enable signal to the amplifier in the multi-mode pixel unit, and the unit architecture of the multi-mode pixel unit is adjusted to the first circuit architecture. The first circuit architecture enables the multi-mode pixel unit array to form a 1R array architecture. Then, the multi-mode intelligent vision device generates a current signal, and the amplifier and transistor maintain the negative voltage of the device through negative feedback and convert the current into voltage output. By providing a disable signal to the amplifier in the multi-mode pixel unit through the voltage source, the unit architecture of the multi-mode pixel unit is adjusted to a second circuit architecture, which enables the multi-mode pixel unit array to form a 1T1R cross array architecture.
3. The system according to claim 1, characterized in that, The unit modes include photoresponsive neuron mode, photoactivated synapse mode, long-term memory synapse mode, and leakage current accumulation discharge neuron mode; The unit negative feedback loop is used to provide different bias voltages for the multi-mode intelligent vision device, so that the multi-mode pixel unit is set to different unit modes, including: The unit negative feedback loop provides a first bias voltage value to the multi-mode intelligent vision device, thereby setting the multi-mode pixel unit to the photoresponse neuron mode. The unit negative feedback loop provides a voltage in the first bias voltage range to the multi-mode intelligent vision device, thereby setting the multi-mode pixel unit to photoactivated synaptic mode. The unit negative feedback loop provides a second bias voltage range to the multi-mode intelligent vision device, thereby setting the multi-mode pixel unit to a long-term memory synaptic mode. The unit negative feedback loop provides a third bias voltage range to the multi-mode intelligent vision device, enabling the multi-mode pixel unit to be set to leakage current accumulation discharge neuron mode, wherein the first bias voltage range, the second bias voltage range and the third bias voltage range do not intersect each other.
4. The system according to claim 3, characterized in that, The visual processing configuration method includes a single-neural morphological perception configuration method, and the single-neural morphological perception configuration method includes an event-driven visual processing configuration algorithm and a frame-based visual processing algorithm. The step of determining the target unit mode corresponding to each multi-mode pixel unit according to a preset visual processing configuration includes: When the visual processing configuration method uses an event-driven visual processing configuration algorithm for visual processing, the target mode corresponding to all the multi-mode pixel units in the multi-mode pixel unit array is set to the light response neuron mode. When the preset visual processing configuration uses a frame-based visual processing algorithm for visual processing, the target mode corresponding to all the multi-mode pixel units in the multi-mode pixel unit array is set to the photoactivated synaptic mode.
5. The system according to claim 4, characterized in that, The visual processing configuration also includes a hybrid configuration based on neuromorphic perception and neuromorphic computing, and the hybrid configuration based on neuromorphic perception and neuromorphic computing includes visual spiking artificial neural network algorithm, visual spiking recurrent neural network algorithm, visual artificial neural network algorithm and visual reservoir processing algorithm. The step of determining the target unit mode corresponding to each multi-mode pixel unit according to a preset configuration method includes: When the visual processing configuration uses the visual impulse artificial neural network algorithm for visual recognition, the multi-mode pixel unit array is divided into a photoresponse neuron array, a long-term memory synapse array, and a leakage current accumulation discharge neuron array, and the multi-mode pixel units contained in different arrays are adjusted in unit mode based on the corresponding array function. When the visual processing configuration uses the visual pulse recurrent neural network algorithm for visual recognition, the multi-mode pixel unit array is divided into a photoresponse neuron array, a long-term memory synapse array, and a leakage current accumulation discharge neuron array, and the multi-mode pixel units contained in different arrays are adjusted in unit mode based on the corresponding array function. When the visual processing configuration uses a visual artificial neural network algorithm or a visual reservoir processing algorithm for visual recognition, the multi-mode pixel unit array is divided into a photoactivated synaptic array and a long-term memory synaptic array, and the multi-mode pixel units contained in different arrays are adjusted in unit mode based on the corresponding array function.
6. A neuromorphic visual image processing method, characterized in that, The method includes: Acquire the visual images to be processed and the image recognition requirements data; Determine the target visual processing configuration that matches the visual image to be processed based on the image recognition requirement data; According to the target visual processing configuration, the multi-mode pixel unit array in the multi-mode neuromorphic visual perception and processing system according to any one of claims 1 to 5 is systematically adjusted to obtain the target neuromorphic visual processing system. The target neuromorphic visual processing system performs image processing on the visual image to be processed.
7. The method according to claim 6, characterized in that, The image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the target visual processing configuration method adopts an event-driven visual processing configuration algorithm for visual processing, the visual image to be processed is input into the target neuromorphic visual processing system, and the visual image to be processed is regarded as a light pulse signal image. The optical pulse signal image is filtered based on the device's illumination intensity threshold and cumulative time threshold to obtain the first target voltage pulse signal image.
8. The method according to claim 6, characterized in that, The image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the preset visual processing configuration uses a frame-based visual processing algorithm for visual processing, the visual image to be processed is input into a photoactivated synapse array of a multi-mode pixel unit array for image processing, resulting in a processed voltage non-pulse signal image.
9. The method according to claim 6, characterized in that, The event-based image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the target visual processing configuration adopts the visual spiking artificial neural network algorithm for visual processing, the multi-mode pixel unit array in the target neuromorphic visual processing system includes a photoresponse neuron array, a long-term memory synapse array, and a leakage current accumulation discharge neuron array. The visual image to be processed is input into the optical response neuron array for image optical signal conversion to obtain a second voltage pulse signal image; The second voltage pulse signal diagram is input into the long-term memory synaptic array for weighted fusion to obtain the third voltage pulse signal diagram; The third voltage pulse signal is input into the leakage current accumulation discharge neuron array for unit accumulation to obtain the target voltage signal.
10. The method according to claim 6, characterized in that, The frame-based image processing of the visual image to be processed according to the target neuromorphic visual processing system includes: When the target visual processing configuration uses a visual artificial neural network algorithm for visual processing, the multi-mode pixel unit array in the target neuromorphic visual processing system includes a photoactivated synapse array and a long-term memory synapse array. The visual image to be processed is input into the photoactivated synapse array for image enhancement, resulting in an enhanced voltage non-pulse signal image. The enhanced voltage non-pulse signal map is input into the long-term memory synaptic array for weighted fusion to obtain the third target voltage non-pulse signal map.