Multi-dimensional spectrum real-time voiceprint recognition system based on bionic physical reserve pool optical calculation
Through ITO/MoS2/Ag heterogeneous memristor array and optical-electrical collaboration strategy, a multi-dimensional spectrum real-time voiceprint recognition system is built, which solves the hardware and algorithm constraints of the existing voiceprint recognition system in cross-species voiceprint processing, and realizes high-precision and low-power cross-species voiceprint recognition, supporting real-time and long-term monitoring.
Patent Information
- Application Number
- CN202510521347.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
When handling multi-dimensional cross-species voiceprints, existing voiceprint recognition systems are limited by the rigid constraints of hardware architecture and the generalization bottleneck of algorithm models, and are difficult to meet the needs of real-time, low power consumption and high precision collaborative optimization. They lack standardized collection of cross-species voiceprint data and fusion mechanisms for multi-source heterogeneous data, and are susceptible to environmental noise interference, and there are problems of high misjudgment rate and trade-offs on energy consumption and real-time.
Using ITO/MoS2/Ag heteromemristor array and combined with the photo-electrical collaboration strategy, a multi-dimensional spectrum real-time voiceprint recognition system is built. Through the MoS2 interlayer exciton effect and the surface plasmon of Ag nanoparticles, wide spectrum light response and ultra-high photoelectric conversion efficiency are realized. A cross-species voiceprint database is designed and the time-frequency joint masking algorithm is used to support real-time stream processing of high-frequency soundprint signals.
It has achieved more than 94% accuracy of cross-species voiceprint recognition, has ultra-high virtual node and multidimensional spectrum resolution capabilities, optical-electrical collaborative programming and ultra-low energy consumption architecture, supports long-term unattended wearable devices and field monitoring nodes, reducing the misjudgment rate and energy consumption.
Smart Images

Figure CN120452448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biometric identification technology, and in particular to a multi-dimensional spectrum real-time voiceprint recognition system based on bionic physical reservoir optical calculation. Background Art
[0002] With the deep integration of artificial intelligence and the Internet of Things (IoT), biometrics are gradually evolving from a single modality to multi-dimensional perception. Voiceprint recognition, with its unique non-contact acquisition method and rich biometric information-carrying capacity, has become a key technological breakthrough in smart security, ecological monitoring, healthcare, and other fields. The essence of voiceprints is to identify the identity of individuals or groups of organisms by analyzing the spectrum, time-frequency modulation, and nonlinear characteristics contained in sound signals. This technology has expanded from traditional human speech recognition to the analysis of the acoustic behavior of plants and animals, providing a new technical path for urban ecological planning, precision agriculture pest warning, and even the protection of endangered species.
[0003] However, existing voiceprint recognition systems are still limited by the rigid constraints of hardware architecture and the generalization bottlenecks of algorithmic models when dealing with multi-dimensional, cross-species voiceprints, making it difficult to meet the requirements of coordinated optimization of real-time performance, low power consumption, and high precision. Current mainstream voiceprint recognition systems rely primarily on traditional von Neumann architecture electronic computing platforms (such as CPUs and GPUs) combined with deep learning algorithms. Their technical approach suffers from multiple inherent flaws. At the feature extraction level, while classic methods based on Mel-Frequency Cepstral Coefficients (MFCCs) or Linear Predictive Coding (LPC) can effectively characterize the formant characteristics of human speech, their spectral resolution capabilities for low-frequency plant or high-frequency animal voiceprints are significantly reduced, and they are susceptible to interference from environmental noise (such as wind, rain, and mechanical vibrations), leading to feature aliasing. At the algorithmic level, while recurrent neural networks (RNNs) and their variants (such as LSTMs and GRUs) can handle long-range dependencies in time series signals, the vanishing gradient problem and a parameter size of up to 106 puts these models under pressure from both storage and computing power when deployed on edge devices. Practical applications often require cloud-based collaborative computing, resulting in data transmission delays (typically exceeding 50ms) and privacy risks. More critically, existing systems' biometric databases are mostly limited to human voice samples. There is a lack of standards for collecting cross-species voiceprint data, and mechanisms for fusing heterogeneous multi-source data are imperfect (e.g., differences in time-frequency resolution). This results in insufficient generalization when identifying African elephant infrasound (14-24Hz) or coral reef biome soundscapes, resulting in recognition accuracy rates dropping by over 30% compared to laboratory settings.
[0004] In recent years, breakthroughs in neuromorphic computing technology have provided a new hardware implementation paradigm for voiceprint recognition. Reservoir computing (RC) based on memristors has attracted considerable attention due to its ability to simulate biological synaptic plasticity and its advantages in parallel multiply-accumulate operations. Research has demonstrated the use of titanium oxide (TiOx) dynamic random access memory (DRRAM) to build a 200-node reservoir network, achieving handwritten digit classification (with 92% accuracy) and arrhythmia detection (with 89% sensitivity) by leveraging the spatiotemporal dynamics of silver nanowire network memristors. However, these memristor systems based on electrical pulse modulation face fundamental limitations. First, the number of virtual nodes is limited by the random nature of oxygen vacancy migration at the device interface, typically enabling only a few hundred (<600) separable intermediate conductance states, making it difficult to achieve the high-resolution encoding required for multidimensional soundprint signals (e.g., a three-dimensional spectrum containing the fundamental frequency, harmonics, and resonance peaks). Second, the upper limit of the optical response frequency is generally below 103 Hz, making it impossible to resolve transient acoustic events in the ultrasonic frequency band in real time. Third, existing memristors mostly use a single metal oxide structure, resulting in weak photoelectric synergy effects. The photocurrent relaxation time and dark current stability make it difficult to support continuous, long-term, cross-species soundprint monitoring. For example, while ZnOx / HfOx heterojunction devices can mimic the integrated sensing, storage, and computing functions of the retina, their photoconductive gain and wavelength sensitivity (responding only to ultraviolet light) limit their applicability in complex lighting environments. Furthermore, they lack circuit designs tailored to low-frequency plant sound waves (e.g., preamplification of sub-hertz signals).
[0005] Furthermore, in terms of security, existing voiceprint recognition systems generally lack liveness detection and anti-forgery attack capabilities. Synthetic speech generation technologies (such as WaveNet and Tacotron) can accurately reproduce the target voiceprint's spectral characteristics through deep learning models, while recording replay attacks based on MEMS microphones can easily bypass traditional spectral threshold detection. Research has shown that commercial voiceprint authentication systems can experience a sudden increase in false acceptance rates from 0.1% to 12% when subjected to adversarial sample attacks, seriously threatening their credibility in financial payment and forensic evidence collection scenarios. Furthermore, the inherent analog-to-digital conversion bottlenecks and clock synchronization errors of traditional CMOS platforms further exacerbate system vulnerabilities. Signal distortion significantly increases the false positive rate when processing millisecond-level voiceprint transients. Another core contradiction in existing technologies is the trade-off between energy consumption and real-time performance. Deep learning voiceprint recognition systems based on GPU clusters consume up to 300W of power per inference and have latency exceeding 5ms, making them difficult to meet the energy constraints of wearable devices or field monitoring nodes. While some studies have attempted to reduce computational overhead through algorithm lightweighting, these generally result in accuracy losses exceeding 15%. On the other hand, although the theoretical energy efficiency ratio of memristor reservoir calculation is significantly better than that of traditional architectures, it is limited by the non-ideal characteristics of the device (such as conductivity relaxation and cycle durability). In actual systems, complex calibration circuits and redundant nodes are still required, which greatly reduces the energy efficiency advantage. For example, although the reservoir system constructed with SnS-based optoelectronic devices can achieve fingerprint classification (with an accuracy of 94%), its optical power density is as high as 103μW / μm. 2 , and requires periodic electrical reset operations, which makes it difficult to support unattended long-term voiceprint monitoring tasks. Summary of the Invention
[0006] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a multi-dimensional spectrum real-time voiceprint recognition system based on bionic physical reservoir optical computing, which breaks through the multi-dimensional constraints of the existing technology through material innovation, architecture optimization and algorithm collaboration.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A multi-dimensional spectrum real-time voiceprint recognition system based on bionic physical reservoir optical computing, the key of which is: including a host computer, a microprocessor and multiple synchronous sampling modules, each synchronous sampling module includes a DAC laser trigger unit, a laser, an ORRAM optical computing array, a signal amplification unit, and an ADC sampling unit, wherein: The host computer is used to pre-process the voiceprint signal to be identified obtained from the multi-biological voiceprint signal database to generate a feature data array; The microprocessor is used to generate a control signal according to the characteristic data array; The DAC laser trigger unit is used to output a timing voltage pulse whose pulse amplitude linearly corresponds to the characteristic intensity of the voiceprint signal to be identified according to the control signal; The laser is used to generate a sequence of optical pulses based on time-sequential voltage pulses; The ORRAM optical computing array is composed of an ITO / MoS2 / Ag heterojunction memristor, which is used to convert the optical pulse sequence into a photocurrent response signal through the synergistic effect of the MoS2 interlayer exciton effect and the surface plasmon resonance of Ag nanoparticles; The signal amplification unit is used to amplify the photocurrent response signal output by the ORRAM optical computing array and convert it into a voltage signal; The ADC sampling unit is used to sample the voltage signal output by the signal amplifying unit in real time and convert it into a digital model; The host computer is further configured to complete voiceprint recognition based on the sampling data of the multiple synchronous sampling modules.
[0008] Furthermore, the data stored in the multi-biological voiceprint signal database are three-dimensional spectrum features of human language, plant voiceprint signals, and animal voiceprint signals.
[0009] Furthermore, the specific steps of the host computer preprocessing the voiceprint signal to be identified are as follows: Downsample and perform preliminary feature extraction on the voiceprint signal to obtain the spatiotemporal characteristics of the signal; The continuous wavelet transform method is used to obtain the frequency characteristics of the voiceprint signal to be identified; The mask length and the number of matrices are constructed, each matrix is filled with random simulation values, and the spatiotemporal features are multiplied with the frequency features to obtain the feature data array.
[0010] Furthermore, the preparation method of the ITO / MoS2 / Ag heterojunction memristor is as follows: ITO array substrate cleaning; Using magnetron sputtering, metal material is deposited on the surface of the ITO array substrate to form a bottom electrode; Molybdenum disulfide is deposited on the bottom electrode to form a MoS2 resistive switching functional layer; In less than 5.0×10 -4 Under the high vacuum environment of Pa, a silver metal material is deposited on the resistive switching functional layer by DC magnetron sputtering to form a top electrode; A multi-electrode array structure is formed on the top electrode using a mask plate to obtain the ITO / MoS2 / Ag heterojunction memristor.
[0011] Furthermore, the steps of obtaining the photocurrent response signal of the ORRAM optical computing array are as follows: Based on the memristor reservoir algorithm, after each light pulse cycle, the reservoir state x(t) is captured and all x(t) values are integrated into the reservoir state matrix X; Get the weight matrix Wout obtained after linear regression training; The output photocurrent response signal Y is calculated according to the formula Y=X×Wout.
[0012] Furthermore, the microprocessor is also used to control the DAC laser trigger unit to output an electrical reset pulse to instantaneously restore the ORRAM optical computing array to an initial state.
[0013] Furthermore, the synchronous sampling module also includes an inverting proportional amplifier unit and a multiplexer. The inverting proportional amplifier unit is used to convert the timing voltage pulse of the positive voltage into an equal-amplitude negative pulse. The multiplexer is used to dynamically switch the output channel to output positive pulses or negative pulses, thereby realizing the regulation of the frequency and light intensity of the laser.
[0014] Furthermore, the host computer and the microcontroller are connected via a USART serial port, and the DAC laser trigger unit and the ADC sampling unit are both connected via a control bus and the USART serial port.
[0015] Furthermore, the microprocessor adopts a minimum system based on an STM32 single-chip microcomputer, and the signal amplification unit adopts a transimpedance amplifier.
[0016] This invention breaks through the multi-dimensional constraints of existing technologies through material innovation, architecture optimization and algorithm collaboration. At the device level, it uses an ITO / MoS2 / Ag heterojunction memristor array, utilizing the interlayer exciton effect of MoS2 and the surface plasmon resonance of Ag nanoparticles to achieve a wide spectrum light response (405nm-850nm) and ultra-high photoelectric conversion efficiency. Combined with the transparent conductive properties of the ITO electrode, a single device can achieve a power output of 0.66μW / μm 2Under high-intensity light, 1458 separable photoconductive states (equivalent to 10-bit precision) can be generated, quadrupling the number of virtual nodes in traditional oxide memristors. At the system architecture level, an innovative opto-electrical synergy strategy is introduced. Simultaneously, the rapid ion migration characteristics of the Ag electrode are utilized to instantaneously reset the device state with a negative voltage pulse, compressing the system refresh cycle to the microsecond level, thereby supporting real-time stream processing of 1MHz high-frequency voiceprint signals. Furthermore, by constructing a cross-species voiceprint database (covering the three-dimensional spectral characteristics of human speech, plant voiceprint signals, and animal voiceprint signals) and designing a joint time-frequency masking algorithm, the system can adaptively extract significant features of different bio-voiceprints (such as the continuity of formants in human speech and the intermittent pulses of plant sounds). Ultimately, linear regression at the output layer of the reservoir achieves a cross-species recognition accuracy of over 94%, providing a new hardware solution for fields such as smart ecology and biomedical monitoring.
[0017] The remarkable effects of the present invention are: 1. This system has ultra-high virtual node and multi-dimensional spectrum analysis capabilities To address the limited spectral resolution caused by the insufficient number of virtual nodes in existing memristor systems (usually <600), this invention uses an ITO / MoS2 / Ag heterojunction memristor array. Through the synergistic effect of the MoS2 interlayer exciton effect and the surface plasmon resonance of Ag nanoparticles, a single device can achieve 1458 separable photoconductive states (equivalent to 10-bit precision), which is more than 5 times higher than that of traditional TiOx or WOx memristors (about 200 nodes). 2 ) driven, the device's optical response frequency is extended to 1MHz, which can analyze the transient characteristics of ultrasonic frequency bands (such as bat echolocation 150kHz) and sub-hertz plant sound waves in real time, and its spectrum coverage is much wider than that of existing optoelectronic devices (up to 10 4 Experiments show that this system significantly optimizes the extraction of three-dimensional spectrum (fundamental frequency, harmonics, and formant) features of cross-species voiceprints compared to traditional MFCC-based methods.
[0018] 2. This system has optical-electrical collaborative programming and ultra-low energy consumption architecture The voiceprint recognition system of traditional CMOS platform (such as GPU) relies on cloud computing, with a single inference power consumption of up to 300W and a delay of more than 1ms. Although the existing memristor system has high theoretical energy efficiency, it is limited by the conductivity relaxation (>10ms) and reset cycle (>1s), and the actual energy efficiency ratio is only 10 3 -10 4TOPS / W. This invention innovatively introduces an optical-electrical collaborative programming strategy: light pulses trigger proton-coupled charge transfer to complete signal encoding. At the same time, the fast ion migration characteristics of the Ag electrode are utilized to instantaneously reset the device state with a -1.5V pulse, compressing the system refresh cycle to the microsecond level. It also supports continuous stream processing of 1MHz high-frequency voiceprint signals, meeting the long-term unattended needs of wearable devices and field monitoring nodes.
[0019] 3. This system realizes cross-species voiceprint fusion and hardware-algorithm collaborative optimization Existing technologies lack a standardized cross-species database and multi-source data fusion mechanism, resulting in insufficient generalization capabilities for plant and animal voiceprint recognition. This invention utilizes a dynamic nonlinear mapping mechanism based on reservoir calculations, combined with a time-frequency domain joint masking algorithm, to extract and identify voiceprint features. Furthermore, the cross-species voiceprint database includes samples of human speech, plant, and animal voiceprint signals (frequency 0.1Hz-1MHz). Through downsampling and wavelet transforms, a transient feature extraction algorithm based on continuous wavelet transforms was developed. At the hardware level, cross-species voiceprint recognition is achieved through parallel computing on a 24-channel ORRAM optical computing array and linear regression weight optimization, maintaining a recognition accuracy of over 94%, significantly outperforming various traditional voiceprint recognition models. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a principle block diagram of the present invention; Figure 2 This is a joint characterization diagram of the ITO / MoS2 / Ag heterojunction memristor prepared by the present invention; Figure 3 1 is a diagram showing the electrical test results of the ITO / MoS2 / Ag heterojunction memristor prepared by the present invention; Figure 4 This is a diagram showing the optical test results of the ITO / MoS2 / Ag heterojunction memristor prepared by the present invention; Figure 5 This is a diagram showing the key performance test results of the ITO / MoS2 / Ag heterojunction memristor prepared by the present invention; Figure 6 It is a system framework diagram of the present invention; Figure 7 2 is a graph showing the classification accuracy test results of the system of the present invention. DETAILED DESCRIPTION
[0021] The specific implementation manner and working principle of the present invention will be further described in detail below with reference to the accompanying drawings. Example
[0022] This embodiment is divided into four core stages: device preparation - performance testing - system integration - system verification. First, a MoS2 thin film is deposited on an ITO substrate by a magnetron sputtering process, and then a DC power supply is used to drive the sputtering target. - 4 Under a high vacuum environment of 1000 Pa, silver (Ag) metal was deposited, and a multi-electrode array structure was formed using a precision mask to ensure high device uniformity. The memristor was then subjected to comprehensive photoelectric performance testing, including resistive cycling (±0.5V sweep), synaptic plasticity tests (SVDP, SNDP, SWDP, etc.), and short-term memory (STM) stability assessments, verifying its 1458 separable photoconductive states (>10-bit accuracy and microsecond reset capability). During the system construction phase, an ORRAM-based reservoir computing hardware architecture was designed, integrating a 24-channel DAC / ADC module, an STM32 master control unit, and a transimpedance amplifier circuit. This architecture enabled coordinated control of optical pulse encoding (405nm laser triggering) and electrical reset (-1.5V pulses). Simultaneously, a multiplexing interface was developed to accommodate the parallel input of cross-species voiceprint signals (0.1Hz-1MHz). At the algorithm level, the voiceprint signal is mapped to a high-dimensional feature space by combining time-frequency mask preprocessing with continuous wavelet transform (CWT). The dynamic nonlinear response of the memristor array is used to generate a reservoir state matrix. Finally, linear regression (Wout weight optimization) is used to complete the three-dimensional spectrum classification of human, plant, and animal voiceprint signals. The details are as follows: Based on the above content, we can know that we need to prepare ITO / MoS2 / Ag heterojunction memristor first. The preparation process is as follows: The first step is to clean the ITO array substrate; Cleaning of the substrate is crucial because it removes organic, inorganic and particle contamination from the surface, thereby ensuring the quality of subsequent thin film growth.
[0023] In the second step, a metal material is deposited on the surface of the ITO array substrate using magnetron sputtering to form a bottom electrode; Magnetron sputtering is a widely used deposition technique for thin films of metals, semiconductors, and insulators. The core principle of this technology lies in the synergistic effect of electric and magnetic fields. During magnetron sputtering, electrons travel along a spiral trajectory near the target material, colliding with argon atoms, ionizing the argon atoms and generating new electrons and positive ions. Driven by the electric field, these positive argon ions impact the target material surface at high speed, sputtering target atoms or molecules, which are then deposited on the substrate surface, forming a thin film. Through magnetron sputtering, the laboratory is able to deposit uniform and continuous metal films under precisely controlled conditions. These metal films are key to constructing the bottom electrode of high-performance memristors.
[0024] The third step is to deposit molybdenum disulfide on the bottom electrode to form a MoS2 resistive switching functional layer; The fourth step is to -4 Under the high vacuum environment of Pa, a silver metal material is deposited on the resistive switching functional layer by DC magnetron sputtering to form a top electrode; This embodiment uses magnetron sputtering technology to prepare the top electrode of the memristor. The sputtering target is driven by a DC power supply and the surface temperature is less than 5.0×10 -4 Silver (Ag) metal materials are deposited under a high vacuum environment of Pa to ensure high purity and low defect rate of the electrode.
[0025] The fifth step is to form a multi-electrode array structure on the top electrode using a mask plate to obtain the ITO / MoS2 / Ag heterojunction memristor.
[0026] Different upper electrode patterns are achieved using a mask produced using a sophisticated process to form a multi-electrode array structure.
[0027] After preparing the ITO / MoS2 / Ag heterojunction memristor, the prepared memristor was first electrically characterized using an electrochemical workstation (CHI760E) and a LakeShore low-temperature probe station (TTPX) test system. These devices ensure the accuracy and repeatability of the test data through high sensitivity and high stability measurements. The electrical tests mainly include cyclic voltammetry (IV) and current-time (IT) tests, which together form the basis for evaluating the performance of the memristor. The combined characterization results are shown in Figure 2. Figure 2 As shown in Figure 2, the multilayer architecture of the ITO / MoS2 / Ag memristor is shown by SEM (Scanning Electron Microscopy) cross-sectional analysis. Figure 2 As shown in Figure a, a 43.88 nm Ag conductive substrate, a 32.91 nm MoS2 thin film layer, and a 175.5 nm ITO top electrode are shown in sequence. The precisely designed silver layer ensures optimal electrical contact, while the MoS2 realizes the voltage-driven resistance switching that is critical to the function of the memristor. Figure 2 As shown in b, all XPS data are calibrated based on the C 1s peak at 284.5 eV to fit the Mo3d orbital peak. The core energy levels of Mo 3d3 / 2 and 3d5 / 2 are located at 235.8 eV and 232.9 eV, respectively, with a corresponding spin-orbit splitting energy of 2.9 eV, indicating the presence of Mo 6 ⁺ features; In addition, another set of peaks was observed at 233.8 eV (Mo 3d3 / 2) and 230.4 eV (Mo 3d5 / 2), with a spin-orbit splitting energy of 3.4 eV between the two peaks, which is speculated to be caused by Mo 4 The electron state contribution of ⁺. Figure 2c shows the UV-visible absorption spectrum of the MoS2 switching functional layer of the ITO / MoS2 / Ag memristor. αhv ) 2 and energy ( hv ) fitting analysis (where a 、 h 、 v Represents absorbance s intensity, Planck constant and frequency respectively), the semiconductor band gap of MoS2 film is obtained to be 2.95eV, as shown Figure 2 As shown in d.
[0028] The fabricated ITO / MoS2 / Ag heterojunction memristor was then electrically characterized to reveal the device's inherent physical mechanisms and performance, providing experimental data. This electrical characterization precisely measures key parameters such as resistance, current, and voltage to investigate the memristor's resistive switching characteristics and mechanisms, while also optimizing the device fabrication process and ultimately improving device performance. The fabricated memristor was electrically characterized using an electrochemical workstation (CHI760E) and a LakeShore cryogenic probe station (TTPX) test system. These instruments ensure accurate and reproducible test data through their highly sensitive and stable measurements. Electrical testing primarily includes cyclic voltammetry (IV) and current-time (IT) measurements, which together form the foundation for evaluating memristor performance.
[0029] The prepared ITO / MoS2 / Ag heterojunction memristor was electrically tested at room temperature. Figure 3The memristor was tested for its durability, and the retention times of both HRS and LRS were characterized by current-time (it) curves at a read voltage of 0.1 V. Both HRS and LRS were well maintained for 10E-6 s, indicating that the programmable information stored in the Ag / MoS2 / ITO memristor exhibited good non-volatility. RS memory behavior in darkness and light. Our previous work has shown that RS memory behavior is modulated by voltage sweep rate and amplitude. The voltage sweep rate dependence of RS memory behavior is attributed to the interaction between electrons and ions at the interface. Voltage sweep rates of ±0.4, ±0.5, ±0.6, ±0.7, ±0.8, and ±0.9 were applied to the memristor, and the RS memory effect was expected to increase with increasing amplitude. The Ag / MoS2 / ITO memristor was operated at various bias voltage sweep rates from 0.1 V / s to 1.0 V / s at a voltage sweep rate of ±0.8 V. RS memory behavior was detectable at all voltage sweep rates. Voltage amplitudes of 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9V were achieved on the prepared memristor. The bipolar RS memory effect is expected to increase with increasing amplitude. Based on basic electrical test results, it is clear that this device is an analog memristor, requiring testing of its usable computational accuracy. By applying different stimulus voltages and reading the memristor, 64 current states were isolated from the numerous test results, demonstrating that the device maintains at least 6 bits of computational accuracy. At a read voltage of 0.2V, the duration of the stimulus pulse was varied, demonstrating that the response current increased with increasing stimulus duration.
[0030] The prepared ITO / MoS2 / Ag heterojunction memristor was optically tested at room temperature. Figure 4As shown. Integrated optoelectronic devices for sensing, storage, and computing are closely related to novel neuromorphic sensors. A novel neuromorphic visual sensor, integrating sensing, memory, and computing, can mimic the functions of the human retina while also possessing the ability to sense light signals, store signals, and perform information preprocessing. A memristor was tuned using light pulses of equal intensity but varying numbers. The relationship between the time interval between two stimulation pulses and the PPF index was studied. The results showed that the closer the two pulses, the larger the PPF index, indicating a greater facilitation of transmitter release from the presynaptic membrane. The learning and memory time of the memristor increased with the number of pulses, as did the current response. Using the same number of light pulses of varying intensities to tune the memristor, the learning and memory time increased with increasing pulse intensity, as did the current response. Using light pulses of the same intensity but varying frequencies to tune the memristor, higher frequencies resulted in more stimulation within a fixed timeframe, longer total stimulation duration, and a larger current response. Positive photoconductivity also exhibited non-volatility when modulated at a power of 50-80 mW. Taking advantage of the properties of positive and negative photoconductivity, memristors can be used to simulate synapses. Under a 60mW light pulse, short-term plasticity (STP) is observed: the current changes when the memristor is stimulated and then returns to its original level after stimulation. Furthermore, increasing the pulse intensity allows for the simulation of long-term plasticity (LTP), with light pulses starting at 50mW and increasing in steps of 5mW. The current can partially retain its original state after stimulation.
[0031] The prepared ITO / MoS2 / Ag heterojunction memristor was tested for key performance at room temperature. Figure 5 As shown. Among them, Figure 5 (a) shows the high-precision encoding capability (>10 bits) for real-time complex analog signal processing; Figure 5 (b) shows the condition at 0.35 μW / μm 2 pulse frequency-dependent plasticity (SFDP) at laser intensities of 100 nm; Figure 5 (c) shows the photoconductive update results. By using six light pulses (0.35 μW / μm 2 , 0.5ms) to increase the conductance from low to high, while the conductance is directly reset back to the initial value by using 1 negative voltage pulse (-1.5V, 0.5ms); Figure 5(d) shows the frequency dependence of the time-accumulated photocurrent at different light intensities. Key performance tests of the ITO / MoS2 / Ag heterojunction memristor reveal its core advantages in voiceprint recognition systems. By manipulating the light pulse parameters (intensity, number, and frequency), the device exhibits remarkable synaptic plasticity. The device achieves >10-bit precision in photoconductive state encoding at a light intensity of 0.35 μW / μm². After encoding with six light pulses, a single negative voltage pulse (-1.5 V) enables efficient updating and resetting of the conductance. Frequency dependence tests of the time-accumulated photocurrent (0.1 Hz to 1 MHz) show that the device maintains a linear response over a wide frequency band, with a photoconductive gain dynamic range of three orders of magnitude. These properties enable accurate analysis of the multidimensional spectral characteristics of human speech, plant, and animal voiceprints, providing a foundation for real-time voiceprint recognition.
[0032] In the implementation of this invention, the establishment of the algorithm architecture of the memristor reservoir (RC) is a key step, which involves mapping complex voiceprint signals into high-dimensional space and utilizing the dynamic characteristics of the memristor for effective information processing. Figure 6 Specifically, the voiceprint signal is first preprocessed by time-frequency mask to extract the spatiotemporal features f1(t)~f n (t), these features are then converted into a series of voltage values, which are used to control the laser to trigger different light doses, thereby driving the memristor nodes to perform reservoir calculations and ultimately output the results.
[0033] Therefore, this embodiment proposes a multi-dimensional spectrum real-time voiceprint recognition system based on bionic physical reservoir optical calculation. For details, please see the attached Figure 1 and attached Figure 6 The system includes a host computer (PC), a microprocessor (MCU), and multiple synchronous sampling modules. The host computer and the microcontroller are connected via a USART serial port. The synchronous sampling modules are connected to the USART serial port via a control bus. Each synchronous sampling module includes a DAC laser trigger unit, a laser, an ORRAM optical computing array, a signal amplification unit, and an ADC sampling unit connected in sequence. The DAC laser trigger unit and the ADC sampling unit are all connected to the USART serial port via a control bus. The host computer is used to pre-process the voiceprint signal to be identified obtained from the multi-organism voiceprint signal database to generate a feature data array; in this example, the data stored in the multi-organism voiceprint signal database are three-dimensional spectrum features of human language, plant voiceprint signals, and animal voiceprint signals; The microprocessor is used to generate a control signal according to the characteristic data array; The DAC laser trigger unit is used to output a timing voltage pulse whose pulse amplitude linearly corresponds to the characteristic intensity of the voiceprint signal to be identified according to the control signal; The laser is used to generate a sequence of optical pulses based on time-sequential voltage pulses; The ORRAM optical computing array is composed of an ITO / MoS2 / Ag heterojunction memristor, which is used to convert the optical pulse sequence into a photocurrent response signal through the synergistic effect of the MoS2 interlayer exciton effect and the surface plasmon resonance of Ag nanoparticles; The signal amplification unit is used to amplify the photocurrent response signal output by the ORRAM optical computing array and convert it into a voltage signal; The ADC sampling unit is used to sample the voltage signal output by the signal amplifying unit in real time and convert it into a digital model; The host computer is further configured to complete voiceprint recognition based on the sampling data of the multiple synchronous sampling modules.
[0034] Furthermore, the synchronous sampling module also includes an inverting proportional amplifier unit and a multiplexer. The inverting proportional amplifier unit is used to convert the timing voltage pulse of the positive voltage into an equal-amplitude negative pulse. The multiplexer is used to dynamically switch the output channel to output positive pulses or negative pulses, thereby realizing the regulation of the frequency and light intensity of the laser.
[0035] Based on the attached Figure 1 As can be seen from the above, this system integrates a host computer, a microcontroller, and a custom PCB. The test script is written in Python. The host computer sends test commands to the microcontroller via serial communication. The microcontroller controls the PCB's DAC output voltage, ADC voltage acquisition, and multiplexer channel switching.
[0036] In its implementation, this system forms a 24-channel synchronous sampling module. Correspondingly, the 24 DAC laser trigger units consist of three 8-channel, 12-bit DAC chips, each with independent output voltage control. A multiplexer connects the DAC laser trigger unit to its inverting proportional amplifier circuit module, allowing selection of positive and negative output voltages via analog switches. The ADC sampling system combines an external 16-channel, 16-bit ADC with the ADC channels integrated in the STM32, forming a 24-channel ADC sampling architecture to ensure synchronous voltage acquisition. Furthermore, the PCB board integrates 24 weak current detection channels, also known as signal amplification units. Each channel uses a transimpedance amplifier to convert current in the nanoamp range to a voltage. All multiplexers are analog multiplexers, with digital signals controlling the switching of different ports. The PCB also includes a multi-channel regulated power supply module, providing sufficient low-ripple positive and negative power for the entire system, enhancing system stability.
[0037] During the implementation of this embodiment, the specific steps of the host computer preprocessing the voiceprint signal to be identified are as follows: The first step is to downsample and extract preliminary features of the voiceprint signal to obtain the temporal and spatial characteristics of the signal; The second step is to use the continuous wavelet transform (CWT) method to obtain the frequency characteristics of the voiceprint signal to be identified; The third step is to construct the mask length and the number of matrices, fill each matrix with random simulation values, and multiply the spatiotemporal features with the frequency features to map to a higher dimensional space to obtain the feature data array F(t) = {f1(t), f2(t) ..., f n (t)}.
[0038] By designing the joint time-frequency masking algorithm described in this embodiment, the system can adaptively extract the salient features of different biometric voiceprints (such as the formant continuity of human speech and the pulse intermittency of plant sounds). The row vectors of these newly formed feature matrices are linearly scaled to simulate the intensity of the pulse train, and these row vectors are sequentially and parallelly input into the memristor nodes in the RC hardware system for efficient encoding and processing of the voiceprint signals.
[0039] In the process of efficiently encoding and processing the voiceprint signal by the memristor node, the optical pulse input and device state value of the memristor node are used as input, and the state value at the next moment is used as output. In this way, the memristor device can simulate the internal state update dynamics and realize the efficient encoding and processing of the voiceprint signal. That is, the steps of the ORRAM optical computing array to obtain the photocurrent response signal are as follows: In the first step, based on the memristor reservoir algorithm, after each light pulse cycle, the reservoir state x(t) is captured and all x(t) values are integrated into the reservoir state matrix X, thereby encapsulating a series of single-sample reservoir states. The second step is to obtain the weight matrix Wout obtained after linear regression training. The specific method of obtaining it is: use the reservoir state matrix X and the target output label Y to train linear regression to determine the connection weight matrix Wout. The formula for the linear regression training output weight matrix Wout is: Wout=(X T X) −1 X T Y, weight matrix Wout matrix connecting the reserve layer and the output layer; In the third step, for a fully trained network, the expected output Y can be obtained by applying X×Wout to the sample processing, that is, the output photocurrent response signal Y can be calculated according to the formula Y=X×Wout.
[0040] By building this algorithmic architecture, the memristor reservoir system can process complex voiceprint signals and achieve high-precision recognition of animal, plant, and human voiceprints. This reservoir calculation method, based on a physical dynamic system, not only enables parallel multiplication and accumulation calculations directly in the memristor crossbar array using Ohm's law and Kirchhoff's law, but also achieves higher data processing efficiency compared to traditional CMOS platform processors such as CPUs, GPUs, and FPGAs.
[0041] Furthermore, to eliminate the cumulative photoconductance effect at nodes in the ORRAM optical computing array, this system employs an optical-electrical synergistic reset strategy. After continuous optical pulse encoding, the microprocessor controls the DAC laser trigger unit to output an electrical reset pulse, which instantly restores the ORRAM optical computing array to its initial state through rapid ion migration at the Ag electrode. Compared to traditional natural decay reset (taking >10ms), this strategy compresses the system refresh cycle to microseconds.
[0042] Preferably, the microprocessor adopts a minimum system based on an STM32 single-chip microcomputer, and the signal amplification unit adopts a transimpedance amplifier.
[0043] In summary, this system builds a high-precision, low-latency voiceprint signal processing hardware platform based on ITO / MoS2 / Ag memristor array and optical-electrical synergistic control technology. Its core modules include DAC laser trigger unit, laser, ORRAM optical computing array, signal amplification unit and ADC sampling unit. The specific workflow is as follows: Step 1: DAC laser triggering and optical pulse encoding, the voiceprint signal to be identified is mapped into high-dimensional spatiotemporal features through front-end preprocessing (time-frequency masking, continuous wavelet transform), and the host computer (PC) transmits the feature data array to the microcontroller (STM32MCU) through the USART serial port; Step 2: the microprocessor generates a control signal according to the characteristic data array; Step 3: The DAC laser trigger unit (TLV5610 chip, 12-bit resolution, 1MHz refresh rate) outputs a timing voltage pulse whose pulse amplitude (0-5V) linearly corresponds to the characteristic intensity of the voiceprint signal to be identified according to the control signal; At the same time, because the ORRAM optical computing array requires the coordinated operation of positive and negative voltages (positive pulses for optical triggering and negative pulses for reset), the timing voltage pulses output by the DAC laser trigger unit are converted into equal-amplitude negative pulses (-5V to +5V) through an inverting proportional amplifier circuit (based on an operational amplifier chip). The output channels are then dynamically switched by a multiplexer (MUX) to achieve precise control of the frequency (1Hz to 1MHz) and light intensity (0.1 to 1.89μW / μm²) of the laser (405nm wavelength). For example, when processing bat ultrasonic signals (150kHz), the DAC drives the laser with 1.5V pulses to generate a 100μs-wide light pulse sequence, ensuring that the ORRAM node undergoes stable photoconductive state transitions under photon excitation. Step 4: the laser generates a light pulse sequence according to the timing voltage pulse; Step 5: The ORRAM optical computing array achieves 1458 separable photoconductive states (equivalent to 10-bit precision) in a single device through the synergistic effect of the MoS2 interlayer exciton effect and the surface plasmon resonance of Ag nanoparticles, converting the optical pulse sequence into a photocurrent response signal. Step 6: The signal amplification unit amplifies the photocurrent response signal output by the ORRAM optical computing array and converts it into a voltage signal. The nonlinear dynamic mapping capability of the photocurrent response signal supports the extraction of voiceprint features. At the same time, the optical pulse interval matches the short-term memory (STM) characteristics of the ORRAM, ensuring the complete preservation of the signal timing characteristics. Step 7: The ADC sampling unit samples the voltage signal output by the signal amplification unit in real time and converts it into a digital model; Step 8. The ADC and the microprocessor's built-in ADC channel work together to form a 24-channel synchronous sampling system to ensure the parallel acquisition of multimodal voiceprint signals (such as human, animal, and plant voiceprint signals). The sampling data of multiple synchronous sampling modules are transmitted to the host computer in real time through the serial port. The host computer completes voiceprint recognition based on the sampling data of multiple synchronous sampling modules combined with the linear regression algorithm (Wout weight matrix optimization).
[0044] After testing, the system, driven by 1MHz optical pulses, meets the real-time processing requirements of complex voiceprints. Based on the high-dimensional feature mapping capability of optical computing in the reserve pool, the system recognizes and classifies the voice signals of 10 testers in a noisy environment. The results are as follows Figure 7As shown. It can be seen that high-precision 1:N verification is achieved, and the recognition accuracy is 100%, which is significantly better than the traditional MFCC method. In particular, it maintains stable performance in complex noise scenes, verifying the spectral selectivity advantage of optoelectronic devices. Optical pulse reservoir encoding is performed on the sounds of tomato plants recorded in the greenhouse and the sounds of tobacco plants recorded in the greenhouse. The system's classification accuracy of the two plant voiceprints is significantly improved compared to the existing LPC scheme. In complex scenarios that integrate human voiceprints, tobacco plant voiceprints, and cat voiceprints, the system uses a joint time-frequency masking algorithm and a parallel optical computing architecture to achieve real-time analysis of multi-dimensional spectral features, with an identification accuracy of over 94%, which is more than 7% lower than the error rate of the traditional RNN model, fully demonstrating its core advantages in wide-band, multi-source voiceprint processing.
[0045] As can be seen from the above, the present invention breaks through the multi-dimensional constraints of existing technologies through material innovation, architecture optimization and algorithm collaboration. At the device level, the ITO / MoS2 / Ag heterojunction memristor array is used to utilize the interlayer exciton effect of MoS2 and the surface plasmon resonance of Ag nanoparticles to achieve a wide spectrum light response (405nm-850nm) and ultra-high photoelectric conversion efficiency. Combined with the transparent conductive properties of the ITO electrode, the single device can achieve a power output of 0.66μW / μm 2 Under high-intensity light, 1458 separable photoconductive states (equivalent to 10-bit precision) can be generated, quadrupling the number of virtual nodes in traditional oxide memristors. At the system architecture level, an innovative opto-electrical synergy strategy is introduced. Simultaneously, the rapid ion migration characteristics of the Ag electrode are utilized to instantaneously reset the device state with a negative voltage pulse, compressing the system refresh cycle to the microsecond level, thereby supporting real-time stream processing of 1MHz high-frequency voiceprint signals. Furthermore, by constructing a cross-species voiceprint database (covering the three-dimensional spectral characteristics of human speech, plant voiceprint signals, and animal voiceprint signals) and designing a joint time-frequency masking algorithm, the system can adaptively extract significant features of different bio-voiceprints (such as formant continuity in human speech and pulse intermittency in plant sounds). Ultimately, linear regression at the output layer of the reservoir achieves a cross-species recognition accuracy of over 94%, providing a new hardware solution for fields such as smart ecology and biomedical monitoring.
[0046] The technical solution provided by the present invention is introduced in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A multi-dimensional spectrum real-time voiceprint recognition system based on bionic physical reservoir optical calculation, characterized by: It includes a host computer, a microprocessor, and multiple synchronous sampling modules. Each synchronous sampling module includes a DAC laser trigger unit, a laser, an ORRAM optical computing array, a signal amplification unit, and an ADC sampling unit, among which: The host computer is used to pre-process the voiceprint signal to be identified obtained from the multi-biological voiceprint signal database to generate a feature data array; The microprocessor is used to generate a control signal according to the characteristic data array; The DAC laser trigger unit is used to output a timing voltage pulse whose pulse amplitude linearly corresponds to the characteristic intensity of the voiceprint signal to be identified according to the control signal; The laser is used to generate a sequence of optical pulses based on time-sequential voltage pulses; The ORRAM optical computing array is composed of an ITO / MoS2 / Ag heterojunction memristor, which is used to convert the optical pulse sequence into a photocurrent response signal through the synergistic effect of the MoS2 interlayer exciton effect and the surface plasmon resonance of Ag nanoparticles; The signal amplification unit is used to amplify the photocurrent response signal output by the ORRAM optical computing array and convert it into a voltage signal; The ADC sampling unit is used to sample the voltage signal output by the signal amplifying unit in real time and convert it into a digital model; The host computer is further configured to complete voiceprint recognition based on the sampling data of the multiple synchronous sampling modules.
2. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to claim 1 is characterized by: The data stored in the multi-biological voiceprint signal database are three-dimensional spectrum features of human language, plant voiceprint signals, and animal voiceprint signals.
3. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to claim 1 is characterized by: The specific steps of the host computer preprocessing the voiceprint signal to be identified are as follows: Downsample and perform preliminary feature extraction on the voiceprint signal to obtain the spatiotemporal characteristics of the signal; The continuous wavelet transform method is used to obtain the frequency characteristics of the voiceprint signal to be identified; The mask length and the number of matrices are constructed, each matrix is filled with random simulation values, and the spatiotemporal features are multiplied with the frequency features to obtain the feature data array.
4. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to claim 1 is characterized by: The preparation method of the ITO / MoS2 / Ag heterojunction memristor is as follows: ITO array substrate cleaning; Using magnetron sputtering, metal material is deposited on the surface of the ITO array substrate to form a bottom electrode; Molybdenum disulfide is deposited on the bottom electrode to form a MoS2 resistive switching functional layer; In less than 5.0×10 -4 Under the high vacuum environment of Pa, a silver metal material is deposited on the resistive switching functional layer by DC magnetron sputtering to form a top electrode; A multi-electrode array structure is formed on the top electrode using a mask plate to obtain the ITO / MoS2 / Ag heterojunction memristor.
5. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to claim 1 is characterized by: The steps of obtaining the photocurrent response signal of the ORRAM optical computing array are as follows: Based on the memristor reservoir algorithm, after each light pulse cycle, the reservoir state x(t) is captured and all x(t) values are integrated into the reservoir state matrix X; Get the weight matrix Wout obtained after linear regression training; The output photocurrent response signal Y is calculated according to the formula Y=X×Wout.
6. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to claim 1 is characterized by: The microprocessor is further configured to control the DAC laser trigger unit to output an electrical reset pulse to instantaneously restore the ORRAM optical computing array to an initial state.
7. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to claim 1 is characterized by: The synchronous sampling module also includes an inverting proportional amplifier unit and a multiplexer. The inverting proportional amplifier unit is used to convert the timing voltage pulse of the positive voltage into a negative pulse of equal amplitude. The multiplexer is used to dynamically switch the output channel to output positive pulses or negative pulses to achieve the regulation of the frequency and light intensity of the laser.
8. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to claim 1 is characterized by: The host computer and the microcontroller are connected via a USART serial port, and the DAC laser trigger unit and the ADC sampling unit are both connected via a control bus and the USART serial port.
9. The multi-dimensional spectrum real-time voiceprint recognition system based on biomimetic physical reservoir optical calculation according to any one of claims 1 to 8, characterized in that: The microprocessor adopts a minimum system based on an STM32 single-chip microcomputer, and the signal amplification unit adopts a transimpedance amplifier.