A noise-resistant speech recognition method and system based on humidity-sensitive memory resistance, a terminal and a medium
By combining humidity-sensitive memristor LIF neural circuits and audio data, multimodal information fusion is achieved, solving the problem of low accuracy of speech recognition systems in noisy environments and improving recognition performance.
Patent Information
- Application Number
- CN202511509770.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Existing speech recognition systems show a significant drop in accuracy in noisy environments, lack the ability to utilize environmental information (such as humidity changes) that accompanies speech generation, and are unable to achieve multimodal information fusion, thus limiting their applicability in real-world environments.
A humidity-sensitive memristor-based LIF neural network is used to collect humidity information and fuse it with audio data. The speech is then recognized through a neural network, using humidity information as a complementary feature to improve recognition accuracy.
It significantly improves the accuracy and robustness of speech recognition in noisy environments, breaking through the limitations of single speech signal recognition.
Smart Images

Figure CN120977299B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electronic science and technology, and in particular to a noise-resistant speech recognition method and system based on humidity-sensitive memristor, a terminal and a medium. BACKGROUND
[0002] Multimodal perception fusion plays a core role in human perception and decision-making, especially in complex or noisy environments. For example, when crossing a busy street, accurate judgments are made by relying on the fusion of visual information and auditory information. However, in acoustically chaotic environments, relying solely on auditory signals is susceptible to interference, leading to a decline in speech recognition performance.
[0003] Existing speech recognition systems are mostly based on single-modal audio signals, such as Mel-Frequency Cepstral Coefficients (MFCC) extraction followed by neural network processing. Such systems perform well in quiet environments but significantly decrease in recognition accuracy in noisy environments. Traditional methods lack the use of environmental information (such as humidity changes in exhaled airflow during speech) that accompanies the speech production process, cannot achieve multimodal information fusion, and limit their applicability in real-world environments.
[0004] Therefore, the prior art still has defects. SUMMARY
[0005] The technical problem to be solved by the present application is to provide a noise-resistant speech recognition method and system based on humidity-sensitive memristor, a terminal and a medium, which addresses the above-mentioned defects of the prior art. The technical solution adopted by the present application is as follows:
[0006] In a first aspect, the present application provides a noise-resistant speech recognition method based on humidity-sensitive memristor, wherein the method comprises:
[0007] Collecting speech signals based on a microphone device to obtain audio data, and collecting humidity information based on a LIF neuron circuit, and outputting voltage pulse data based on humidity information by the LIF neuron circuit, wherein the LIF neuron circuit includes a humidity-sensitive memristor;
[0008] Applying random noise to the audio data, and fusing and encoding the audio data containing random noise with the voltage pulse data to obtain a fusion data vector group;
[0009] Inputting the fusion data vector group into a pre-trained neural network to output a speech recognition result, wherein the neural network is a network model trained based on the corresponding relationship between feature vector samples and speech recognition samples, and the feature vector samples are vector samples obtained by previously fusing and encoding audio samples and voltage pulse samples.
[0010] In an implementation, the LIF neuron circuit is composed of the humidity sensitive memristor, a current-limiting resistor, a load resistor, and a parallel capacitor. When the voice signal is accompanied by humid airflow, the resistance of the humidity sensitive memristor decreases with the increase of humidity, causing the charging process of the parallel capacitor to accelerate, so that the humidity sensitive memristor reaches the threshold faster and triggers a pulse discharge.
[0011] In an implementation, the preparation process of the humidity sensitive memristor includes:
[0012] 1 milliliter of Nafion solution is measured, mixed with 3.6 milliliters of anhydrous ethanol, and shaken for 5 minutes to uniformly mix the Nafion solution with the anhydrous ethanol to obtain a Nafion diluent;
[0013] The glass substrate with a 185-nanometer-thick ITO electrode is cleaned with deionized water by ultrasonic cleaning, then dried with nitrogen, and dried at 120°C for 30 minutes;
[0014] 200 microliters of the Nafion diluent are taken and spin-coated on the cleaned glass substrate, then annealed at 100°C for 45 minutes to obtain a 30-nanometer-thick Nafion polymer medium layer;
[0015] The electrode is prepared using a metal mask plate with a pre-designed pattern to obtain a humidity sensitive memristor.
[0016] In an implementation, the spin-coating speed conditions are 500 rpm for 5 seconds and then 1500 rpm for 40 seconds.
[0017] In an implementation, the electrode is prepared using a metal mask plate with a pre-designed pattern to obtain a humidity sensitive memristor, including:
[0018] The metal mask plate is covered on the spin-coated glass substrate, 50 nanometers of silver are evaporated on the glass substrate by thermal evaporation, and then the metal mask plate is removed to obtain a humidity sensitive memristor.
[0019] In an implementation, the audio data containing random noise is fused and encoded with the voltage pulse data to obtain a set of fusion data vectors, including:
[0020] The voltage pulse data is widened, and the widened voltage pulse data is multiplied by a scaling factor to obtain the magnitude-matched voltage pulse data;
[0021] The audio data containing random noise is point-by-point added to the magnitude-matched voltage pulse data to obtain voice humidity fusion data;
[0022] perform MFCC encoding on the voice-humidity fusion data to obtain a fusion data vector group.
[0023] In an implementation manner, the method further includes:
[0024] perform MFCC encoding on the audio data containing random noise to obtain an audio data vector group;
[0025] input the audio data vector group and the fusion data vector group into the trained neural network respectively to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group, wherein the accuracy of the second recognition result is higher than that of the first recognition result.
[0026] In a second aspect, the embodiments of the present application further provide an anti-noise voice recognition system based on humidity-sensitive memristor, wherein the system is configured to implement the steps of the anti-noise voice recognition based on humidity-sensitive memristor in the above-mentioned solution, and the system includes:
[0027] a multi-modal information acquisition module configured to acquire voice signals based on a microphone device to obtain audio data, and acquire humidity information based on a LIF neuron circuit, and output voltage pulse data based on the humidity information by the LIF neuron circuit, wherein the LIF neuron circuit includes a humidity-sensitive memristor;
[0028] a fusion and encoding module configured to apply random noise to the audio data, and fuse and encode the audio data containing random noise and the voltage pulse data to obtain a fusion data vector group;
[0029] a voice recognition module configured to input the fusion data vector group into a pre-trained neural network to output a voice recognition result, wherein the neural network is a network model trained based on a corresponding relationship between a feature vector sample and a voice recognition sample, and the feature vector sample is a vector sample obtained by pre-fusing and encoding an audio sample and a voltage pulse sample.
[0030] In a third aspect, the embodiments of the present application further provide a terminal, wherein the terminal includes a memory, a processor, and a humidity-sensitive memristor-based anti-noise voice recognition program stored in the memory and executable on the processor, and the processor implements the steps of the humidity-sensitive memristor-based anti-noise voice recognition method of any one of the above-mentioned solutions when executing the humidity-sensitive memristor-based anti-noise voice recognition program.
[0031] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, wherein the computer readable storage medium stores a humidity-sensitive memristor-based anti-noise speech recognition program, and the humidity-sensitive memristor-based anti-noise speech recognition program implements the steps of the humidity-sensitive memristor-based anti-noise speech recognition method according to any one of the above solutions on the computer readable storage medium.
[0032] Advantageously, compared with the prior art, the present application provides a humidity-sensitive memristor-based anti-noise speech recognition method. The present application first collects a voice signal based on a microphone device to obtain audio data, and collects humidity information based on a LIF neuron circuit, and outputs voltage pulse data based on the humidity information by the LIF neuron circuit, wherein the LIF neuron circuit includes a humidity-sensitive memristor. Then, random noise is applied to the audio data, and the audio data containing the random noise is fused and encoded with the voltage pulse data to obtain a fusion data vector group. Then, the fusion data vector group is input into a pre-trained neural network to output a speech recognition result, wherein the neural network is a network model trained based on a corresponding relationship between a feature vector sample and a speech recognition sample, and the feature vector sample is a vector sample obtained by fusing and encoding an audio sample and a voltage pulse sample in advance.
[0033] The present application utilizes the high sensitivity of the LIF neuron circuit based on the humidity-sensitive memristor to humidity fluctuations to convert humidity information into corresponding voltage pulse data, thereby providing complementary feature information for speech recognition, breaking through the recognition limit of a single voice signal, and significantly improving the speech recognition accuracy and robustness in a noisy environment by utilizing multi-modal information fusion. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 The flowchart of a preferred embodiment of the humidity-sensitive memristor-based anti-noise speech recognition method provided by the embodiments of the present application.
[0035] Figure 2 The working principle diagram of the humidity-sensitive memristor-based anti-noise speech recognition method provided by the embodiments of the present application.
[0036] Figure 3 The LIF neuron circuit diagram in the embodiments of the present application.
[0037] Figure 4 The spike pulse diagram output by the embodiments of the present application under different humidity.
[0038] Figure 5 The structure diagram of the humidity-sensitive memristor in the humidity-sensitive memristor-based anti-noise speech recognition method provided by the embodiments of the present application.
[0039] Figure 6 The threshold voltage and current change of the humidity sensitive memristor in the embodiment of the present application under different humidity.
[0040] Figure 7 The schematic diagram of the influence of humidity and voltage on the opening speed of the humidity sensitive memristor.
[0041] Figure 8 The schematic diagram of the humidity field programming experimental results of the humidity sensitive memristor.
[0042] Figure 9 The schematic diagram of the pulse neural network provided by the embodiment of the present application.
[0043] Figure 10 The schematic diagram of the comparison of the speech recognition accuracy under different noise levels.
[0044] Figure 11 The principle block diagram of the anti-noise speech recognition system based on the humidity sensitive memristor provided by the embodiment of the present application.
[0045] Figure 12 The principle block diagram of the terminal provided by the embodiment of the present application. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical scheme and effect of the present application more clear and explicit, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0047] The flowchart shown in the drawings is only an example and does not necessarily include all the contents and operations or steps, nor does it necessarily execute in the order described. For example, some operations or steps can be further divided, combined or partially combined, so the actual execution order may be changed according to the actual situation.
[0048] It should be understood that the terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless the context clearly indicates otherwise, the singular form "a", "an" and "the" is intended to include the plural form.
[0049] It should be understood that in order to facilitate the clear description of the technical scheme of the embodiments of the present application, in the embodiments of the present application, the same items or similar items with basically the same function and role are distinguished by using "first", "second" and the like. For example, the first control information and the second control information are only used to distinguish different control information and do not limit the order.
[0050] Those skilled in the art can understand that the terms "first", "second" and the like do not limit the number and execution order, and the terms "first", "second" and the like do not necessarily mean different.
[0051] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application means one or more of the associated listed items and all possible combinations thereof, and includes these combinations.
[0052] To solve the problems of the prior art, the embodiment of the present application provides an anti-noise speech recognition method based on humidity-sensitive memristor. Based on the method of the embodiment, the high sensitivity of the LIF (Leaky Integrate-and-Fire) neuron circuit based on the humidity-sensitive memristor to humidity fluctuations can convert humidity information into corresponding voltage pulse data, thereby providing complementary feature information for speech recognition, breaking through the recognition limit of a single speech signal, and improving the speech recognition accuracy and robustness in a noisy environment. In specific applications, the embodiment first collects a speech signal based on a microphone device to obtain audio data, and collects humidity information based on a LIF neuron circuit, and outputs voltage pulse data based on the humidity information by the LIF neuron circuit, wherein the LIF neuron circuit includes a humidity-sensitive memristor. Then, random noise is applied to the audio data, and the audio data containing random noise is fused and encoded with the voltage pulse data to obtain a fusion data vector group. Next, the fusion data vector group is input into a pre-trained neural network to output a speech recognition result, wherein the neural network is a network model trained based on the corresponding relationship between a feature vector sample and a speech recognition sample, and the feature vector sample is a vector sample obtained by previously fusing and encoding an audio sample and a voltage pulse sample. In the embodiment, a LIF neuron circuit of a humidity-sensitive memristor is constructed, and the humidity information of the airflow accompanying the speech signal can be detected based on the LIF neuron circuit to output voltage pulse data, and then the audio data and the voltage pulse data are comprehensively analyzed to obtain a speech recognition result. Compared with the technical solution of the prior art in which only the audio data is analyzed to realize speech recognition, the speech recognition result of the embodiment is more accurate.
[0053] The anti-noise speech recognition method based on the humidity-sensitive memristor of the embodiment can be applied to a terminal for processing and analyzing the collected audio data and voltage pulse data to realize speech recognition. The terminal can be a smart product terminal such as a mobile phone or a computer. Figure 1 The anti-noise speech recognition method based on the humidity-sensitive memristor of the embodiment includes the following steps:
[0054] In step S100, a voice signal is collected based on a microphone device to obtain audio data, and humidity information is collected based on a LIF neuron circuit, and voltage pulse data is output by the LIF neuron circuit based on the humidity information, wherein the LIF neuron circuit comprises a humidity-sensitive memristor.
[0055] Traditional speech recognition technology mainly relies on processing methods based on audio data, such as extracting Mel-frequency cepstral coefficients (MFCC) and combining neural networks for recognition and classification. However, in a noisy environment, background noise seriously interferes with the quality of the audio signal, resulting in a significant decrease in recognition rate. Therefore, the present embodiment can introduce humidity information naturally accompanied by the speech signal as an auxiliary perception modality, effectively overcoming the limitations brought by pure audio data analysis. Specifically, when a person speaks, they exhale humid airflow, and the humidity change is highly correlated with the speech content in terms of timing and intensity. The LIF neuron circuit is sensitive to humidity fluctuations, so the humidity can be converted into corresponding voltage pulse data, thereby providing complementary feature information for speech recognition.
[0056] To this end, the present embodiment constructs a multi-modal signal acquisition system, which combines Figure 2 As shown in the multi-modal signal acquisition system, the multi-modal signal acquisition system comprises a microphone device and a LIF neuron circuit. The microphone device is placed 10 cm away from the mouth to collect the speech signal and obtain the audio data. The LIF neuron circuit is located 5 cm away from the mouth to detect the humidity information of the exhaled humid airflow and output voltage pulse data through the LIF neuron circuit. Specifically, as shown in the LIF neuron circuit of the present embodiment, the LIF neuron circuit is composed of a humidity-sensitive memristor, a current-limiting resistor Figure 3 (200k ), a load resistor (20k ), and a parallel capacitor (0.33 ). When the speech signal is accompanied by humid airflow, the resistance of the humidity-sensitive memristor decreases with the increase of humidity, causing the charging process of the parallel capacitor to accelerate, so that the humidity-sensitive memristor reaches the threshold value faster and triggers pulse discharge. As shown in the LIF neuron circuit, the higher the humidity, the greater the output pulse frequency and amplitude, which can realize real-time encoding of the speech-related humidity fluctuations. The LIF neuron circuit directly outputs voltage pulse data. Figure 4
[0057] Traditional humidity sensing systems usually adopt a serial architecture with separate sensors, memories and computing units, resulting in high energy consumption and delay in signal conversion and transmission. Existing humidity-sensitive nanowire-based memristors face the problem of high integration complexity. The embodiment adopts a polymer material highly sensitive to humidity, prepares a humidity-sensitive memristor, and builds a LIF neuron circuit for humidity signal sensing and processing based on the humidity-sensitive memristor. When the environmental humidity increases, the opening threshold voltage of the humidity-sensitive memristor decreases, the opening time shortens, and the device current significantly increases. By using this humidity-dependent electrical characteristic, the LIF neuron circuit built can generate differentiated pulse discharge behavior according to humidity changes: under high humidity conditions, the output pulse frequency and amplitude increase; under low humidity conditions, they decrease accordingly. The embodiment uses the mapping relationship between humidity and pulse characteristic parameters to realize direct sensing and preliminary processing of humidity signals. The humidity-sensitive memristor has the advantages of simple preparation process, easy large-scale integration, fast response speed, low threshold voltage and low power consumption.
[0058] Specifically, in the preparation of the humidity-sensitive memristor, the following steps are included:
[0059] Step 1, measure 1 milliliter of Nafion solution, mix the Nafion solution with 3.6 milliliters of anhydrous ethanol, and shake for 5 minutes to mix the Nafion solution and anhydrous ethanol uniformly to obtain a Nafion diluent.
[0060] Step 2, ultrasonically clean the glass substrate with a 185-nanometer-thick ITO (Indium Tin Oxide) electrode with deionized water for 10 minutes to remove surface contaminants. Then, dry the surface with high-purity nitrogen and dry at 120°C for 30 minutes. The glass substrate of the embodiment can also be selected from a silicon wafer or a flexible substrate.
[0061] Step 3, take 200 microliters of the Nafion diluent, spin-coat the Nafion diluent on the cleaned glass substrate, and set the spin-coating program to include two stages. The first stage is spin-coated at a speed of 500 revolutions per minute for 5 seconds to ensure solution spreading. The second stage is spin-coated at a speed of 1500 revolutions per minute for 40 seconds to form a uniform film. Then, anneal at 100 degrees Celsius for 45 minutes to remove residual solvents and enhance film density, obtaining a 30-nanometer-thick Nafion polymer dielectric layer.
[0062] Step 4, using a metal mask with a pre-made pattern to prepare the electrode, to obtain a humidity sensitive memristor. Specifically includes: covering the metal mask on the spin-coated glass substrate, by thermal evaporation method, in a vacuum condition on the glass substrate evaporation of 50 nanometer silver, evaporation rate control at 0.1 angstrom / s, after removing the metal mask, to obtain a humidity sensitive memristor. The structure of the humidity sensitive memristor is shown in Figure 5 In other implementations, copper, platinum, nickel, aluminum can also be used as the electrode material of the humidity sensitive memristor.
[0063] The performance of the humidity sensitive memristor is studied in this embodiment, mainly around two aspects: first, the threshold transition characteristics of the humidity sensitive memristor under different humidity conditions are studied, and the influence of humidity on the operating voltage and the opening time of the device is analyzed; The second is to study the resistance state programming method based on the humidity field control. The specific experimental content and results are as follows:
[0064] (1) The semiconductor parameter analyzer is used to test the electrical performance of the device in cooperation with the humidity control system. First, in the environment with relative humidity of 30%, 50%, 70% and 90%, respectively, a direct current voltage of 0 V to 1.5 V is applied to the device for back sweep, and the scanning rate is 0.01 V / s. The resistance change behavior of the humidity sensitive memristor is monitored. The experimental results show that: under lower humidity (30%), the device does not have obvious resistance change switching; while under higher humidity conditions (50%, 70% and 90%), the device shows typical resistance change characteristics. Further analysis shows that the threshold voltage of the device decreases significantly with the increase of the environmental humidity, and the high resistance state current increases with the increase of the humidity, as shown in Figure 6 .
[0065] (2) The semiconductor parameter analyzer is used to apply a pulse voltage signal to the humidity sensitive memristor to evaluate its fastest dynamic response characteristics. In 600 nanoseconds, a 1 microsecond 5 volt pulse signal is applied, and the humidity sensitive memristor generates a current response under the action of the voltage. Taking 70% of the maximum current as the opening point, the fastest response time of the device is calculated, as shown in Figure 7 .
[0066] (3) Humidity field programming experiment: under the condition of fixed bias 0.3 V, the resistance state writing and erasing of the device is realized by adjusting the environmental humidity. When the environmental humidity increases, 0.3 V exceeds the threshold voltage of the humidity sensitive memristor, which promotes the humidity sensitive memristor to switch from the high resistance state to the low resistance state; while the humidity decreases, the voltage is lower than the threshold voltage of the humidity sensitive memristor, and the humidity sensitive memristor automatically returns to the high resistance state. This result shows that reversible and non-volatile resistance state control can be realized by adjusting the environmental humidity only, thereby verifying the feasibility of the humidity field programming, and the experimental results are shown in Figure 8 .
[0067] The embodiment collects humidity information of airflow accompanying the voice signal through the LIF neuron circuit based on the humidity-sensitive memristor. The resistance of the humidity-sensitive memristor decreases with the increase of humidity, which causes the charging process of the parallel capacitor to accelerate, so that the humidity-sensitive memristor reaches the threshold and triggers the pulse discharge faster. The higher the humidity, the greater the output pulse frequency and amplitude, thereby realizing real-time coding of the voice-related humidity fluctuation. The LIF neuron circuit directly outputs voltage pulse data, which is recorded through an oscilloscope, and the voltage pulse data and audio data are used together to realize voice recognition. It can be seen that the LIF neuron circuit of the embodiment can not only perceive humidity information, but also convert the perception result into a spike pulse signal that is easier for the circuit to process.
[0068] Step S200, random noise is applied to the audio data, and the audio data containing random noise is fused and coded with the voltage pulse data to obtain a fusion data vector group.
[0069] After obtaining the above-mentioned audio data and voltage pulse data, the embodiment can apply random noise to the audio data. Then the audio data containing random noise is fused and coded with the voltage pulse data to obtain a fusion data vector group. When performing fusion and coding, the embodiment first widens the voltage pulse data, which can be widened by interpolation method to widen the data dimension of the voltage pulse data, so that it is completely aligned with the audio data on the time axis. For example, if the audio data is 1000 sampling points, the voltage pulse data should be widened to 1000 corresponding points. Then, the widened voltage pulse data is multiplied by a scaling factor to obtain voltage pulse data with matched magnitude. The scaling factor in the embodiment can be set and adjusted according to the amplitude range of the voltage pulse data and the audio data, such as 0.1 or 0.2. By multiplying the widened voltage pulse data by the scaling factor, the widened voltage pulse data and the audio data can be matched in magnitude, avoiding the data magnitude of the widened voltage pulse data being too large or too small, which may be covered by the audio data or cover the audio data during subsequent fusion. Then, the audio data containing random noise is added point by point with the voltage pulse data with matched magnitude to obtain voice humidity fusion data. Then, the voice humidity fusion data is MFCC coded to realize feature extraction, and a fusion data vector group is obtained. The MFCC coding of the embodiment is a common feature extraction method in the field of voice recognition technology, which is mainly used to convert time-domain signals into frequency-domain feature vectors. The embodiment adopts the mode of first fusion and then coding, which not only simplifies the data processing process, but also reduces the coding time and complexity.
[0070] Step S300, inputting the fusion data vector group into a pre-trained neural network to output a voice recognition result.
[0071] After obtaining the fusion data vector group, the embodiment can input the fusion data vector group into the pre-trained neural network to output the speech recognition result. In the embodiment, the neural network is a network model trained based on the correspondence between the feature vector sample and the speech recognition sample. The feature vector sample is a vector sample obtained by pre-fusing and encoding the audio sample and the voltage pulse sample. In the training process, first, a plurality of audio samples are collected. The audio samples can be audio recorded by a plurality of volunteers in a quiet environment. At the same time of collecting the audio samples, the humidity information is collected by the LIF neuron circuit and the voltage pulse sample is output. Then, random noise is added to the collected audio samples. For example, the noise level can be set to 10%, 20%, and 30%, respectively. In this way, three groups of audio samples containing different noise can be generated to simulate different degrees of background interference in reality. Next, the audio sample and the voltage pulse sample are fused and encoded to obtain the feature vector sample. In the fusion and encoding process, the voltage pulse sample is first widened, then multiplied by a scaling factor, and then added point by point with the audio sample to obtain the fusion sample data. Finally, the fusion sample data is MFCC encoded to obtain the feature vector sample. Then, the embodiment can pre-identify the audio sample to determine the recognition result corresponding to the audio sample, thereby obtaining the speech recognition sample. Then, the correspondence between the feature vector sample and the speech recognition sample is established, and finally the neural network is trained according to the correspondence to obtain the trained neural network.
[0072] Based on this, when the embodiment inputs the fusion data vector group into the trained neural network, the trained neural network can automatically recognize the fusion data vector group to output the speech recognition result. In one implementation, the neural network of the embodiment can be a spiking neural network. As shown in Figure 9 , the spiking neural network can be used to recognize and classify the fusion data vector group to output the speech classification result, such as Figure 9 category 1, category 2, and category 3 output in , thereby realizing speech recognition. In other implementations, the neural network of the embodiment can also be replaced by a binary neural network, a feedforward neural network, a recurrent neural network, etc. The embodiment is not limited thereto.
[0073] In addition, in order to verify that the humidity information has an optimization effect on speech recognition in a noisy environment, the experiment in Figure 2As shown in the above embodiments, the audio data containing random noise can be encoded by MFCC to realize feature extraction and obtain an audio data vector group. Then, the audio data vector group and the fusion data vector group are respectively input into the trained neural network to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group. Specifically, as shown in the above embodiments, the audio data vector group and the fusion data vector group are respectively input into the trained neural network to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group. Figure 10 As shown in the above embodiments, the audio data containing random noise can be encoded by MFCC to realize feature extraction and obtain an audio data vector group. Then, the audio data vector group and the fusion data vector group are respectively input into the trained neural network to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group. Specifically, as shown in the above embodiments, the audio data vector group and the fusion data vector group are respectively input into the trained neural network to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group. Figure 10 As shown in the above embodiments, the audio data containing random noise can be encoded by MFCC to realize feature extraction and obtain an audio data vector group. Then, the audio data vector group and the fusion data vector group are respectively input into the trained neural network to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group. Specifically, as shown in the above embodiments, the audio data vector group and the fusion data vector group are respectively input into the trained neural network to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group.
[0074] In summary, the present embodiment first acquires a voice signal based on a microphone device to obtain audio data, and acquires humidity information based on a LIF neuron circuit, and outputs voltage pulse data based on the humidity information by the LIF neuron circuit, wherein the LIF neuron circuit includes a humidity-sensitive memristor. Then, random noise is applied to the audio data, and the audio data containing random noise is fused and encoded with the voltage pulse data to obtain a fusion data vector group; then, the fusion data vector group is input into a pre-trained neural network to output a voice recognition result. The present embodiment utilizes the high sensitivity of the LIF neuron circuit based on the humidity-sensitive memristor to humidity fluctuations to convert humidity information into corresponding voltage pulse data, thereby providing complementary feature information for voice recognition, breaking through the recognition limit of a single voice signal, and significantly improving the voice recognition accuracy and robustness in a noisy environment by utilizing multi-modal information fusion.
[0075] Based on the above embodiments, the present application also provides a noise-resistant voice recognition system based on a humidity-sensitive memristor. The system of the present embodiment is used to implement the steps in the above method embodiments. Specifically, as shown in the above embodiments, the system of the present embodiment is used to implement the steps in the above method embodiments. Figure 11As shown, the system includes a multimodal information acquisition module 10, a fusion and encoding module 20, and a speech recognition module 30. Specifically, the multimodal information acquisition module is used to acquire speech signals based on a microphone device to obtain audio data, and to acquire humidity information based on a LIF neuron circuit, and the LIF neuron circuit outputs voltage pulse data based on the humidity information, wherein the LIF neuron circuit includes a humidity-sensitive memristor. The fusion and encoding module 20 is used to apply random noise to the audio data, and to fuse and encode the audio data containing random noise with the voltage pulse data to obtain a fused data vector group. The speech recognition module 30 is used to input the fused data vector group into a pre-trained neural network and output a speech recognition result, wherein the neural network is a network model trained based on the correspondence between feature vector samples and speech recognition samples, and the feature vector samples are vector samples obtained by pre-fusing and encoding audio samples and voltage pulse samples.
[0076] The humidity-sensitive memristor-based noise-resistant speech recognition system in this embodiment is based on the same principle as the steps in the above method embodiments, and will not be repeated here.
[0077] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 12 As shown. The terminal may include one or more processors 100 ( Figure 12 (Only one is shown in the diagram), memory 101, and computer program 102 stored in memory 101 and executable on one or more processors 100. For example, a noise-resistant speech recognition program based on humidity-sensitive memristors. When one or more processors 100 execute computer program 102, they can implement various steps in the embodiments of the noise-resistant speech recognition method based on humidity-sensitive memristors. Alternatively, when one or more processors 100 execute computer program 102, they can implement the functions of various modules / units in the embodiments of the noise-resistant speech recognition system based on humidity-sensitive memristors, which is not limited here.
[0078] In one embodiment, the processor 100 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0079] In one embodiment, the memory 101 can be an internal storage unit of the electronic device, such as a hard disk or a memory of the electronic device. The memory 101 can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 101 can include both the internal storage unit and the external storage device of the electronic device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 can also be used to temporarily store data that has been output or will be output.
[0080] Those skilled in the art can understand that, Figure 12 The block diagram shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific terminal can include more or less components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0081] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, operating database or other medium used in the embodiments of the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0082] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method of noise-robust speech recognition based on humidity-sensitive memristors, characterized by, The method comprises: Based on the microphone device, the voice signal is collected to obtain the audio data, and the humidity information is collected based on the LIF neuron circuit, and the voltage pulse data is outputted by the LIF neuron circuit based on the humidity information, wherein the LIF neuron circuit comprises a humidity-sensitive memristor; Random noise is applied to the audio data, and the audio data containing random noise is fused and encoded with the voltage pulse data to obtain a fusion data vector group; The fusion data vector group is inputted into a pre-trained neural network to output a voice recognition result, wherein the neural network is a network model trained based on the corresponding relationship between a feature vector sample and a voice recognition sample, and the feature vector sample is a vector sample obtained by pre-fusing and encoding an audio sample and a voltage pulse sample.
2. The noise-robust speech recognition method based on the humidity-sensitive memristor according to claim 1, wherein, The LIF neuron circuit is composed of the humidity-sensitive memristor, the current-limiting resistor, the load resistor and the parallel capacitor. When the voice signal is accompanied by humid airflow, the resistance of the humidity-sensitive memristor decreases with the increase of humidity, which accelerates the charging process of the parallel capacitor, so that the humidity-sensitive memristor reaches the threshold value faster and triggers pulse discharge.
3. The noise-robust speech recognition method based on the humidity-sensitive memristor according to claim 2, wherein, The preparation process of the humidity-sensitive memristor comprises: 1 milliliter of Nafion solution is measured, the Nafion solution is mixed with 3.6 milliliters of anhydrous ethanol, and oscillation is performed for 5 minutes to uniformly mix the Nafion solution and the anhydrous ethanol to obtain a Nafion diluent; The glass substrate with a 185-nanometer-thick ITO electrode is ultrasonically cleaned with deionized water, then dried with nitrogen, and dried at 120°C for 30 minutes; 200 microliters of the Nafion diluent are taken, the Nafion diluent is spin-coated on the cleaned glass substrate, then annealed at 100°C for 45 minutes to obtain a 30-nanometer-thick Nafion polymer medium layer; An electrode is prepared using a metal mask plate with a pre-prepared pattern to obtain a humidity-sensitive memristor.
4. The noise-robust speech recognition method based on the humidity-sensitive memristor according to claim 3, characterized in that, The spin-coating speed condition is 500 rpm first, maintained for 5 seconds, and then 1500 rpm, maintained for 40 seconds.
5. The noise-robust speech recognition method based on the humidity-sensitive memristor according to claim 3, wherein, An electrode is prepared using a metal mask plate with a pre-prepared pattern to obtain a humidity-sensitive memristor, comprising: The metal mask plate is covered on the spin-coated glass substrate, 50 nanometers of silver are evaporated on the glass substrate by the method of thermal evaporation, and then the metal mask plate is removed to obtain the humidity-sensitive memristor.
6. The noise-robust speech recognition method based on the humidity-sensitive memristor according to claim 1, wherein, The audio data containing random noise is fused and encoded with the voltage pulse data to obtain a fusion data vector group, comprising: The voltage pulse data is widened, and the widened voltage pulse data is multiplied by a scaling factor to obtain voltage pulse data with a matched order of magnitude; The audio data containing random noise is point-by-point added to the voltage pulse data with a matched order of magnitude to obtain voice humidity fusion data; The voice humidity fusion data is MFCC encoded to obtain a fusion data vector group.
7. The noise-robust speech recognition method based on the humidity-sensitive memristor according to claim 1, wherein, The method further comprises: The audio data containing random noise is MFCC encoded to obtain an audio data vector group; The audio data vector group and the fusion data vector group are input into a trained neural network respectively, to obtain a first recognition result corresponding to the audio data vector group and a second recognition result corresponding to the fusion data vector group, wherein the accuracy of the second recognition result is higher than that of the first recognition result.
8. A noise-robust speech recognition system based on humidity-sensitive memristors, characterized by The system is used to implement the steps of the anti-noise speech recognition based on the humidity-sensitive memristor in any one of claims 1-7, and the system comprises: A multi-modal information acquisition module is configured to acquire a voice signal based on a microphone device to obtain audio data, and acquire humidity information based on a LIF neuron circuit, and output voltage pulse data based on the humidity information by the LIF neuron circuit, wherein the LIF neuron circuit comprises a humidity-sensitive memristor; A fusion and encoding module is configured to apply random noise to the audio data, and fuse and encode the audio data containing random noise and the voltage pulse data to obtain a fusion data vector group; A speech recognition module is configured to input the fusion data vector group into a pre-trained neural network to output a speech recognition result, wherein the neural network is a network model trained based on a corresponding relationship between a feature vector sample and a speech recognition sample, and the feature vector sample is a vector sample obtained by previously fusing and encoding an audio sample and a voltage pulse sample.
9. A terminal, characterized by comprising: The terminal comprises a memory, a processor, and a humidity-sensitive memristor-based anti-noise speech recognition program stored in the memory and executable on the processor, and when the processor executes the humidity-sensitive memristor-based anti-noise speech recognition program, the steps of the humidity-sensitive memristor-based anti-noise speech recognition method in any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a humidity-sensitive memristor-based anti-noise speech recognition program, and the humidity-sensitive memristor-based anti-noise speech recognition program implements the steps of the humidity-sensitive memristor-based anti-noise speech recognition method in any one of claims 1-7 on the computer-readable storage medium.
Citation Information
Patent Citations
Hybrid input device for touchless user interface
CN105051812A
Silent speech recognition method and system
CN118173117A