Flexible liquid metal sensor and voiced / unvoiced speech recognition system

By combining flexible liquid metal sensors with machine learning, mixed-mode acoustic and vibration signals from the larynx are acquired, solving the speech recognition challenges in existing technologies such as speech impairment and noisy environments, and achieving highly efficient speech semantic recognition.

WO2026113045A1PCT designated stage Publication Date: 2026-06-04CHINA ACAD OF LAUNCH VEHICLE TECH

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CHINA ACAD OF LAUNCH VEHICLE TECH
Filing Date
2024-12-04
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing speech recognition technologies have low accuracy in speech impairments and noisy environments, and struggle to effectively identify vocal vibration signals from the throat.

Method used

By using a flexible liquid metal sensor attached to the throat and combining it with machine learning methods, a throat sensor and signal acquisition and analysis system are designed to acquire the acoustic-vibration mixed-mode signals of the throat, thereby achieving both spoken and silent speech recognition.

Benefits of technology

It achieves efficient speech semantic recognition in environments with speech barriers and noise, improving the accuracy and reliability of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136909_04062026_PF_FP_ABST
    Figure CN2024136909_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A flexible liquid metal sensor and a voiced / unvoiced speech recognition system. The flexible liquid metal sensor attached to the throat is configured to detect a hybrid-modal signal consisting of laryngeal low-frequency muscle movements and acoustic vibrations during voiced or unvoiced speech of a person. By designing discrete and continuous sensor layouts to accurately capture the frequency of laryngeal vibrations during phonation and spatial distribution characteristics of muscle movements, pressure information in different directions and at different positions is measured. Laryngeal data acquisition and filtering are achieved by means of a signal acquisition and analysis system, so as to construct a relational dataset between laryngeal acoustic-vibration hybrid-modal signals and expressed speech semantics. By establishing a deep neural network training model based on time-frequency features to train the relational dataset, the content expressed by a person during voiced or unvoiced speech can be recognized, thereby meeting the requirements of persons with speech impairments and the general population for conveying the expressed content in environments where speech is impaired.
Need to check novelty before this filing date? Find Prior Art

Description

A flexible liquid metal sensor and a voice recognition system with and without sound

[0001] This application claims priority to Chinese Patent Application No. 2024117085392, filed on November 27, 2024, entitled "A Flexible Liquid Metal Sensor and a Voice Recognition System", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to a voice and silent speech recognition system based on a flexible liquid metal sensor, belonging to the field of speech recognition technology. Background Technology

[0003] In recent years, researchers have explored numerous experiments to elucidate the correspondence between subtle movements and vibrations on the skin surface and human activity patterns. Speech, in particular, is formed by the combined movements of the throat and facial muscles controlling the shape changes of the resonant cavities used for sound production. Therefore, the recognizability of speech is implicit in muscle movement patterns and skin surface vibrations. Flexible sensors possess high conformability, can directly contact the skin surface, and are highly sensitive to movement and deformation. "Non-acoustic" recognition algorithms based on flexible sensors typically acquire vibrations of the vocal cords or facial skin / muscles through sensors, and after signal enhancement and denoising, use machine learning methods to classify the data and obtain the subject's intended expression.

[0004] In 2021, Rekimoto et al. from the University of Tokyo proposed Derma, which uses two small 6-DOF accelerometers / angular velocity sensors to acquire 12 channels of motion of the skin near the jaw (Proceedings of the Augmenmented Humans International Conference, 2021). After standardization, the data is fed into a one-dimensional convolutional neural network to predict 35 commands with an accuracy of 94.49%.

[0005] In 2022, Ravenscroft et al. from the University of Cambridge fabricated a graphene strain sensor (Sensor, 2022, 22, 299). During the experiment, they attached it to the throat and measured the change in graphene resistance to obtain the amplitude of vocal cord vibration. The authors then extracted six temporal features and Mel-Frequency Cepstral Coefficients (MFCC) features, concatenated them, and input them into a classification network consisting of a convolutional layer and an LSTM layer. The network achieved a classification accuracy of 54.6% for 15 words.

[0006] In 2023, Yang et al. from Tsinghua University proposed a new generation of artificial larynx (AT), published in *Nature Machine Intelligence* (Nature Machine Intelligence, 2023, 5, 169-180). They used graphene sensors to capture vocal cord vibrations and converted the voltage signal into a standardized 48Kps, 16-bit Pulse Code Modulation (PCM) signal using a downsampling method. The authors then performed a Fast Fourier Transform (FFT) on the signal with a window size of 4096 and a length of 2048, obtaining a 227×227×3-dimensional spectrogram. This spectrogram was input into an AlexNet convolutional network to extract features, and the ReliefF algorithm was used to rank the features. The top 7 features were then input into a Support Vector Machine (SVM) for classification. The prediction accuracy was 98.13% in 6 word classification tasks and 81.25% in 6 sentence classification tasks.

[0007] In contrast, the flexible liquid metal sensor exhibits a minimum single-point pressure threshold of 5 mN and a maximum of up to 900 mN, with a strong linear correlation between load and resistance change rate. This indicates that the sensor is highly sensitive to small loads while also being suitable for applications with a wide range of load variations, maintaining local sensitivity even during these range changes. The flexible liquid metal sensor combines the stretchability and high conformability of PDMS flexible microtubes with the strong conductivity of liquid metal, resulting in highly sensitive pressure response and excellent skin adhesion, making it suitable for wearable sensor fabrication. The continuous extraction and curing process used to produce PDMS flexible microtubes eliminates the need for complex manufacturing techniques and expensive equipment such as photolithography, significantly reducing manufacturing costs. Summary of the Invention

[0008] The technical problem solved by this application is to overcome the shortcomings of the prior art and provide a voice and silent speech recognition system based on a flexible liquid metal sensor. This system combines high-efficiency manufacturing technology with artificial intelligence technology to systematically analyze the throat acoustic vibration signal with mixed modes, thereby solving the speech semantic recognition problem with high efficiency. It can be widely applied in environments with speech impairments.

[0009] This application proposes a non-acoustic speech recognition system that utilizes the characteristics of a flexible liquid metal sensor adapted to measure laryngeal vibration signals. Specifically, it acquires mixed-mode acoustic and vibration signals from the larynx when a person is speaking or speaking silently by designing a laryngeal sensor and a signal acquisition and analysis system, and then uses machine learning methods to recognize the content expressed by the person speaking or speaking silently.

[0010] The technical solution provided in this application is as follows:

[0011] A flexible liquid metal sensor is attached to the throat and generates an acoustic vibration resistance analog signal based on the vibration of the throat when words and sentences are spoken or unspoken. The flexible liquid metal sensor includes at least one sensor unit, which includes a PDMS microtube filled with liquid metal. The two ends of the PDMS microtube are encapsulated with wires. The liquid metal is a pure metal or alloy with a melting point in the temperature range of -20℃ to 25℃. The sensor can sense pressure at a single point in the range of 1mN to 1000mN. The upper limit of the frequency response of the sensor unit is >5000Hz.

[0012] The PDMS microtube has an inner diameter of 10–400 μm, a length exceeding 50 cm, and a wall thickness of 5–600 μm.

[0013] The PDMS microtubes are produced using a continuous extraction and curing process.

[0014] The method for preparing the PDMS microtube includes: vertically passing a heated metal wire through a PDMS precursor, the PDMS precursor forming a PDMS coating around the metal wire, heating and curing the metal wire with the PDMS coating in a heating unit, the PDMS coating being cured in situ to form a tubular PDMS microtube, and separating the PDMS microtube from the metal wire by ultrasonic treatment to obtain the PDMS microtube.

[0015] The temperature of the heated metal wire is 90°C to 110°C, and the residence time of the metal wire in the PDMS precursor is 2-6 minutes.

[0016] The conditions for heat curing the metal wire with the PDMS coating in the heating unit are: 90℃~100℃, 3-5 minutes.

[0017] It also includes a PDMS flexible substrate, on which multiple sensor units are disposed. The PDMS microtubes of the multiple sensor units are connected side by side to the surface of the PDMS flexible substrate to form a two-dimensional sensor.

[0018] The fabrication method of the two-dimensional sensor includes: uniformly dispensing PDMS onto a silicon wafer and performing homogenization to obtain a flexible substrate blank with uniform thickness; performing plasma treatment on the flexible substrate blank to obtain a PDMS flexible substrate; performing plasma treatment on the PDMS microtubes to obtain treated microtubes; sequentially arranging the treated microtubes on the surface of the PDMS flexible substrate, and then heating to 90℃~110℃ and holding for 3-5 minutes, so that the treated microtubes adhere to the surface of the PDMS flexible substrate.

[0019] A voice recognition system based on a flexible liquid metal sensor includes: a flexible liquid metal sensor, a sound vibration signal sampling circuit board, a server, an audio output circuit board, and an artificial voice reproduction module. The server is equipped with a voice recognition module.

[0020] A flexible liquid metal sensor is attached to the throat to acquire the acoustic resistance analog signal of the throat during both spoken and silent speech, and sends the acoustic resistance analog signal to the acoustic signal sampling circuit board.

[0021] The acoustic vibration signal sampling circuit board is used to receive the acoustic vibration resistance analog signal obtained by the flexible liquid metal sensor, convert the acoustic vibration resistance analog signal into a voltage analog signal, amplify, filter and sample the voltage analog signal by ADC to obtain an acoustic vibration voltage digital signal, and upload the acoustic vibration voltage digital signal to the server.

[0022] The speech recognition module provides a pre-trained deep learning model, receives the acoustic voltage digital signal sent by the acoustic signal sampling circuit board, performs data processing and feature extraction on the acoustic voltage digital signal, and uses the time-frequency features obtained from the feature extraction as input to the pre-trained deep learning model; based on the input time-frequency features, the pre-trained deep learning model outputs the corresponding words and sentences expressed by the speech, and converts the words and sentences into audio digital signals; the speech recognition module sends the audio digital signals to the audio output circuit board;

[0023] An audio output circuit board receives the digital audio signal, modulates the digital audio signal into multiple analog audio signals, amplifies the multiple analog audio signals, and outputs them to the artificial voice reproduction module.

[0024] The artificial voice reproduction module is used to receive the amplified multi-channel audio analog signals output by the audio output circuit board, and apply the amplified multi-channel audio analog signals to various points of the artificial voice to achieve pronunciation.

[0025] The flexible liquid metal sensor is attached to the throat, including:

[0026] Flexible liquid metal sensors include multiple sensor units and / or multiple two-dimensional sensors;

[0027] Multiple sensor units are discretely placed at different parts of the throat to collect raw data of electrical signal changes from the sensor units. The raw data of electrical signal changes is subjected to spectrum analysis and noise reduction. The time-domain and frequency-domain signals of each sensor unit are analyzed, and the frequency distribution and amplitude of the frequency-domain signal and the distribution of the time-domain signal are observed. The spacing and position of the sensor units are adjusted, and the measurements and analyses are performed again to determine the sensor unit placement position that makes the difference between the time-domain signal and the frequency-domain signal most significant.

[0028] Multiple two-dimensional sensors were attached to different positions in the throat. The change of signal χ along the x-axis of the microtube arrangement direction, dχ / dx, was analyzed. After adjusting the spacing and position of the two-dimensional sensors, the measurement and analysis were performed again to determine the two-dimensional sensor placement position that maximizes the absolute value of dχ / dx.

[0029] The flexible liquid metal sensor is attached to the throat by combining the sensor unit arrangement position and / or the two-dimensional sensor arrangement position.

[0030] The acquisition of the acoustic vibration resistance simulation signal of the larynx during spoken and unspoken speech includes: having the subject utter words and sentences, and a flexible liquid metal sensor acquiring the acoustic vibration resistance simulation signal;

[0031] The throat and / or head are moved, and a flexible liquid metal sensor acquires non-speech vibration analog resistance signals; the throat and / or head movements include coughing, swallowing, yawning, and nodding; to exclude non-speech vibration characteristic resistance signals mixed in with the acoustic vibration resistance analog signals based on the characteristics of the non-speech vibration analog signals. The trained deep learning model is obtained through the following method:

[0032] A dataset relating the acoustic vibration voltage digital signal of the larynx to the vocabulary and sentences expressed in speech is constructed. The time-frequency features of the acoustic vibration voltage digital signal of the larynx are extracted. Based on this dataset, where the vocabulary and sentences are composed of English phonetic symbols or Chinese phonemes, the correspondence between the time-frequency features of the vibration voltage digital signal and the English phonetic symbols or Chinese phonemes in the vocabulary and sentences is obtained. A deep learning training model is then established based on this correspondence between the time-frequency features of the vibration voltage digital signal and the English phonetic symbols or Chinese phonemes.

[0033] The dataset consisting of multiple acoustic, vibration, and voltage digital signals is divided into a training set and a test set. The deep learning training model is trained using k-fold cross-validation by combining the training set and the test set until the deep learning training model meets the accuracy requirements, thus obtaining a well-trained deep learning training model.

[0034] The acoustic vibration signal sampling circuit board includes a first host computer communication module, a first microprocessor, a sensor excitation circuit, and an ADC sampling and filtering circuit. The sensor excitation circuit applies a voltage to the flexible liquid metal sensor. When the sensor is stimulated by a mixture of slight muscle movements in the throat and sound vibrations, the resistance change of the flexible liquid metal sensor becomes an acoustic vibration resistance analog signal. The sensor excitation circuit can convert the resistance change into a voltage signal change, which becomes an acoustic vibration voltage analog signal. The acoustic vibration voltage analog signal is then sent to the ADC sampling and filtering circuit.

[0035] An ADC sampling and filtering circuit receives the acoustic vibration voltage analog signal, amplifies and filters the analog signal using the built-in amplification and bandpass filter circuit of the ADC chip, then converts the analog signal into a digital signal using the ADC chip to obtain the acoustic vibration voltage digital signal, and sends the digital signal to a first microprocessor; the first microprocessor receives the digital signal and sends it to a host computer communication module; it is used to control the operation of the sensor excitation circuit, the ADC sampling and filtering circuit, and the host computer communication module;

[0036] The host computer communication module receives the acoustic vibration voltage digital signal and sends it to the server through a pass-through mode.

[0037] The audio output circuit board includes a second host computer communication module, a second microprocessor, an audio decoder, and an audio excitation circuit.

[0038] The second host computer communication module receives the audio digital signal obtained by the voice recognition module and transmits the audio digital signal to the second microprocessor through a transparent transmission mode.

[0039] The second microprocessor receives the audio digital signal and sends the audio digital signal to the audio decoder;

[0040] An audio decoder receives the digital audio signal, modulates the digital audio signal into multiple analog audio signals, and sends the multiple analog audio signals to an audio excitation circuit.

[0041] An audio excitation circuit receives the multiple audio analog signals and amplifies them to obtain amplified multiple audio analog signals.

[0042] Data transmission between the first host computer communication module and the server, and between the second host computer communication module and the server, is conducted via wireless or wired communication. The wireless communication methods include Bluetooth, WIFI, ZigBee, and cellular network technologies, while the wired communication methods include Ethernet, USB, and serial port connection technologies.

[0043] The extraction of the time-frequency characteristics of the acoustic vibration voltage digital signal of the throat includes: using the short-time Fourier transform time-varying spectrum or Mel time-varying frequency cepstral coefficients as the extracted time-frequency characteristics.

[0044] The deep learning training model is: ResNet50 network, EfficientFormerV2 and / or connection-based temporal classifier CTC.

[0045] When the deep learning training model is a ResNet50 network, the input of the ResNet50 network is the time-frequency feature of the acoustic-vibration voltage digital signal, i.e., the short-time Fourier transform time-varying spectrum or the time-varying Mel frequency cepstral coefficients; the short-time Fourier transform time-varying spectrum or the time-varying Mel frequency cepstral coefficients are passed through 49 convolutional layers to obtain a feature vector of dimension M; the feature vector is input into a fully connected layer, and the output is a vector of dimension N. The vector of dimension N is passed through Softmax and the index with the highest probability is selected. The index with the highest probability indicates that the English phonetic symbol or Chinese phoneme corresponding to the index is the English phonetic symbol or Chinese phoneme at that time point.

[0046] When the deep learning training model is EfficientFormerV2, the input of EfficientFormerV2 is the time-frequency feature of the acoustic-vibration voltage digital signal, that is, the short-time Fourier transform time-varying spectrum or the time-varying Mel frequency cepstral coefficients; after the short-time Fourier transform time-varying spectrum or the time-varying Mel frequency cepstral coefficients are passed through a depthwise separable convolution and multi-head attention intermediate layer, a vector of dimension N is output through a fully connected layer. After the vector of dimension N is passed through Softmax, the index with the highest probability is selected. The index with the highest probability indicates that the English phonetic symbol or Chinese phoneme corresponding to the index is the English phonetic symbol or Chinese phoneme at that time point.

[0047] The ResNet50 or EfficientFormerV2 outputs the English phonetic symbols or Chinese phonemes for all time points within the time period to the temporal classifier CTC; finally, the CTC algorithm arranges the English phonetic symbols and / or Chinese phonemes into English / Chinese words or sentences based on the consecutive English phonetic symbols, phonemes, and English and Chinese vocabulary; the output English / Chinese words or sentences are used as the speech text output by the deep learning training model.

[0048] In summary, this application includes at least the following beneficial technical effects:

[0049] Current automatic speech recognition technology has limitations. Speech-based interaction is easily affected by human physiological limitations and environmental interference, such as environmental noise, which reduces recognition accuracy. Human vocalization is a complex process involving precise coordination between multiple organs, including the tongue, lips, jaw, vocal cords, and lungs. Damage to any organ or part can lead to the interruption of related pathways, resulting in speech disorders. This application proposes a speech recognition system based on a flexible liquid metal sensor to address the above problems, enabling people with speech disorders to speak and achieving accurate speech transmission in noisy environments. Attached Figure Description

[0050] Figure 1. Schematic diagram of the experimental setup for preparing PDMS microtubes;

[0051] Figure 2. Physical images of PDMS microtubes with different inner diameters;

[0052] Figure 3. Resistance change of the flexible liquid metal sensor under pressure;

[0053] Figure 4. Time domain diagram (a) and frequency domain diagram (b) of the flexible liquid metal sensor measurement at an input frequency of 5500 Hz;

[0054] Figure 5. A physical diagram of the distribution of microtubes arranged side by side on a two-dimensional surface;

[0055] Figure 6. Schematic diagram of continuous distributed sensor application;

[0056] Figure 7. Schematic diagram of the signal acquisition and analysis system;

[0057] Figure 8. Speech recognition method based on phoneme decomposition;

[0058] Figure 9. Schematic diagram of STFT calculation steps;

[0059] Figure 10. Schematic diagram of MFCC feature extraction process;

[0060] Figure 11. ResNet50 architecture diagram;

[0061] Figure 12. Schematic diagram of residual element;

[0062] Figure 13. Architecture diagram of EfficientFormerV2;

[0063] Figure 14. Short-time Fourier transform image used for identification;

[0064] Figure 15. Classification confusion matrix for the test set;

[0065] Figure 16. Recognition accuracy and loss function. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments disclosed in this application will be described in further detail below with reference to the accompanying drawings.

[0067] This application discloses a flexible liquid metal sensor, as shown in Figure 1. A continuous extraction and curing process is used to produce flexible PDMS microtubes. Specifically, a metal wire heated to 90℃~110℃ is vertically passed through a PDMS precursor (i.e., a dimethylsiloxane precursor). The viscosity and surface tension cause the PDMS precursor to form a PDMS coating around the metal wire. Then, the PDMS-coated metal wire is heated to 90℃~100℃ in an electric heating unit, where the PDMS is cured in situ, forming a tubular PDMS microtube. The PDMS microtube is then separated from the metal wire by ultrasonic treatment. The resulting flexible PDMS microtube, as shown in Figure 2, has an inner diameter of 10~400μm, a length exceeding 50cm, and uniform inner and outer diameters with an outer diameter-to-inner diameter ratio of 2:1, 3:1, or 4:1. The produced PDMS microtube has a high aspect ratio, thin walls, and is tough, preventing sagging or collapse during handling and use. The aspect ratio is length / diameter, and the wall thickness is 5-600μm.

[0068] Liquid metal is poured into the produced PDMS microtubes, and wires are encapsulated at both ends of the microtubes to form sensor units. Each sensor unit is a single flexible liquid metal sensor. The liquid metal is a pure metal or alloy with a melting point in the temperature range of -20℃ to 25℃, such as gallium-indium alloy. Flexible liquid metal sensors made of gallium-indium alloy possess characteristics such as high sensitivity, a large measurement threshold range, good conformality, excellent frequency response, and immunity to humidity. PDMS flexible microtubes are easily deformable and can be used to design and fabricate devices with different geometries, exhibiting good conformality. The liquid metal is encapsulated in flexible microtubes, isolating it from ambient water, and has excellent resistance to changes in ambient humidity. A single-point compression experiment was conducted on the gallium-indium alloy flexible liquid metal sensor, as shown in Figure 3. The experimental results show that the minimum single-point sensing pressure threshold of the flexible liquid metal sensor developed in this application is 5mN, and the maximum is up to 900mN, with a good linear correlation between load and resistance change rate. This indicates that the sensor is highly sensitive to small loads and can be applied to working conditions with a large range of load changes, while maintaining local sensitivity during the range change process.

[0069] Traditional piezoresistive sensors are limited by their inherent frequency response characteristics, typically resulting in a low upper limit for measurement frequency. By attaching a flexible liquid metal sensor to a forced vibration table, the sensor vibrates at a certain frequency with the table. The fixed end exerts relative motion on the sensor, creating a single-point compression, thus achieving a point vibration effect. By testing vibration at a frequency of 5500Hz and analyzing the frequency domain diagram of the filtered resistance change of the flexible liquid metal sensor (Figure 4), it can be observed that the frequency signal measured by the flexible liquid metal sensor is essentially consistent with the frequency signal output by the vibration table. The experiment demonstrates that the frequency response of the flexible liquid metal sensor is extremely excellent among piezoresistive sensors, making it suitable for measuring throat acoustic vibration signals.

[0070] Multiple microtube sensors are connected to form a two-dimensional continuous distributed flexible sensor. Using a spin coater, a fixed amount of PDMS is uniformly dropped onto a silicon wafer. The spin coater is then rotated at 500 RPM for 10 seconds, followed by 1000 RPM for 10 seconds, to create a flexible substrate blank with uniform thickness. The flexible substrate blank and PDMS microtubes are then subjected to plasma treatment to improve surface activity, resulting in a PDMS flexible substrate and treated microtubes. The treated microtubes are then sequentially arranged at designated positions on the PDMS flexible substrate, forming a two-dimensional structure as shown in Figure 5. The substrate is heated to 90℃~110℃ and held for 3-5 minutes, causing the treated microtubes to adhere to the PDMS flexible substrate. Liquid metal is then injected into the treated microtubes, and wires are encapsulated at both ends of the treated microtubes to create a two-dimensional continuous distributed flexible sensor.

[0071] Single flexible liquid metal sensors were attached to different parts of the throat in a discrete layout, and raw data of electrical signal changes were collected. Spectral analysis and noise reduction were performed on the sensors. The time and frequency domain signals of each sensor were analyzed, and the frequency distribution, amplitude, and time domain signal distribution were observed. After adjusting the spacing and position of the sensors, measurements and analyses were performed again to explore the sensor layout that resulted in the most significant differences between the time and frequency domain signals.

[0072] As shown in Figure 5 above, a two-dimensional continuous distributed flexible sensor is fabricated by arranging n microtubes side-by-side into a continuous distributed two-dimensional structure. As shown in Figure 6, when this two-dimensional continuous distributed flexible sensor is attached to the subject's larynx, the two-dimensional pressure distribution in the larynx during vocalization can be measured. By attaching the two-dimensional continuous distributed flexible sensor to different positions on the larynx, the change in signal χ along the x-axis of the microtube arrangement direction (dχ / dx) is analyzed to explore the sensor placement position that maximizes the absolute value of dχ / dx.

[0073] Discrete distributed sensors (i.e., single flexible liquid metal sensors) and continuous distributed sensors (i.e., two-dimensional continuous distributed flexible sensors) are attached to different locations in the larynx. Subjects speak words and sentences, and the signals are sampled, filtered, and noise-reduced by a signal acquisition and analysis system to obtain digital signals of larynx vibration.

[0074] To eliminate interference from non-speech vibration feature signals mixed into the detected mixed modal signals, a dataset of sensor data collected during laryngeal / head movements was established, including behaviors such as coughing, swallowing, yawning, and nodding, in order to distinguish the differences in modal signals detected by different behaviors.

[0075] The next step is to extract features based on the acoustic and vibration digital signals. The following methods can be used for feature extraction:

[0076] Based on the time-domain features of acoustic and vibration digital signals, such as root mean square, variance, bias, kurtosis, shape factor, and entropy, a feature array is constructed. This array is then trained or predicted using machine learning algorithms to remove redundant information and retain important features. The time-frequency domain features of acoustic and vibration digital signals can be obtained using short-time Fourier transform and time-varying Mel-frequency cepstral coefficient extraction methods.

[0077] After feature extraction, feature analysis is performed on the vocabulary to summarize the correspondence between features and English phonetics or phonemes. Based on machine learning algorithms, a mapping relationship is established between the digital signal of laryngeal vocalization and the vocabulary and sentences expressed in speech. A connection-time classification network system based on phoneme decomposition is constructed, and the network is trained with a large amount of data. The trained neural network is then used to test the test set to identify specific instruction words.

[0078] This embodiment also provides a spoken and silent speech recognition system based on a flexible liquid metal sensor, as shown in Figure 7. It includes a flexible liquid metal sensor, an acoustic vibration signal sampling circuit board, a server, an audio output circuit board, and an artificial voice reproduction module. The server is equipped with a speech recognition module. It completes three levels of functions: acquisition of mixed-mode acoustic vibration signals from the throat, signal processing, and audio reproduction.

[0079] Among them, the flexible liquid metal sensor is used to be attached to the throat to obtain analog resistance signals (which are acoustic vibration signals).

[0080] The acoustic vibration signal sampling circuit board includes a first host computer communication module, a first microprocessor, a sensor excitation circuit, and an ADC sampling and filtering circuit. The excitation power supply of the sensor excitation circuit applies voltage to the sensor. When a slight muscle movement of the throat and a mixed signal of sound vibration stimulate the flexible liquid metal sensor, the impedance of the flexible liquid metal sensor will change. The sensor excitation circuit can effectively convert this resistance change into a voltage signal change. The ADC sampling and filtering circuit receives the changing voltage signal, which is the acoustic vibration voltage analog signal. The ADC chip has a built-in amplification circuit and a bandpass filter circuit. Through the built-in amplification circuit and bandpass filter circuit of the ADC chip, the acoustic vibration voltage analog signal is amplified and filtered. Then, the ADC chip samples and converts the acoustic vibration voltage analog signal into a digital signal. The first microprocessor receives the acoustic vibration voltage digital signal and sends it to the host computer communication module. The host computer communication module is used to receive the acoustic vibration voltage digital signal and simultaneously send it to the server through a pass-through mode.

[0081] The acoustic and vibration signal sampling circuit board also includes a power management system, which consists of a charging circuit, a battery, a sensor excitation voltage converter, and a step-down regulator (LDO). The charging circuit can be wired or wirelessly charged using a coil. It can charge a 3.7V battery and has a 4.2V maximum voltage protection threshold to prevent damage from overvoltage. The battery can be a polymer lithium battery, offering advantages such as high power density and customizable form factor.

[0082] As shown in Figure 8, the server speech recognition module provides a pre-trained deep learning model for receiving the acoustic vibration voltage digital signal from the throat, performing data processing and feature extraction on the digital signal, and using the time-frequency features obtained from the feature extraction as input to the pre-trained deep learning model; based on the input time-frequency features, the pre-trained deep learning model outputs the corresponding words and sentences expressed by the speech, and converts the words and sentences into audio digital signals; the speech recognition module sends the audio digital signals to the audio output circuit board.

[0083] The trained deep learning training model is obtained through the following methods: constructing a dataset of the relationship between the digital signal of vocal vibration voltage of the larynx and the expressed words and sentences; obtaining the correspondence between the time-frequency characteristics of the digital signal of vocal vibration voltage and the English phonetic symbols or Chinese phonemes in the words and sentences according to the dataset of the relationship between the digital signal of vocal vibration voltage of the larynx and the expressed words and sentences, and the words and sentences being composed of English phonetic symbols or Chinese phonemes; establishing a deep learning training model based on the correspondence between the time-frequency characteristics of the digital signal of vocal vibration voltage and the English phonetic symbols or Chinese phonemes. A phoneme is the smallest speech unit divided according to the natural attributes of speech, and all speech can be composed of phonemes. For example, the pinyin of "中" is "zhong", which is composed of three phonemes: "zh", "o", and "ng". In "General Theory of Modern Chinese", all pronunciation situations in Chinese Mandarin are represented by 38 phonemes, as shown in Table 1.

[0084] This application decomposes the speech text into combinations of phonemes, and uses deep learning methods to predict English phonetic symbols or Chinese phonemes to achieve general vocal / non-vocal recognition of speech.

[0085] Divide the dataset composed of multiple digital signals of vocal vibration voltage into a training set and a test set; combine the training set and the test set, and use the k-fold cross-validation method to train the deep learning training model until the deep learning training model meets the accuracy requirements, and obtain the trained deep learning training model.

[0086] After the model is trained, input the digital signal of vocal vibration voltage detected by this system into the speech recognition module, and the words and sentence text expressed by the recognized speech can be output and converted into an audio digital signal.

[0087] Table 1 Chinese Phoneme List

[0088] The audio output circuit board includes a second host computer communication module, a second microprocessor, an audio decoder, and an audio excitation circuit. The second host computer communication module is used to receive the audio digital signal in the lossless modulation format output by the speech recognition module, and at the same time transmit the audio digital signal to the second microprocessor in the transparent transmission mode. The second microprocessor sends the audio digital signal to the audio decoder; The audio decoder modulates the audio digital signal into multiple channels of audio analog signals and sends the multiple channels of audio analog signals to the audio excitation circuit. The audio excitation circuit amplifies the multiple channels of audio analog signals and outputs them to the artificial voice reproduction module.

[0089] The artificial voice reproduction module is used to receive the amplified multiple channels of audio analog signals output by the audio output circuit board and apply the multiple channels of audio analog signals to each point of the artificial throat to achieve pronunciation.

[0090] This system is suitable for acquiring laryngeal acoustic vibration signals and speech recognition in both audible and non-audible situations. It can also help patients with vocal cord damage or hoarseness to regain their voice.

[0091] To train a deep learning model for speech recognition, a method for acquiring acoustic vibration signals was designed. During data acquisition using a flexible liquid metal sensor, a display randomly shows m words or sentences. Subjects are required to accurately record speech and acquire laryngeal signals based on on-screen prompts, whether speaking or remaining silent. Before formal data acquisition, each subject needs training to ensure proficiency in spoken / silent speech interaction with different words or sentences. During formal acquisition, each subject repeats the data acquisition six times, acquiring 50 commands each time. The computer (server) synchronously saves the data sampled by the flexible sensor to folders with corresponding labels according to the on-screen instructions.

[0092] The speech recognition method based on phoneme decomposition is shown in Figure 8.

[0093] The flexible sensor signals acquired during data acquisition contain noisy signals. By observing the spectral characteristics of the signals, a notch filter is used to filter out power frequency electromagnetic interference.

[0094] Feature extraction is performed through a speech recognition algorithm module. The main task of feature extraction is to obtain a set of multi-dimensional feature vectors that can effectively distinguish different English phonetic symbols or phonemes. The greater the difference between the multi-dimensional feature vectors corresponding to different letters or phonemes, the better the recognition performance of the model trained based on these feature vectors.

[0095] The short-time Fourier transform frequency spectrum and Mel frequency cepstral coefficients were selected as the extracted features. The accuracy of the two features in speech recognition was tested respectively, and the feature with higher accuracy was selected as the final extracted feature.

[0096] The Short Time Fourier Transform (STFT) estimates the power of a signal at each specific time and frequency point in the time-frequency domain to obtain the time-varying spectrum.

[0097] Figure 9 shows the basic steps for calculating the STFT: ① Select a window function of finite length; ② Place the window on the signal starting from the signal start point; ③ Divide the signal into segments and weight them using the window to generate a series of data segments; ④ Calculate the spectrum of the window data segments using the Fast Fourier Transform (FFT); ⑤ Slide the window along the time axis; ⑥ Repeat steps ③ to ⑤ until the window reaches the end of the signal.

[0098] Mel frequency cepstral coefficients:

[0099] The Mel scale frequency is proposed based on the special nonlinear characteristics of human hearing and can be calculated using formula (1):

[0100] Where Mel(f) is the Mel scale frequency; f is the frequency.

[0101] The extraction process of Mel-Frequency Cepstral Coefficients (MFCC) features in this application is shown in Figure 10. First, the acoustic vibration digital signal is pre-emphasized, then framed and windowed, followed by Fast Fourier Transform, then filtered and logarithmically obtained using the Mel filter bank, thus obtaining the Mel-Frequency Spectral Coefficients (MFSC) features. Finally, the MFSC is subjected to Discrete Cosine Transform to obtain the time-varying MFCC features.

[0102] The pre-emphasis process mainly amplifies high-frequency signals through a pre-emphasis filter. After pre-emphasis, the steps and functions of framing, windowing, and FFT of the signal are similar to those of the short-time Fourier transform process. The window function used for each frame of signal in the FFT process is the Hamming window, and its expression is shown in equation (2):

[0103] Where w[n] is the value of the Hamming window at the nth sampling point; n is the index of the sampling point in the window; and N is the length of the window.

[0104] Next, the signal after the short-time Fourier transform is filtered using a Mel filter bank. Each filter in the Mel filter bank is a triangle with a center frequency response of 1, and both sides decrease linearly towards 0 at the same rate until the response reaches 0, which is the center frequency of two adjacent filters. The number of Mel filter banks used in this application is N. Mel The expression for the corresponding filter is shown in equation (3):

[0105] Among them, H m f(k) is the frequency response of the m-th mel filter; k is the frequency index of the m-th mel filter; f(m) is the center frequency of the m-th mel filter.

[0106] A deep learning model is trained using a dataset composed of STFT time-varying frequency domain features or time-varying MFCC features, and the training continues until the accuracy requirements are met.

[0107] The deep learning training models that can be used are: ResNet50 network, EfficientFormerV2 and / or CTC, as detailed below.

[0108] Convolutional Neural Networks (CNNs), a popular deep neural network algorithm in recent years, are a special type of feedforward neural network that uses backpropagation to train parameters. However, when the number of layers in a CNN network reaches a certain depth, further increases in the number of layers can lead to the vanishing gradient problem, which also causes a decrease in network accuracy. The emergence of residual networks has solved the gradient vanishing problem and also has the advantage of being able to extract deep features from the input, thus resulting in stronger performance in detection or classification.

[0109] The ResNet50 network consists of 49 convolutional layers and 1 fully connected layer, as shown in Figure 11. Its key feature lies in the residual units within the structure, as shown in Figure 12. These residual units contain cross-layer connections; the broken lines in the figure allow the input to be directly passed across layers and added to the output of the convolutional layers. Assuming the input image is x and the output is H(x), and the output after convolution is a nonlinear function F(x), then the final output is H(x) = F(x) + x. The network is then transformed into finding the residual function F(x) = H(x) - x, which is easier to optimize than F(x) = H(x).

[0110] In this application, the input to ResNet50 is the spectral power distribution map or Mel frequency cepstral coefficient map extracted in the previous step. After passing through 49 convolutional layers, a feature vector of dimension M is obtained. This feature vector is then input into the final fully connected layer, which outputs a vector of dimension N. After passing through Softmax, the index with the highest probability is selected, which represents the English phonetic symbol or Chinese phoneme corresponding to that index at that time point.

[0111] EfficientFormerV2 is a type of visual Transformer. The basic principle of Transformer is the self-attention mechanism. Self-attention calculates the attention weights of its own input features and then applies weighted enhancement to those features.

[0112] The overall framework diagram of EfficientFormerV2 is shown in Figure 13 below.

[0113] In this application, the input of EfficientFormerV2 is the spectral power distribution map or Mel frequency cepstral coefficient map extracted in the previous step. After passing through intermediate layers such as depthwise separable convolution and multi-head attention, the output is a vector of dimension N through a fully connected layer. After passing through Softmax, the sequence number with the highest probability is selected, which means that the English phonetic symbol or Chinese phoneme at that time point is the English phonetic symbol or Chinese phoneme corresponding to that sequence number.

[0114] Connectionist Temporal Classification (CTC) can automatically align labels on time series, making it suitable for tasks such as speech recognition and image-text recognition. CTC can automatically align the frame sequences of speech with the corresponding transcribed text sequences during model training, eliminating the need to label the start and end times of each phoneme or phoneme, thus enabling direct classification on the time series.

[0115] Suppose the input is a symbolic sequence X = [x1, x2, ..., x...]. T The symbol ] represents the corresponding output, which is represented by the symbol sequence Y = [y1, y2, ..., y]. U The expression ] indicates that, for a given input X, the posterior probability p(Y|X) of Y should be maximized when training the model. The CTC algorithm introduces "space placeholders" during alignment to represent the alignment result, usually denoted by "∈". Using space placeholders, strict alignment between the input and output can be achieved.

[0116] For a pair of inputs and outputs (X,Y), the goal of CTC is to maximize the sum of probabilities of all valid paths. However, the number of paths corresponding to the output is enormous. The CTC algorithm uses dynamic programming to solve for the conditional probabilities of the outputs. If two paths have the same output at the same time step, these two paths are merged at that time step to reduce the amount of subsequent computation.

[0117] Since solving for conditional probabilities only involves addition and multiplication, it is differentiable and can be optimized using gradient descent.

[0118] The CTC algorithm uses an improved Beam Search method to find the maximum value of the conditional probability. That is, at each time step, it finds the n characters with the highest probability, uses these n characters as a prefix, merges duplicates and deletes ∈ characters, stores the prefix, and continues to find the highest probability character in the next time step.

[0119] In this application, all the English phonetic symbols or Chinese phonemes at all time points within the said time period output by the ResNet50 or EfficientFormerV2 are input into the Connectionist Temporal Classification (CTC), and the English phonetic symbols and / or Chinese phonemes corresponding to the vocabulary or sentences within the said time period are obtained; finally, through the CTC algorithm and based on the composition of consecutive English phonetic symbols, phonemes, and Chinese and English vocabulary before and after, the English phonetic symbols and / or Chinese phonemes are arranged into English / Chinese vocabulary or sentences; and the output English / Chinese vocabulary or sentences are used as the speech text output by the deep learning training model.

[0120] After the model training is completed, this application uses the test set to test the CTC algorithm, and verifies the performance of the algorithm through indicators such as accuracy rate, confusion matrix, and Character Error Rate (CER). Among them, the accuracy rate refers to the proportion of the classification results recognized by the algorithm that are the same as the true results, evaluating the classification accuracy; the confusion matrix respectively counts the proportion of observed values misclassified and correctly classified by the classification model, and displays them through a matrix to evaluate the classification effect of each English phonetic symbol or Chinese phoneme; CER represents the ratio of the number of characters replaced, deleted, and inserted when converting the predicted text into the reference example sentence to the total number of words in the reference example sentence, evaluating the error rate between the predicted text and the standard text.

[0121] According to the above solution, this application uses a multi-channel signal acquisition device based on flexible liquid metal sensors to simultaneously collect the resistance change signals of 6 distributed sensors to extract more features, and uses the method of short-time Fourier transform to obtain the time-frequency spectrum image, as shown in Figure 14. At frequencies of 2500 Hz (which can collect most human voice signals except for individual consonants) and 8000 Hz (which can collect all human voice signals) respectively, a large amount of data has been collected for the letters a, b, c, d, e, the vowels a, o, e, i, u, and the Chinese characters 你 (nǐ), 我 (wǒ), 来 (lái), 听 (tīng), 看 (kàn). And the deep learning training model used in this application is used to identify and classify abcde. The classification confusion matrix of the test set is shown in Figure 15, and the recognition accuracy and loss function of the training set and the test set are shown in Figure 16. The recognition accuracy rate has reached 97%.

[0122] The content not described in detail in the specification of this application belongs to the well-known technology of those skilled in the art.

[0123] The above has described this application in detail in combination with specific implementation manners and exemplary examples, but these descriptions should not be construed as limitations on this application. Those skilled in the art understand that without departing from the spirit and scope of this application, various equivalent replacements, modifications or improvements can be made to the technical solutions and their implementation manners of this application, and these all fall within the scope of this application. The protection scope of this application is subject to the appended claims.

Claims

1. A flexible liquid metal sensor, characterized in that: A flexible liquid metal sensor is attached to the throat and generates an acoustic resistance analog signal based on the vibration of the throat when words and sentences are spoken, whether audibly or silently. The flexible liquid metal sensor includes at least one sensor unit, which includes a PDMS microtube filled with liquid metal. The two ends of the PDMS microtube are encapsulated with wires. The liquid metal is a pure metal or alloy with a melting point in the temperature range of -20℃ to 25℃. The sensor can sense pressure at a single point from 1mN to 1000mN. The upper limit of the frequency response of the sensor unit is >5000Hz.

2. The flexible liquid metal sensor according to claim 1, characterized in that: The PDMS microtube has an inner diameter of 10–400 μm, a length exceeding 50 cm, and a wall thickness of 5–600 μm.

3. A flexible liquid metal sensor according to claim 1, characterized in that, The PDMS microtubes are produced using a continuous extraction and curing process.

4. A flexible liquid metal sensor according to claim 1 or 3, characterized in that, The method for preparing the PDMS microtube includes: vertically passing a heated metal wire through a PDMS precursor, the PDMS precursor forming a PDMS coating around the metal wire, heating and curing the metal wire with the PDMS coating in a heating unit, the PDMS coating being cured in situ to form a tubular PDMS microtube, and separating the PDMS microtube from the metal wire by ultrasonic treatment to obtain the PDMS microtube.

5. A flexible liquid metal sensor according to claim 4, characterized in that: The temperature of the heated metal wire is 90°C to 110°C, and the residence time of the metal wire in the PDMS precursor is 2-6 minutes.

6. A flexible liquid metal sensor according to claim 4, characterized in that, The conditions for heat curing the metal wire with the PDMS coating in the heating unit are: 90℃~100℃, 3-5 minutes.

7. A flexible liquid metal sensor according to claim 1, characterized in that: It also includes a PDMS flexible substrate, on which multiple sensor units are disposed. The PDMS microtubes of the multiple sensor units are connected side by side to the surface of the PDMS flexible substrate to form a two-dimensional sensor.

8. A flexible liquid metal sensor according to claim 7, characterized in that, The fabrication method of the two-dimensional sensor includes: uniformly dispensing PDMS onto a silicon wafer and performing homogenization to obtain a flexible substrate blank with uniform thickness; performing plasma treatment on the flexible substrate blank to obtain a PDMS flexible substrate; performing plasma treatment on the PDMS microtubes to obtain treated microtubes; sequentially arranging the treated microtubes on the surface of the PDMS flexible substrate, and then heating to 90℃~110℃ and holding for 3-5 minutes, so that the treated microtubes adhere to the surface of the PDMS flexible substrate.

9. A voice recognition system based on a flexible liquid metal sensor according to any one of claims 1-8, characterized in that, include: The system includes a flexible liquid metal sensor, an acoustic and vibration signal sampling circuit board, a server, an audio output circuit board, and an artificial voice reproduction module. The server is equipped with a voice recognition module. A flexible liquid metal sensor is installed and fitted to the throat to acquire the acoustic resistance analog signal of the throat when speaking with or without sound, and sends the acoustic resistance analog signal to the acoustic signal sampling circuit board. The acoustic vibration signal sampling circuit board is used to receive the acoustic vibration resistance analog signal obtained by the flexible liquid metal sensor, convert the acoustic vibration resistance analog signal into a voltage analog signal, amplify, filter and sample the voltage analog signal by ADC to obtain an acoustic vibration voltage digital signal, and upload the acoustic vibration voltage digital signal to the server. The server receives the aforementioned acoustic vibration voltage digital signal and uses the acoustic vibration voltage digital signal as input to the speech recognition module; The speech recognition module provides a pre-trained deep learning model, receives the acoustic vibration voltage digital signal as input, performs data processing and feature extraction on the acoustic vibration voltage digital signal, and uses the time-frequency features obtained from the feature extraction as input to the pre-trained deep learning model; based on the input time-frequency features, the pre-trained deep learning model outputs the corresponding words and sentences expressed by the speech, and converts the words and sentences into audio digital signals; the speech recognition module sends the audio digital signals to the audio output circuit board; An audio output circuit board receives the audio digital signal, modulates the audio digital signal into multiple audio analog signals, amplifies the multiple audio analog signals to obtain amplified multiple audio analog signals, and outputs the amplified multiple audio analog signals to the artificial voice reproduction module. The artificial voice reproduction module is used to receive the amplified multi-channel audio analog signals output by the audio output circuit board, and apply the amplified multi-channel audio analog signals to various points of the artificial voice to achieve pronunciation.

10. The spoken and silent speech recognition system according to claim 9, characterized in that, The flexible liquid metal sensor is attached to the throat, including: Flexible liquid metal sensors include multiple sensor units and / or multiple two-dimensional sensors; Multiple sensor units are discretely placed at different parts of the throat to collect raw data of electrical signal changes from the sensor units. The raw data of electrical signal changes is subjected to spectrum analysis and noise reduction. The time-domain and frequency-domain signals of each sensor unit are analyzed, and the frequency distribution and amplitude of the frequency-domain signal and the distribution of the time-domain signal are observed. The spacing and position of the sensor units are adjusted, and the measurements and analyses are performed again to determine the sensor unit placement position that makes the difference between the time-domain signal and the frequency-domain signal most significant. Multiple two-dimensional sensors were attached to different positions in the throat. The change of signal χ along the x-axis of the microtube arrangement direction, dχ / dx, was analyzed. After adjusting the spacing and position of the two-dimensional sensors, the measurement and analysis were performed again to determine the two-dimensional sensor placement position that maximizes the absolute value of dχ / dx. The flexible liquid metal sensor is installed in the throat by combining the sensor unit arrangement position and / or the two-dimensional sensor arrangement position.

11. The spoken and silent speech recognition system according to claim 9, characterized in that, The acquisition of the acoustic vibration resistance simulation signal of the larynx during spoken and unspoken speech includes: having the subject utter words and sentences, and a flexible liquid metal sensor acquiring the acoustic vibration resistance simulation signal; The throat and / or head are moved, and a flexible liquid metal sensor acquires a non-voice vibration analog resistance signal; the throat and / or head movements include coughing, swallowing, yawning, and nodding; to exclude non-voice vibration characteristic resistance signals mixed in with the acoustic vibration resistance analog signal based on the characteristics of the non-voice vibration analog signal.

12. The spoken and silent speech recognition system according to claim 9, characterized in that: The trained deep learning model is obtained through the following method: A dataset relating the acoustic vibration voltage digital signal of the larynx to the vocabulary and sentences expressed in speech is constructed. The time-frequency features of the acoustic vibration voltage digital signal of the larynx are extracted. Based on this dataset, where the vocabulary and sentences are composed of English phonetic symbols or Chinese phonemes, the correspondence between the time-frequency features of the vibration voltage digital signal and the English phonetic symbols or Chinese phonemes in the vocabulary and sentences is obtained. A deep learning training model is then established based on this correspondence between the time-frequency features of the vibration voltage digital signal and the English phonetic symbols or Chinese phonemes. The dataset consisting of multiple acoustic, vibration, and voltage digital signals is divided into a training set and a test set. The deep learning training model is trained using k-fold cross-validation by combining the training set and the test set until the deep learning training model meets the accuracy requirements, thus obtaining a well-trained deep learning training model.

13. The spoken and silent speech recognition system according to claim 9, characterized in that: The acoustic and vibration signal sampling circuit board includes a first host computer communication module, a first microprocessor, a sensor excitation circuit, and an ADC sampling and filtering circuit; The sensor excitation circuit applies a voltage to the flexible liquid metal sensor. When a slight muscle movement in the throat and a mixed signal of sound vibration stimulate the flexible liquid metal sensor, the resistance of the flexible liquid metal sensor changes. This change in resistance is a simulated acoustic vibration resistance signal. The sensor excitation circuit converts this change in resistance into a change in voltage signal, which is a simulated acoustic vibration voltage signal. The simulated acoustic vibration voltage signal is then sent to the ADC sampling and filtering circuit. The ADC sampling and filtering circuit receives the acoustic vibration voltage analog signal, amplifies and filters the acoustic vibration voltage analog signal through the built-in amplification circuit and bandpass filter circuit of the ADC chip, and then converts the acoustic vibration voltage analog signal into a digital signal through the ADC chip sampling to obtain the acoustic vibration voltage digital signal, and sends the acoustic vibration voltage digital signal to the first microprocessor. The first microprocessor receives the acoustic vibration voltage digital signal and sends the acoustic vibration voltage digital signal to the host computer communication module; it is used to control the operation of the sensor excitation circuit, the ADC sampling and filtering circuit, and the host computer communication module. The host computer communication module receives the acoustic vibration voltage digital signal and sends the acoustic vibration voltage digital signal to the server.

14. The spoken and silent speech recognition system according to claim 9, characterized in that: The audio output circuit board includes a second host computer communication module, a second microprocessor, an audio decoder, and an audio excitation circuit. The second host computer communication module receives the audio digital signal obtained by the voice recognition module and transmits the audio digital signal to the second microprocessor through a transparent transmission mode. The second microprocessor receives the audio digital signal and sends the audio digital signal to the audio decoder; An audio decoder receives the digital audio signal, modulates the digital audio signal into multiple analog audio signals, and sends the multiple analog audio signals to an audio excitation circuit. The audio excitation circuit receives the multiple audio analog signals, amplifies the multiple audio analog signals to obtain amplified multiple audio analog signals, and outputs the amplified multiple audio analog signals to the artificial voice reproduction module.

15. The spoken and silent speech recognition system according to claim 9, characterized in that: Data transmission between the first host computer communication module and the server, and between the second host computer communication module and the server, is conducted via wireless or wired communication. The wireless communication methods include Bluetooth, WIFI, ZigBee, and cellular network technologies, while the wired communication methods include Ethernet, USB, and serial port connection technologies.

16. The spoken and silent speech recognition system according to claim 9, characterized in that, The extraction of the time-frequency characteristics of the acoustic vibration voltage digital signal of the throat includes: using the short-time Fourier transform time-varying spectrum or Mel time-varying frequency cepstral coefficients as the extracted time-frequency characteristics.

17. The spoken and silent speech recognition system according to claim 12, characterized in that: The deep learning training model is: ResNet50 network, EfficientFormerV2 and / or connection-based temporal classifier CTC.

18. The spoken and silent speech recognition system according to claim 17, characterized in that: When the deep learning training model is a ResNet50 network, the input of the ResNet50 network is the time-frequency feature of the acoustic-vibration-voltage digital signal, namely the short-time Fourier transform time-varying spectrum or time-varying Mel frequency cepstral coefficients; the short-time Fourier transform time-varying spectrum or time-varying Mel frequency cepstral coefficients are passed through 49 convolutional layers to obtain a feature vector of dimension M; the feature vector is input into a fully connected layer, and the output is a vector of dimension N. The vector of dimension N is passed through Softmax and the index with the highest probability is selected. The index with the highest probability indicates that the English phonetic symbol or Chinese phoneme corresponding to the index is the English phonetic symbol or Chinese phoneme at that time point.

19. The spoken and silent speech recognition system according to claim 17, characterized in that: When the deep learning training model is EfficientFormerV2, the input of EfficientFormerV2 is the time-frequency characteristics of the acoustic-vibration-voltage digital signal, namely the short-time Fourier transform time-varying spectrum or time-varying Mel frequency cepstral coefficients. After the short-time Fourier transform time-varying spectrum or time-varying Mel frequency cepstral coefficients are passed through a depthwise separable convolution and multi-head attention intermediate layer, a vector of dimension N is output through a fully connected layer. After the vector of dimension N is passed through Softmax, the index with the highest probability is selected. The index with the highest probability indicates that the English phonetic symbol or Chinese phoneme corresponding to the index is the English phonetic symbol or Chinese phoneme at that time point.

20. The spoken and silent speech recognition system according to claim 17, characterized in that: The ResNet50 or EfficientFormerV2 outputs the English phonetic symbols or Chinese phonemes for all time points within the time period to the temporal classifier CTC; finally, the CTC algorithm arranges the English phonetic symbols and / or Chinese phonemes into English / Chinese words or sentences based on the consecutive English phonetic symbols, phonemes, and English and Chinese vocabulary; the output English / Chinese words or sentences are used as the speech text output by the deep learning training model.