Voice signal testing method and system based on throat vibration and electroencephalogram electrical stimulation
By combining speech signal testing methods with laryngeal vibration and EEG stimulation, the PageRank algorithm and dual-modal SPCA technology are used to solve the challenges of existing equipment in distinguishing homophones and different meanings, achieving higher recognition accuracy and accuracy.
Patent Information
- Application Number
- CN202510480149.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing wearable devices face challenges in distinguishing sentences or words, and it is difficult to effectively solve the problem of homophones and different meanings in the pronunciation process of deaf and mute people.
A speech signal testing method based on laryngeal vibration and EEG stimulation is adopted. By obtaining the user's throat vibration signal and EEG signal, combined with the channel selection of PageRank algorithm, dual-modal SPCA processing and SVM classification model, the precise classification of speech signals is achieved.
It improves the recognition accuracy of voice information for deaf and dumb people, enhances the recognition accuracy in complex environments, and effectively solves the problem of homophones and different meanings.
Smart Images

Figure CN120011897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech classification, and in particular to a speech signal testing method and system based on laryngeal vibration and electroencephalographic electrical stimulation. Background Art
[0002] There are two ways for deaf-mute people to communicate with healthy people. One is to learn sign language, and the other is to use interactive devices for the deaf-mute. However, sign language is difficult to learn, and healthy people rarely use sign language, which makes them easy to forget. At the same time, sign language is not only difficult for healthy people to learn, but also for deaf-mute people.
[0003] Wearable devices have developed rapidly in interactive devices for the deaf and dumb, and their size has become smaller and smaller, making them more convenient for users to carry. Therefore, wearable devices have received widespread attention in interactive devices for the deaf and dumb. And due to the development of machine learning, the status of human-computer interaction has become very important. However, most of the existing wearable devices rely on a single throat sensor to capture throat vibration signals, which faces challenges in distinguishing sentences or words, and it is difficult to effectively solve the problem of homophones and different meanings in the pronunciation process of the deaf and dumb. Summary of the invention
[0004] The purpose of this application is to propose a speech signal testing method and system based on laryngeal vibration and EEG electrical stimulation in response to the above-mentioned technical problems.
[0005] In a first aspect, the present invention provides a method for testing a speech signal based on laryngeal vibration and electroencephalographic electrical stimulation, comprising the following steps:
[0006] Acquire the basic EEG signal of the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal;
[0007] In the acquisition speech execution phase, the user's laryngeal vibration signal and EEG signal are synchronously collected and data preprocessed respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed EEG signal; the preprocessed EEG signal is subjected to channel selection using the PageRank algorithm to obtain the EEG signal after channel selection;
[0008] The laryngeal vibration signal after preprocessing and the EEG signal after channel selection are subjected to feature extraction and dimensionality reduction processing respectively to obtain the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, and the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction are input into the dual-modal SPCA with L1 regularization to obtain sparse features;
[0009] A speech signal classification model based on SVM is constructed and trained to obtain a trained speech signal classification model, and the sparse features are input into the trained speech signal classification model to obtain a speech classification recognition result.
[0010] Preferably, judging whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal specifically includes:
[0011] The basic EEG signal is analyzed by using time-frequency analysis technology to obtain the average amplitude and spectrum energy of the basic EEG signal. The time-frequency analysis technology includes short-time Fourier transform and wavelet transform.
[0012] In response to determining that the average amplitude of the basic EEG signal in the theta-beta frequency band associated with language processing is less than the amplitude threshold, and the spectrum energy is less than the product of the normal reference value and the energy percentage threshold, an electrical stimulation control signal is generated during the speech preparation stage, and the frequency range of the EEG electrical stimulation corresponding to the electrical stimulation control signal is 20-40 Hz and the intensity range is 0.5-3 mA;
[0013] In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band related to language processing is greater than or equal to the amplitude threshold or the spectral energy is greater than or equal to the product of the normal reference value and the energy percentage threshold, no electrical stimulation control signal is generated in the speech preparation stage.
[0014] Preferably, basic EEG data is collected when the user does not generate a laryngeal vibration signal and is not subjected to EEG electrical stimulation. In the speech preparation stage, the time period is corresponding to the time threshold range before the user generates a laryngeal vibration signal; in the speech execution stage, the time period is corresponding to the time when the user generates a laryngeal vibration signal and is not subjected to EEG electrical stimulation.
[0015] Preferably, the data preprocessing methods of laryngeal vibration signals include pre-emphasis processing, gradual in-and-out processing and smoothing processing; the data preprocessing methods of EEG signals include bandpass filtering processing, processing using a dynamic notch filter and artifact removal, wherein the artifacts must meet the following requirements: the instantaneous fluctuation between adjacent sampling points exceeds ±50μV, the difference between adjacent peaks exceeds 200μV and the time span is within 200ms, the difference between the maximum amplitude and the minimum amplitude exceeds ±100μV, the fluctuation is less than 0.5μV within a continuous 100 millisecond time interval, or the power spectral density within the EEG electrical stimulation frequency bandwidth exceeds 3 times the standard deviation of the baseline value.
[0016] Preferably, the preprocessed EEG signal is subjected to channel selection using a PageRank algorithm to obtain the EEG signal after channel selection, which specifically includes:
[0017] Performing dimension transformation on the preprocessed EEG signal to obtain a transformed EEG signal;
[0018] Calculate the Pearson correlation coefficient between every two channels of the transformed EEG signal and construct a correlation matrix. Based on the percentile of the correlation matrix, divide the channels into and Pearson correlation coefficient between Mapped to the element in the i-th row and j-th column of the adjacency matrix , as shown below:
[0019] ;
[0020] in, , and They represent the 25th, 50th, and 75th percentiles of all Pearson correlation coefficients, respectively;
[0021] Each of the 32 channels of the transformed EEG signal is used as a node in the graph structure. The links between the nodes exist in the form of directed edges in the graph structure. The PageRank algorithm is used to calculate the link relationship between the nodes in the graph structure to determine the importance score of each node. In the process of calculating the importance score, if the clustering algorithm is used, the clustering coefficient of each node is calculated using the following formula:
[0022] ;
[0023] in, Indicates nodes, Indicates that after The number of triangles formed by a node and its neighboring nodes, Indicates The degree of the node, The degree of a node is equal to The in-degree and The sum of the out-degrees of the nodes;
[0024] If the degree algorithm is used, the degree centrality is calculated using the following formula:
[0025] ;
[0026] in, Indicates The degree centrality of the nodes, Indicates The in-degree of the node, The in-degree of a node is the first The sum of the elements of a column, Indicates The out-degree of the node, The out-degree of a node is the first The sum of the elements of a row;
[0027] No. The calculation formula for the importance score of a node is as follows:
[0028] ;
[0029] in, Indicates The importance score of each node, represents the normalization function, represents the algorithm weight, and N represents the total number of nodes;
[0030] The adaptive threshold is calculated based on the importance scores of all nodes, as shown in the following formula:
[0031] ;
[0032] in, represents the adaptive threshold, represents the similarity parameter;
[0033] Nodes whose importance scores are greater than or equal to the adaptive threshold are screened out and their corresponding channels are taken as important channels, and the EEG signals corresponding to the important channels are taken as the EEG signals after channel selection.
[0034] Preferably, the laryngeal vibration features extracted from the laryngeal vibration signal after preprocessing include Mel frequency cepstrum coefficients, root mean square energy and chromaticity features; the EEG features extracted from the EEG signal after channel selection include root mean square energy, energy of multiple frequency bands and average spectral entropy, wherein the multiple frequency bands include five frequency bands of delta, theta, alpha, beta and gamma; the dimensionality reduction method of the laryngeal vibration features and the EEG features adopts principal component analysis;
[0035] The objective function of the bimodal SPCA with L1 regularization is as follows:
[0036] ;
[0037] in, Represents the first projection matrix when the objective function is maximized and the second projection matrix , and satisfy =I and =I, I represents the identity matrix, T represents the transpose, ( ) represents the trace of the matrix, and Respectively and The L1 norm of , C represents the covariance matrix of the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, represents the regularization coefficient;
[0038] By optimizing the objective function of the dual-modal SPCA with L1 regularization, the first projection matrix and the second projection matrix are obtained. The first projection matrix and the second projection matrix are used to project the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features respectively to obtain the projected laryngeal vibration features and the projected EEG features. The projected laryngeal vibration features and the projected EEG features are concatenated to obtain sparse features.
[0039] In a second aspect, the present invention provides a speech signal testing system based on laryngeal vibration and electroencephalographic electrical stimulation, comprising:
[0040] The electrical stimulation control signal generation module is configured to obtain the basic EEG signal collected from the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal;
[0041] The data processing module is configured to obtain the laryngeal vibration signal and the EEG signal of the user synchronously collected during the speech execution phase and perform data preprocessing respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed EEG signal; and select the channel of the preprocessed EEG signal using the PageRank algorithm to obtain the EEG signal after the channel selection;
[0042] The feature extraction and dimensionality reduction module is configured to perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the channel-selected EEG signal, respectively, to obtain the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, and input the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction into the dual-modal SPCA with L1 regularization to obtain sparse features;
[0043] The classification and recognition module is configured to construct and train a speech signal classification model based on SVM to obtain a trained speech signal classification model, input sparse features into the trained speech signal classification model, and obtain a speech classification and recognition result.
[0044] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0045] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0046] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] (1) The speech signal testing method based on laryngeal vibration and EEG stimulation proposed in the present invention can select EEG channels based on the EEG signal channel selection method based on the PageRank algorithm, which can better represent speech-specific features. In the subsequent data processing and feature extraction process, the features extracted based on these key channels can more accurately reflect the EEG patterns related to speech of the deaf-mute, making the features more representative and distinguishable, and effectively improving the recognition accuracy of the speech information of the deaf-mute.
[0049] (2) The sparse features obtained by the speech signal testing method based on laryngeal vibration and EEG electrical stimulation proposed in the present invention after L1 regularized dual-modal SPCA processing have better anti-interference ability, can highlight key features, and suppress the noise characteristics caused by environmental interference, so as to maintain a high recognition accuracy in complex environments and improve the generalization ability. With the help of dual-modal SPCA and other technologies, the intrinsic connection between laryngeal vibration signals and EEG signals is deeply explored, and speech features are comprehensively extracted to ensure the accurate grasp of various types of speech information.
[0050] (3) The speech signal testing method based on laryngeal vibration and EEG stimulation proposed in the present invention can effectively realize the simultaneous acquisition of laryngeal vibration signals and EEG signals when the deaf-mute speaks. By supplementing the laryngeal vibration signals with EEG signals, the problem of homonyms but different meanings is effectively solved. The output of intuitive results is used to translate the speech information expressed by the deaf-mute, thereby achieving more accurate recognition of the sensor signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0052] Figure 1 A flow chart of a speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation according to an embodiment of the present application;
[0053] Figure 2 A schematic diagram of the data collection and interaction flow of the speech signal testing method based on laryngeal vibration and EEG electrical stimulation according to an embodiment of the present application;
[0054] Figure 3 A schematic diagram of a laryngeal vibration signal of the word "airport" in a speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation according to an embodiment of the present application;
[0055] Figure 4 A schematic diagram of a laryngeal vibration signal of the word "down" in a speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation according to an embodiment of the present application;
[0056] Figure 5 A schematic diagram of a laryngeal vibration signal of the easily confused word "bough" in a speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation according to an embodiment of the present application;
[0057] Figure 6 A schematic diagram of a laryngeal vibration signal of an easily confused word "bow" in a speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation according to an embodiment of the present application;
[0058] Figure 7 A schematic diagram of an EEG signal of the easily confused word bough in a speech signal testing method based on laryngeal vibration and EEG electrical stimulation according to an embodiment of the present application;
[0059] Figure 8 A schematic diagram of an EEG signal of an easily confused word "bow" in a speech signal testing method based on laryngeal vibration and EEG electrical stimulation according to an embodiment of the present application;
[0060] Fig. 9 A schematic diagram of a confusion matrix of a single laryngeal vibration signal of easily confused words in a speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation according to an embodiment of the present application;
[0061] Fig.10 A schematic diagram of a confusion matrix of a mixed signal consisting of laryngeal vibration signals and EEG signals of easily confused words in a speech signal testing method based on laryngeal vibration and EEG electrical stimulation according to an embodiment of the present application;
[0062] Fig.11 A schematic diagram of a speech signal testing system based on laryngeal vibration and EEG electrical stimulation according to an embodiment of the present application;
[0063] Fig.12 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] Figure 1 A speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation provided by an embodiment of the present application is shown, comprising the following steps:
[0066] S1, acquiring the basic EEG signal of the user, determining whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and applying EEG electrical stimulation to the user according to the electrical stimulation control signal.
[0067] In a specific embodiment, determining whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal specifically includes:
[0068] The basic EEG signal is analyzed by using time-frequency analysis technology to obtain the average amplitude and spectrum energy of the basic EEG signal. The time-frequency analysis technology includes short-time Fourier transform and wavelet transform.
[0069] In response to determining that the average amplitude of the basic EEG signal in the theta-beta frequency band associated with language processing is less than the amplitude threshold, and the spectrum energy is less than the product of the normal reference value and the energy percentage threshold, an electrical stimulation control signal is generated during the speech preparation stage, and the frequency range of the EEG electrical stimulation corresponding to the electrical stimulation control signal is 20-40 Hz and the intensity range is 0.5-3 mA;
[0070] In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band related to language processing is greater than or equal to the amplitude threshold or the spectral energy is greater than or equal to the product of the normal reference value and the energy percentage threshold, no electrical stimulation control signal is generated in the speech preparation stage.
[0071] In a specific embodiment, basic EEG data is collected when the user does not generate a laryngeal vibration signal and is not subjected to EEG electrical stimulation. In the speech preparation stage, the basic EEG data is collected when the user generates a laryngeal vibration signal and is subjected to a time period corresponding to a time threshold range before the user generates a laryngeal vibration signal. In the speech execution stage, the basic EEG data is collected when the user generates a laryngeal vibration signal and is not subjected to EEG electrical stimulation.
[0072] Specifically, refer to Figure 2In the embodiments of the present application, an EEG cap can be used to collect EEG signals. The EEG cap integrates an electrical stimulation module. Due to the long-term lack of language output training in the deaf-mute population, the intensity of neural activity in language-related brain areas (such as Broca area and Wernicke area) is significantly lower than that in healthy people. Studies have shown that the power of the gamma band decreases by about 40-60%. EEG electrical stimulation (such as tACS / tDCS) can compensate for the signal attenuation problem by regulating the membrane potential of neurons and enhancing the synchronized activity of the target brain area. When the user is ready to speak, the basic EEG signal is first collected through the EEG cap at a sampling rate of 256Hz with 32 channels of the international 10 / 20 system, and the signal characteristics and stability are analyzed using time-frequency analysis technology (including short-time Fourier transform (STFT), wavelet transform) methods, and the average amplitude and spectral energy of the basic EEG signal are obtained.
[0073] Clinical studies have proven that if the following conditions are detected:
[0074] 1. The average amplitude of EEG signals in the θ-β frequency band (4-30Hz) related to language processing is less than 30μV, and the θ wave (4-7Hz) is related to language comprehension;
[0075] 2. The spectral energy of EEG signals in the θ-β frequency band related to language processing is reduced by more than 40% compared with the normal reference value.
[0076] Then, an electrical stimulation control signal is automatically generated and sent to the electrical stimulation function module to generate EEG stimulation. EEG stimulation is only applied in the speech preparation stage, and the parameter information of the electrical stimulation is recorded at the same time, including the stimulation frequency, intensity, etc. In one embodiment, the speech preparation stage is 300-500ms before the laryngeal vibration signal is triggered, the stimulation frequency ranges from 20-40Hz, and the intensity is 0.5-3mA. In other embodiments, the frequency and intensity can also be adjusted according to the electrode area, and can be adjusted automatically according to the signal characteristics and stability. There is no EEG stimulation in the EEG signal in the speech execution stage. Subsequently, the laryngeal vibration signal and the EEG signal are collected synchronously. In addition, during the collection process, different collection test scenarios can be set, such as different environments, different users, etc., to comprehensively evaluate the system performance.
[0077] A laryngeal pressure sensor is used to collect laryngeal vibration signals. In one embodiment, the laryngeal pressure sensor adopts a hydrogel pressure sensor and is connected to an electrochemical workstation (PalmSens4) to collect laryngeal vibration signals using an open circuit potentiometry method.
[0078] The user wore an EEG cap (Emotive) to record EEG signals and a hydrogel pressure sensor on the larynx to record laryngeal (AT) vibration signals through an electrochemical workstation (PalmSens 4), while looking at the word pictures displayed on the computer screen for visual stimulation. EEG signals and laryngeal vibration signals were synchronously acquired with a unified timestamp to ensure the time alignment of multimodal data. In each experimental cycle, the user read the displayed words within 3 seconds, with more than 10 seconds between words for rest, and an interval of about 5 minutes after each experimental session. EEG signals were acquired from 32 channels at a sampling rate of 256 Hz (according to the international 10 / 20 system), while laryngeal vibration signals were acquired from the larynx (thyroid cartilage) at a sampling rate of 250 Hz. The Fz channel was used as the ground and the CP2 channel as the reference.
[0079] S2, acquiring the laryngeal vibration signal and EEG signal of the user synchronously collected during the speech execution phase and performing data preprocessing respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed EEG signal; performing channel selection on the preprocessed EEG signal using the PageRank algorithm to obtain the EEG signal after channel selection.
[0080] In a specific embodiment, the data preprocessing methods of laryngeal vibration signals include pre-emphasis processing, gradual in-and-out processing, and smoothing processing; the data preprocessing methods of EEG signals include bandpass filtering, processing using a dynamic notch filter, and artifact removal, wherein the artifacts must meet the following requirements: the instantaneous fluctuation between adjacent sampling points exceeds ±50μV, the difference between adjacent peaks exceeds 200μV and the time span is within 200ms, the difference between the maximum amplitude and the minimum amplitude exceeds ±100μV, the fluctuation is less than 0.5μV within a continuous 100 millisecond time interval, or the power spectral density within the EEG electrical stimulation frequency bandwidth exceeds 3 times the standard deviation of the baseline value.
[0081] Specifically, the preprocessing process of multimodal data is as follows:
[0082] The laryngeal signal is pre-emphasized to keep the low-frequency part of the signal unchanged and enhance the high-frequency part of the signal. The pre-emphasized laryngeal signal is gradually stepped in and out to remove the sudden change caused by the truncated signal after pre-emphasis. Then the signal is smoothed to reduce the noise in the data, so as to enhance the signal quality and prevent abnormal signal disturbance.
[0083] The EEG signal is filtered through a bandpass filter of 0.5-50Hz to remove the interference of the mains, and the frequency interference of EEG stimulation is eliminated in real time by combining with a dynamic notch filter. Then, artifact removal (noise removal, removal of electrooculography, electromyography, and electrical stimulation residue) is performed. In EEG signal processing, if any of the following five artifact identification conditions is met, it will be marked as an artifact, as follows:
[0084] 1. The instantaneous fluctuation between adjacent sampling points exceeds ±50μV;
[0085] 2. The difference between adjacent peaks exceeds 200μV and the time span is within 200ms;
[0086] 3. The difference between the maximum amplitude and the minimum amplitude exceeds ±100μV;
[0087] 4. Fluctuation less than 0.5μV within a continuous 100 millisecond time interval;
[0088] 5. Within the stimulation frequency bandwidth, the power spectral density exceeds the baseline value by 3 standard deviations.
[0089] In a specific embodiment, the preprocessed EEG signal is subjected to channel selection using the PageRank algorithm to obtain the EEG signal after channel selection, which specifically includes:
[0090] Performing dimension transformation on the preprocessed EEG signal to obtain a transformed EEG signal;
[0091] Calculate the Pearson correlation coefficient between every two channels of the transformed EEG signal and construct a correlation matrix. Based on the percentile of the correlation matrix, divide the channels into and Pearson correlation coefficient between Mapped to the element in the i-th row and j-th column of the adjacency matrix , as shown below:
[0092] ;
[0093] in, , and They represent the 25th, 50th, and 75th percentiles of all Pearson correlation coefficients, respectively;
[0094] Each of the 32 channels of the transformed EEG signal is used as a node in the graph structure. The links between the nodes exist in the form of directed edges in the graph structure. The PageRank algorithm is used to calculate the link relationship between the nodes in the graph structure to determine the importance score of each node. In the process of calculating the importance score, if the clustering algorithm is used, the clustering coefficient of each node is calculated using the following formula:
[0095] ;
[0096] in, Indicates nodes, Indicates that after The number of triangles formed by a node and its neighboring nodes, Indicates The degree of the node, The degree of a node is equal to The in-degree and The sum of the out-degrees of the nodes;
[0097] If the degree algorithm is used, the degree centrality is calculated using the following formula:
[0098] ;
[0099] in, Indicates The degree centrality of the nodes, Indicates The in-degree of the node, The in-degree of a node is the first The sum of the elements of a column, Indicates The out-degree of the node, The out-degree of a node is the first The sum of the elements of a row;
[0100] No. The calculation formula for the importance score of a node is as follows:
[0101] ;
[0102] in, Indicates The importance score of each node, represents the normalization function, represents the algorithm weight, and N represents the total number of nodes;
[0103] The adaptive threshold is calculated based on the importance scores of all nodes, as shown in the following formula:
[0104] ;
[0105] in, represents the adaptive threshold, represents the similarity parameter;
[0106] Nodes whose importance scores are greater than or equal to the adaptive threshold are screened out and their corresponding channels are taken as important channels, and the EEG signals corresponding to the important channels are taken as the EEG signals after channel selection.
[0107] Specifically, the PageRank algorithm is used to select channels for the preprocessed EEG signal. Specifically, the PageRank algorithm is based on a link analysis algorithm of a directed graph. The importance of each node is determined by iteratively calculating the link relationship between nodes. The algorithm evaluates their relative importance by analyzing the link (directed edge) relationship between nodes. Specifically, the PageRank algorithm regards links as votes. When a node links to another node, it is equivalent to voting for the linked node. Different nodes have different weights for votes on linked nodes, and the weight depends on the importance of the voting node, that is, the importance score. The embodiment of the present application regards 32 channels as nodes in a directed graph. The specific steps of the PageRank algorithm are as follows:
[0108] 1. First, transform the dimension of the preprocessed EEG signal as shown in the following formula: (trail, channel, seq_len) −> (channel, trail × seq_len), where trail represents the number of data, channel represents the channel, and seq_len represents the length of each data. Calculate the Pearson correlation coefficient between any two channels;
[0109] 2. Divide the elements in the adjacency matrix by the 25th, 50th, and 75th percentiles. The elements in the interval (0, 25) are assigned 0, the elements in the interval (25, 50) are assigned 1, the elements in the interval (50, 75) are assigned 2, and the elements in the interval (75, 100) are assigned 3. Finally, we get a 32×32 adjacency matrix A, in which each element Indicates channel arrive The larger the value, the stronger the correlation.
[0110] 3. Calculate the clustering coefficient and degree centrality respectively. Use the clustering algorithm to calculate the clustering coefficient. The goal of the clustering coefficient is to measure the clustering tightness of a group. Use the degree algorithm to calculate the degree centrality. Degree centrality counts the number of edges directly connected to a node, including out-degree and in-degree, which can be simply understood as the size of a node's access opportunity. The out-degree of a node is The sum of the weights of all directed edges starting from the node, that is, the first node in the adjacency matrix The sum of the elements of the row; The in-degree of a node is The sum of the weights of all directed edges with nodes as endpoints, that is, the first node in the adjacency matrix The degree centrality is the sum of the out-degree and in-degree. The two coefficients are reconciled to get the importance score of each channel (node). The larger the value, the more attention is paid to the overall situation of the network. When it is 1, it is a clustering algorithm. The smaller the network, the more attention it pays to the local part of the network. When it is 0, it is the degree algorithm.
[0111] 4. Calculate an adaptive threshold according to the importance score, and filter out nodes in the corresponding network through the adaptive threshold. In one embodiment, 16 important channels in the EEG signal are finally filtered out, and the shape of the filtered EEG signal is (channel: 16, seq_len: 750).
[0112] S3, respectively perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the channel-selected EEG signal to obtain the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features, and input the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features into the dual-modal SPCA with L1 regularization to obtain sparse features.
[0113] In a specific embodiment, the laryngeal vibration features extracted from the laryngeal vibration signal after preprocessing include Mel frequency cepstral coefficients, root mean square energy and chromaticity features; the EEG features extracted from the EEG signal after channel selection include root mean square energy, energy of multiple frequency bands and average spectrum entropy, wherein the multiple frequency bands include five frequency bands of delta, theta, alpha, beta and gamma; the dimension reduction method of the laryngeal vibration features and the EEG features adopts the principal component analysis method;
[0114] Applying L1 regularization to the objective function of the dual-modal SPCA with L1 regularization, the objective function of the dual-modal SPCA with L1 regularization after applying L1 regularization is obtained, as shown in the following formula:
[0115] ;
[0116] in, Represents the first projection matrix when the objective function is maximized and the second projection matrix , and satisfy =I and =I, I represents the identity matrix, T represents the transpose, ( ) represents the trace of the matrix, and Respectively and The L1 norm of , C represents the covariance matrix of the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, represents the regularization coefficient;
[0117] By optimizing the objective function of the dual-modal SPCA with L1 regularization, the first projection matrix and the second projection matrix are obtained, and the first projection matrix and the second projection matrix are used to project the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features respectively to obtain the projected laryngeal vibration features and the projected EEG features, and the projected laryngeal vibration features and the projected EEG features are spliced to obtain sparse features.
[0118] Specifically, the laryngeal vibration signal after preprocessing and the EEG signal after channel selection are subjected to feature extraction and dimension reduction processing to reduce the amount of calculation without reducing the accuracy, improve the recognition speed, and obtain multimodal features. The extracted laryngeal vibration features include: Mel frequency cepstral coefficients (MFCC), root mean square energy (RMS) and chromaticity features (CHROMA) in the laryngeal vibration signal; the extracted EEG features include the root mean square energy (RMS) of the EEG signal, the energy and average spectral entropy of the five frequency bands of delta, theta, alpha, beta, and gamma.
[0119] The principal component analysis method was used for the laryngeal vibration features and EEG features respectively, and the principal components with a variance contribution rate of more than 95% were retained for dimensionality reduction to obtain the laryngeal vibration features and EEG features after dimensionality reduction; then the bimodal SPCA (Bimodal-sSPCA) with LI regularization was used for processing, and finally 64-dimensional sparse features with speech specificity were obtained. Bimodal SPCA can simultaneously take into account the correlation between the two different modal data, laryngeal vibration signals and EEG signals. It constructs a joint objective function to find the projection direction that can enable the two modal data to maintain their respective characteristics in a low-dimensional space and maximize the correlation between the modalities. Applying L1 regularization in the objective function of bimodal SPCA can obtain a sparse projection matrix and feature representation. Sparse features have better interpretability and can highlight features that contribute more to speech specificity while suppressing noise and irrelevant features. This helps to improve the generalization ability and computational efficiency of the model. Ordinary PCA does not have this sparsity constraint, and the features obtained may contain more redundant information.
[0120] The objective function of the original bimodal SPCA is:
[0121] ;
[0122] in, The covariance matrix representing the two-modal features describes the second-order statistical relationship between the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features, which is obtained by calculating the expectation.
[0123] The objective function of the bimodal SPCA with L1 regularization is:
[0124] ;
[0125] pass =I and =I constraint, so that the first projection matrix and the second projection matrix is an orthogonal matrix.
[0126] The objective function of the dual-modal SPCA with L1 regularization is optimized by combining alternating optimization and Lagrange multiplier method. The specific process is as follows:
[0127] Fixed second projection matrix , the objective function of the dual-modal SPCA with L1 regularization is about the first projection matrix Derivative (considering constraints and L1 regularization). Let the Lagrangian function be:
[0128] L( ) = ;
[0129] in is the Lagrange multiplier matrix for L( ) Take the derivative and set it to 0, and finally find the first projection matrix ; Similarly, fix the first projection matrix , update the second projection matrix .
[0130] Then, the first projection matrix is used to project the reduced-dimensional laryngeal vibration features to obtain the projected laryngeal vibration features; the second projection matrix is used to project the reduced-dimensional EEG features to obtain the projected EEG features, and then the projected laryngeal vibration features are concatenated with the projected EEG features to obtain the final 64-dimensional speech-specific sparse features.
[0131] S4, constructing and training a speech signal classification model based on SVM to obtain a trained speech signal classification model, inputting the sparse features into the trained speech signal classification model to obtain a speech classification recognition result.
[0132] Specifically, the embodiment of the present application adopts a support vector machine (SVM) classifier as a speech signal classification model, divides sparse features as input signals into training data and test data, trains the support vector machine (SVM) classifier, obtains a trained speech signal classification model, and inputs sparse features into the trained speech signal classification model to obtain speech classification and recognition results. At the same time, the performance under different test conditions is evaluated, such as under different environments, different subjects, etc., by calculating indicators such as recognition accuracy, recall rate, and F1 value to comprehensively measure the performance of the system. As shown in Table 1, the recognition accuracy, F1 value, recall rate, and precision rate of the ten easily confused words "Meet, Meat, Bow, Bough, Hear, Here, Costume, Custom, Implicit, Explicit" by two different users are compared. It can be seen that the generalization ability of the entire system is strong, and the accuracy rate does not fluctuate much for different users.
[0133] Table 1
[0134]
[0135] Specifically, when a person speaks, the larynx vibrates significantly. Different speech contents have different larynx vibration signals. For example, when reading the words "airport" and "down", Figure 3 and 4 As shown in the figure, the waveforms of the two words are obviously different. However, when it comes to easily confused words, such as Figure 5 and 6 As shown in the figure, due to the similar pronunciation of "bough" and "bow", the laryngeal vibration signals are very similar and difficult to distinguish. Compared with laryngeal signals, EEG signals have great advantages in identifying easily confused words. They contain a lot of information, such as Figure 7 and 8 As shown in the figure, the waveforms of the Bough and Bow EEG signals are obviously different, which is of great significance for the recognition and realization of speech signals. Therefore, the speech signal testing method based on laryngeal vibration and EEG electrical stimulation proposed in the embodiment of the present application can overcome the problem that a single information source cannot identify easily confused words and has a low recognition accuracy. The accuracy in the easily confused word recognition and classification task is greatly improved from 69.33% of a single laryngeal signal to 91.50% of a fused signal. Fig. 9 and 10 shown.
[0136] Similarly, a robust recognition rate of 96.26% (Mult) was obtained for words or phrases in other groups. Tables 2 and 3 are for words in other groups, and Table 4 is for the recognition accuracy of words in other groups.
[0137] Table 2:
[0138]
[0139] Table 3:
[0140]
[0141] Table 4:
[0142]
[0143] Further references Fig.11 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a speech signal testing system based on laryngeal vibration and EEG electrical stimulation. Figure 1 Corresponding to the method embodiment shown, the system can be specifically applied to various electronic devices.
[0144] The embodiment of the present application provides a speech signal testing system based on laryngeal vibration and electroencephalographic electrical stimulation, comprising:
[0145] The electrical stimulation control signal generating module 1 is configured to obtain the basic EEG signal collected from the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal;
[0146] The data processing module 2 is configured to obtain the laryngeal vibration signal and the EEG signal of the user synchronously collected during the speech execution phase and perform data preprocessing respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed EEG signal; and select the channel of the preprocessed EEG signal using the PageRank algorithm to obtain the EEG signal after the channel selection;
[0147] The feature extraction and dimension reduction module 3 is configured to perform feature extraction and dimension reduction processing on the preprocessed laryngeal vibration signal and the channel-selected EEG signal, respectively, to obtain the laryngeal vibration features after dimension reduction and the EEG features after dimension reduction, and input the laryngeal vibration features after dimension reduction and the EEG features after dimension reduction into the dual-modal SPCA with L1 regularization to obtain sparse features;
[0148] The classification and recognition module 4 is configured to construct and train a speech signal classification model based on SVM to obtain a trained speech signal classification model, input sparse features into the trained speech signal classification model, and obtain a speech classification and recognition result.
[0149] Fig.12 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Fig.12As shown, the electronic device of this embodiment includes: a processor 1201 and a memory 1202; wherein the memory 1202 is used to store computer-executable instructions; the processor 1201 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0150] Optionally, the memory 1202 may be independent or integrated with the processor 1201 .
[0151] When the memory 1202 is independently provided, the electronic device further includes a bus 1203 for connecting the memory 1202 and the processor 1201 .
[0152] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 1201 executes the computer execution instructions, the above method is implemented.
[0153] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 1201, the above method is implemented.
[0154] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0155] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.
[0156] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The unit formed by the above modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0157] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 1201 to perform some steps of the methods of various embodiments of the present application.
[0158] It should be understood that the processor 1201 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor 1201 may be any conventional processor 1201, etc. The steps of the method disclosed in the invention may be directly embodied in the hardware processor 1201 for execution, or may be executed by a combination of hardware and software modules in the processor 1201.
[0159] The memory 1202 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.
[0160] The bus 1203 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 1203 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 1203 in the drawings of the present application is not limited to only one bus 1203 or one type of bus 1203.
[0161] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0162] An exemplary storage medium is coupled to the processor 1201, so that the processor 1201 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 1201. The processor 1201 and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor 1201 and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0163] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech signal testing method based on laryngeal vibration and electroencephalographic electrical stimulation, characterized in that: The following steps are involved: Acquire the basic EEG signal of the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal; Acquiring the laryngeal vibration signal and the electroencephalogram signal of the user synchronously during the speech execution phase and performing data preprocessing respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed electroencephalogram signal; Using PageRank algorithm to select channels for the preprocessed EEG signals to obtain EEG signals after channel selection; Respectively performing feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the channel-selected EEG signal to obtain reduced-dimensional laryngeal vibration features and reduced-dimensional EEG features, and inputting the reduced-dimensional laryngeal vibration features and reduced-dimensional EEG features into a dual-modal SPCA with L1 regularization to obtain sparse features; A speech signal classification model based on SVM is constructed and trained to obtain a trained speech signal classification model, and the sparse features are input into the trained speech signal classification model to obtain a speech classification recognition result.
2. The voice signal testing method based on laryngeal vibration and brain electrical stimulation according to claim 1, characterized in that: Determining whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal specifically includes: Analyzing the basic EEG signal using a time-frequency analysis technique to obtain an average amplitude and spectrum energy of the basic EEG signal, wherein the time-frequency analysis technique includes short-time Fourier transform and wavelet transform; In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band associated with language processing is less than the amplitude threshold, and the spectrum energy is less than the product of the normal reference value and the energy percentage threshold, generating the electrical stimulation control signal in the speech preparation stage, the frequency range of the EEG electrical stimulation corresponding to the electrical stimulation control signal being 20-40 Hz and the intensity range being 0.5-3 mA; In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band related to language processing is greater than or equal to the amplitude threshold or the spectral energy is greater than or equal to the product of the normal reference value and the energy percentage threshold, the electrical stimulation control signal is not generated in the speech preparation stage.
3. The voice signal testing method based on laryngeal vibration and brain electrical stimulation according to claim 1, characterized in that: The basic EEG data is collected when the user does not generate a laryngeal vibration signal and is not subjected to EEG electrical stimulation. In the speech preparation stage, it is a time period corresponding to a time threshold range before the user generates a laryngeal vibration signal; in the speech execution stage, it is a time period corresponding to a user generating a laryngeal vibration signal and not subjected to EEG electrical stimulation.
4. The voice signal testing method based on laryngeal vibration and brain electrical stimulation according to claim 1, characterized in that: The data preprocessing methods of the laryngeal vibration signal include pre-emphasis processing, gradual in-and-out processing and smoothing processing; the data preprocessing methods of the EEG signal include bandpass filtering processing, processing using a dynamic notch filter and artifact removal, wherein the artifact must meet the following requirements: the instantaneous fluctuation between adjacent sampling points exceeds ±50μV, the difference between adjacent peaks exceeds 200μV and the time span is within 200ms, the difference between the maximum amplitude and the minimum amplitude exceeds ±100μV, the fluctuation is less than 0.5μV within a continuous 100 millisecond time interval, or the power spectral density within the EEG electrical stimulation frequency bandwidth exceeds 3 times the standard deviation of the baseline value.
5. The voice signal testing method based on laryngeal vibration and brain electrical stimulation according to claim 1, characterized in that: The preprocessed EEG signal is subjected to channel selection using the PageRank algorithm to obtain the EEG signal after channel selection, specifically comprising: Performing dimension transformation on the preprocessed EEG signal to obtain a transformed EEG signal; Calculate the Pearson correlation coefficient between every two channels of the transformed EEG signal and construct a correlation matrix. Based on the percentiles of the correlation matrix, separate the channels into and Pearson correlation coefficient between Mapped to the element in the i-th row and j-th column of the adjacency matrix , as shown below: ; in, , and They represent the 25th, 50th, and 75th percentiles of all Pearson correlation coefficients, respectively; Each of the 32 channels of the transformed EEG signal is used as a node in the graph structure. The links between the nodes exist in the form of directed edges in the graph structure. The PageRank algorithm is used to calculate the link relationship between the nodes in the graph structure to determine the importance score of each node. In the calculation process of the importance score, if a clustering algorithm is used, the clustering coefficient of each node is calculated using the following formula: ; in, Indicates nodes, Indicates that after The number of triangles formed by a node and its neighboring nodes, Indicates The degree of the node, The degree of a node is equal to The in-degree and The sum of the out-degrees of the nodes; If the degree algorithm is used, the degree centrality is calculated using the following formula: ; in, Indicates The degree centrality of the nodes, Indicates The in-degree of the node, The in-degree of the node is the first The sum of the elements of a column, Indicates The out-degree of the node, The out-degree of the node is the first The sum of the elements of a row; No. The calculation formula for the importance score of a node is as follows: ; in, Indicates The importance score of each node, represents the normalization function, represents the algorithm weight, and N represents the total number of nodes; The adaptive threshold is calculated based on the importance scores of all nodes, as shown in the following formula: ; in, represents the adaptive threshold, represents the similarity parameter; Nodes whose importance scores are greater than or equal to the adaptive threshold are screened out and their corresponding channels are used as important channels, and the EEG signals corresponding to the important channels are used as the EEG signals after channel selection.
6. The voice signal testing method based on laryngeal vibration and brain electrical stimulation according to claim 1, characterized in that: The laryngeal vibration features extracted from the preprocessed laryngeal vibration signal include Mel frequency cepstral coefficients, root mean square energy and chromaticity features; The EEG features extracted from the EEG signal after the channel selection include root mean square energy, energy of multiple frequency bands and average spectral entropy, wherein the multiple frequency bands include five frequency bands of delta, theta, alpha, beta and gamma; the dimension reduction method of the laryngeal vibration features and the EEG features adopts the principal component analysis method; The objective function of the dual-modal SPCA with L1 regularization is as follows: ; in, Represents the first projection matrix when the objective function is maximized and the second projection matrix , and satisfy =I and =I, I represents the identity matrix, T represents the transpose, ( ) represents the trace of the matrix, and Respectively and The L1 norm of , C represents the covariance matrix of the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, represents the regularization coefficient; By optimizing the objective function of the dual-modal SPCA with L1 regularization, the first projection matrix and the second projection matrix are obtained, and the first projection matrix and the second projection matrix are used to project the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features respectively to obtain the projected laryngeal vibration features and the projected EEG features, and the projected laryngeal vibration features and the projected EEG features are spliced to obtain sparse features.
7. A speech signal testing system based on laryngeal vibration and brain electrical stimulation, characterized in that: include: The electrical stimulation control signal generating module is configured to obtain the basic EEG signal collected from the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal; The data processing module is configured to obtain the laryngeal vibration signal and the electroencephalogram signal of the user synchronously collected during the speech execution phase and perform data preprocessing respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed electroencephalogram signal; Using PageRank algorithm to select channels for the preprocessed EEG signals to obtain EEG signals after channel selection; The feature extraction and dimensionality reduction module is configured to perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the channel-selected EEG signal, respectively, to obtain the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, and input the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction into the dual-modal SPCA with L1 regularization to obtain sparse features; The classification and recognition module is configured to construct and train a speech signal classification model based on SVM to obtain a trained speech signal classification model, input the sparse features into the trained speech signal classification model, and obtain a speech classification and recognition result.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Mild cognitive impairment auxiliary diagnosis system and method based on brain network multi-feature analysis
CN111009324A
Silent speech recognition method and system
CN111723717A
Voice signal processing method and related device therefor
CN114072875A
Fatigue driving identification method based on multi-channel weighted multi-scale permutation entropy
CN115346196A
Voice decoding recognition method and system for multi-mode throat vibration signal and lip moving point data
CN119068870A