Speech Signal Testing Method and System Based on Laryngeal Vibration and Electroencephalogram Electrical Stimulation
By combining laryngeal vibration and electroencephalopathy, the PageRank algorithm and L1 regularized dual-modal SPCA are used for feature extraction to construct an SVM model, which solves the problem of homophones and different meanings in deaf and mute communication, and improves the accuracy and accuracy of speech recognition.
Patent Information
- Application Number
- CN202510480149.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing wearable devices rely on a single throat sensor to effectively solve the problem of homophones and different meanings in the pronunciation process of deaf and dumb people, and the learning of sign language is difficult, and there are fewer occasions for people to use sign language, which leads to communication difficulties.
Combining laryngeal vibration and EEG stimulation, the speech preparation stage is determined by obtaining basic EEG signals to generate electrical stimulation control signals, and the laryngeal vibration signals and EEG signals are synchronized. The PageRank algorithm is used for channel selection, and the L1 regularized dual-modal SPCA is used for feature extraction and dimensionality reduction, and an SVM model is built for speech signal classification.
It improves the accuracy and accuracy of voice information recognition for deaf and dumb people, can maintain a high recognition rate in complex environments, effectively solve the problems of homophones and different meanings, and realizes accurate translation of voice information for deaf and dumb people.
Smart Images

Figure CN120011897B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech classification, and particularly to a method and system for testing speech signals based on laryngeal vibration and electroencephalogram electrical stimulation. Background Art
[0002] There are two ways for deaf-mutes to communicate with normal people. One is to communicate by learning sign language, and the other is to use a deaf-mute interaction device. However, sign language is difficult to learn, and normal people use sign language in fewer situations, resulting in easy forgetting. At the same time, sign language is not only difficult for normal people to learn, but also difficult for deaf-mutes to learn.
[0003] In deaf-mute interaction devices, the development of wearable devices is rapid, and the size has become smaller and smaller, which is convenient for users to carry. Therefore, wearable devices in deaf-mute interaction devices have received extensive attention. And due to the development of machine learning, the status of human-computer interaction has become crucial. However, most of the existing wearable devices rely on a single laryngeal sensor to capture laryngeal vibration signals, which faces challenges in distinguishing sentences or words and is difficult to effectively solve the problem of homophones with different meanings during the pronunciation of deaf-mutes. Summary of the Invention
[0004] The purpose of this application is to propose a method and system for testing speech signals based on laryngeal vibration and electroencephalogram electrical stimulation for the above-mentioned technical problems.
[0005] In the first aspect, the present invention provides a method for testing speech signals based on laryngeal vibration and electroencephalogram electrical stimulation, including the following steps:
[0006] Obtain the basic electroencephalogram signal of the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic electroencephalogram signal, and apply electroencephalogram electrical stimulation to the user according to the electrical stimulation control signal;
[0007] Obtain the laryngeal vibration signal and electroencephalogram signal of the user synchronously collected in the speech execution stage and perform data preprocessing on them respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed electroencephalogram signal; use the PageRank algorithm to select channels for the preprocessed electroencephalogram signal to obtain the electroencephalogram signal after channel selection;
[0008] Extract features and perform dimensionality reduction processing on the preprocessed laryngeal vibration signal and the electroencephalogram signal after channel selection respectively to obtain the laryngeal vibration features after dimensionality reduction and the electroencephalogram features after dimensionality reduction, and input the laryngeal vibration features after dimensionality reduction and the electroencephalogram features after dimensionality reduction into the dual-modal SPCA with L1 regularization to obtain sparse features;
[0009] Construct a speech signal classification model based on SVM and train it to obtain a trained speech signal classification model. Input the sparse features into the trained speech signal classification model to obtain the speech classification and recognition results.
[0010] Preferably, it is determined whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic electroencephalogram (EEG) signal, specifically including:
[0011] Use time-frequency analysis technology to analyze the basic EEG signal to obtain the average amplitude and spectral energy of the basic EEG signal. The time-frequency analysis technology includes short-time Fourier transform and wavelet transform;
[0012] In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band related to language processing is less than the amplitude threshold, and the spectral energy is less than the product of the normal reference value and the energy percentage threshold, an electrical stimulation control signal is generated in the speech preparation stage. The frequency range of the EEG electrical stimulation corresponding to the electrical stimulation control signal is 20 - 40 Hz, and the intensity range is 0.5 - 3 mA;
[0013] In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band related to language processing is greater than or equal to the amplitude threshold or the spectral energy is greater than or equal to the product of the normal reference value and the energy percentage threshold, no electrical stimulation control signal is generated in the speech preparation stage.
[0014] Preferably, the basic EEG data is collected when the user does not generate a laryngeal vibration signal and the user is not subjected to EEG electrical stimulation. In the speech preparation stage, it is the time period corresponding to the time threshold range before the user generates a laryngeal vibration signal. In the speech execution stage, it is the time period corresponding to the user generating a laryngeal vibration signal and the user not being subjected to EEG electrical stimulation.
[0015] Preferably, the data preprocessing methods for the laryngeal vibration signal include pre-emphasis processing, gradual in-and-out processing, and smoothing processing; the data preprocessing methods for the EEG signal include band-pass filtering processing, processing using a dynamic notch filter, and artifact removal, where the artifacts need to meet the following conditions: the instantaneous fluctuation between adjacent sampling points exceeds ±50 μV, the difference between adjacent peaks exceeds 200 μV and the time span is within 200 ms, the difference between the maximum amplitude and the minimum amplitude exceeds ±100 μV, the fluctuation is lower than 0.5 μV within a continuous 100 - millisecond time interval, or the power spectral density within the EEG electrical stimulation frequency bandwidth exceeds 3 times the standard deviation of the baseline value.
[0016] Preferably, the PageRank algorithm is used to select channels for the preprocessed EEG signal to obtain the EEG signal after channel selection, specifically including:
[0017] Perform dimensionality transformation on the preprocessed EEG signal to obtain the transformed EEG signal;
[0018] Calculate the Pearson correlation coefficient between every two channels of the transformed EEG signals, and form a correlation matrix. Based on the percentiles of the correlation matrix, map the Pearson correlation coefficient between channel and to the element in the i-th row and j-th column of the adjacency matrix as shown in the following formula: ;
[0019]
[0020] where, , and represent the 25th percentile, 50th percentile, and 75th percentile of all Pearson correlation coefficients, respectively;
[0021] Take each of the 32 channels of the transformed EEG signals as a node in the graph structure. The links between the nodes exist in the form of directed edges in the graph structure. Use the PageRank algorithm to calculate the link relationships between the nodes in the graph structure to determine the importance score of each node. During the calculation of the importance score, if a clustering algorithm is used, use the following formula to calculate the clustering coefficient of each node:
[0022] ;
[0023] where, represents the -th node, represents the number of triangles formed by passing through the -th node and its neighbor nodes, represents the degree of the -th node. The degree of the -th node is equal to the sum of the in-degree of the -th node and the out-degree of the -th node;
[0024] If a degree algorithm is used, use the following formula to calculate the degree centrality:
[0025] ;
[0026] where, represents the degree centrality of the -th node, represents the in-degree of the -th node. The in-degree of the -th node is the sum of the elements in the -th column of the adjacency matrix, represents the out-degree of the -th node, The out-degree of a node is the sum of the elements in the th row of the adjacency matrix;
[0027] The importance score of the node is calculated as follows:
[0028] ;
[0029] where represents the importance score of the th node, represents the normalization function, represents the algorithm weight, and N represents the total number of nodes;
[0030] The adaptive threshold is calculated based on the importance scores of all nodes, as shown in the following formula:
[0031] ;
[0032] where represents the adaptive threshold, represents the similarity parameter;
[0033] Nodes with importance scores greater than or equal to the adaptive threshold are selected, and the corresponding channels are used as important channels. The EEG signals corresponding to the important channels are used as the EEG signals after channel selection.
[0034] Preferably, the laryngeal vibration features extracted from the preprocessed laryngeal vibration signals include Mel-frequency cepstral coefficients, root mean square energy, and chromatic features; the EEG features extracted from the EEG signals after channel selection include root mean square energy, energy in multiple frequency bands, and average spectral entropy, where the multiple frequency bands include five frequency bands: delta, theta, alpha, beta, and gamma; the dimensionality reduction method for laryngeal vibration features and EEG features is the principal component analysis method;
[0035] The objective function of the dual-modal SPCA with L1 regularization is shown in the following formula:
[0036] ;
[0037] where represents the first projection matrix when the objective function takes the maximum value and the second projection matrix , and satisfy = I and = I, where I represents the identity matrix, T represents the transpose, ( ) represents the trace of the matrix, and respectively represent and The L1 norm, where C represents the covariance matrix of the reduced laryngeal vibration features and the reduced EEG features, represents the regularization coefficient;
[0038] By optimizing the objective function of the bimodal SPCA with L1 regularization, the first projection matrix and the second projection matrix are obtained. The reduced laryngeal vibration features and the reduced EEG features are projected using the first projection matrix and the second projection matrix respectively to obtain the projected laryngeal vibration features and the projected EEG features. The projected laryngeal vibration features and the projected EEG features are concatenated to obtain sparse features.
[0039] In a second aspect, the present invention provides a speech signal test system based on laryngeal vibration and EEG electrical stimulation, including:
[0040] An electrical stimulation control signal generation module, configured to acquire the user's basic EEG signal, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal;
[0041] A data processing module, configured to acquire the user's laryngeal vibration signal and EEG signal synchronously collected in the speech execution stage and perform data preprocessing on them respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed EEG signal; perform channel selection on the preprocessed EEG signal using the PageRank algorithm to obtain the EEG signal after channel selection;
[0042] A feature extraction and dimensionality reduction module, configured to perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the EEG signal after channel selection respectively to obtain the reduced laryngeal vibration features and the reduced EEG features, and input the reduced laryngeal vibration features and the reduced EEG features into the bimodal SPCA with L1 regularization to obtain sparse features;
[0043] A classification and recognition module, configured to construct and train a speech signal classification model based on SVM to obtain a trained speech signal classification model, and input the sparse features into the trained speech signal classification model to obtain a speech classification and recognition result.
[0044] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect.
[0045] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0046] Fifthly, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] (1) The electroencephalogram (EEG) channels selected by the EEG signal channel selection method based on the PageRank algorithm in the speech signal test method based on laryngeal vibration and EEG electrical stimulation proposed by the present invention can better represent speech-specific features. In the subsequent data processing and feature extraction processes, the features extracted based on these key channels can more accurately reflect the EEG patterns related to speech of deaf-mutes, making the features more representative and discriminative, and effectively improving the recognition accuracy of the speech information of deaf-mutes.
[0049] (2) The sparse features obtained by the proposed speech signal test method based on laryngeal vibration and EEG electrical stimulation after applying L1-regularized bimodal SPCA processing have better anti-interference ability, can highlight key features, and suppress noise features brought by environmental interference, so that a high recognition accuracy can still be maintained in a complex environment, improving the generalization ability. With the help of technologies such as bimodal SPCA, the internal relationship between laryngeal vibration signals and EEG signals is deeply explored, and speech features are comprehensively extracted to ensure the accurate grasp of various speech information.
[0050] (3) The proposed speech signal test method based on laryngeal vibration and EEG electrical stimulation of the present invention can effectively collect the laryngeal vibration signals and EEG signals of deaf-mutes when speaking. Through the supplementation of EEG signals to laryngeal vibration signals, the problem of homophones with different meanings is effectively solved, and intuitive results are output for translating the speech information expressed by deaf-mutes, realizing more accurate recognition of sensing signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a schematic flowchart of the speech signal test method based on laryngeal vibration and EEG electrical stimulation for the embodiments of the present application;
[0053] Figure 2 Schematic diagram of data acquisition and interaction of the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0054] Figure 3 Schematic diagram of the laryngeal vibration signal of the word "airport" in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0055] Figure 4 Schematic diagram of the laryngeal vibration signal of the word "down" in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0056] Figure 5 Schematic diagram of the laryngeal vibration signal of the homophone "bough" in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0057] Figure 6 Schematic diagram of the laryngeal vibration signal of the homophone "bow" in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0058] Figure 7 Schematic diagram of the electroencephalogram signal of the homophone "bough" in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0059] Figure 8 Schematic diagram of the electroencephalogram signal of the homophone "bow" in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0060] Figure 9 Schematic diagram of the confusion matrix of the single laryngeal vibration signal of the homophones in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0061] Figure 10 Schematic diagram of the confusion matrix of the mixed signal composed of the laryngeal vibration signal and the electroencephalogram signal of the homophones in the voice signal test method based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0062] Figure 11 Schematic diagram of the voice signal test system based on laryngeal vibration and electroencephalogram electrical stimulation in the embodiments of the present application;
[0063] Figure 12 Schematic diagram of the hardware structure of the electronic device provided in the embodiments of the present invention. Detailed implementation manners
[0064] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] Figure 1 A voice signal testing method based on laryngeal vibration and electroencephalogram (EEG) electrical stimulation provided by an embodiment of the present application is shown, including the following steps:
[0066] S1. Obtain the basic EEG signal of the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal.
[0067] In a specific embodiment, determining whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal specifically includes:
[0068] Analyze the basic EEG signal by using time-frequency analysis technology to obtain the average amplitude and spectral energy of the basic EEG signal. The time-frequency analysis technology includes short-time Fourier transform and wavelet transform;
[0069] In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band related to language processing is less than the amplitude threshold, and the spectral energy is less than the product of the normal reference value and the energy percentage threshold, an electrical stimulation control signal is generated in the speech preparation stage. The frequency range of the EEG electrical stimulation corresponding to the electrical stimulation control signal is 20 - 40 Hz, and the intensity range is 0.5 - 3 mA;
[0070] In response to determining that the average amplitude of the basic EEG signal in the θ-β frequency band related to language processing is greater than or equal to the amplitude threshold or the spectral energy is greater than or equal to the product of the normal reference value and the energy percentage threshold, no electrical stimulation control signal is generated in the speech preparation stage.
[0071] In a specific embodiment, the basic EEG data is collected when the user does not generate a laryngeal vibration signal and the user is not subjected to EEG electrical stimulation. The speech preparation stage is the time period corresponding to the time threshold range before the user generates a laryngeal vibration signal, and the speech execution stage is the time period corresponding to the user generating a laryngeal vibration signal and the user not being subjected to EEG electrical stimulation.
[0072] Specifically, refer to Figure 2, in the embodiments of the present application, an electroencephalogram (EEG) cap can be used to collect EEG signals. The EEG cap integrates an electrical stimulation function module. Since the deaf-mute population lacks long-term language output training, the neural activity intensity in their language-related brain regions (such as Broca area and Wernicke area) is significantly lower than that of healthy people. Research shows that the power in the gamma frequency band drops by about 40-60%. Electrical brain stimulation (such as tACS / tDCS) can adjust the neuron membrane potential and enhance the synchronous activity in the target brain region, thereby compensating for the signal attenuation problem. When the user is about to speak, first, the EEG cap collects the basic EEG signals at a sampling rate of 256 Hz through 32 channels of the international 10 / 20 system. The time-frequency analysis technology (including the short-time Fourier transform (STFT) and wavelet transform) methods are used to analyze the signal characteristics and stability, and the average amplitude and spectral energy of the basic EEG signals are obtained.
[0073] Verified by clinical research, if the following conditions are detected:
[0074] 1. The average amplitude of the EEG signal in the θ-β frequency band (4~30 Hz) related to language processing is lower than 30 μV. The θ wave (4~7 Hz) is related to language understanding.
[0075] 2. The spectral energy of the EEG signal in the θ-β frequency band related to language processing is reduced by more than 40% compared with the normal reference value.
[0076] Then an electrical stimulation control signal is automatically generated and sent to the electrical stimulation function module to generate electrical brain stimulation. The electrical brain stimulation is only applied during the speech preparation stage. At the same time, the parameter information of the electrical stimulation is recorded, including the stimulation frequency, intensity, etc. In one of the embodiments, the speech preparation stage is 300~500 ms before the triggering of the laryngeal vibration signal. The stimulation frequency range is 20~40 Hz, and the intensity is 0.5~3 mA. In other embodiments, the frequency and intensity can also be adjusted according to the electrode area and adjusted by itself according to the signal characteristics and stability. There is no electrical brain stimulation in the EEG signals during the speech execution stage. Subsequently, the laryngeal vibration signal and the EEG signal are synchronously collected. And during the collection process, different collection test scenarios can be set, such as different environments, different users, etc., to comprehensively evaluate the system performance.
[0077] The laryngeal vibration signal is collected by using a laryngeal pressure sensor. In one of the embodiments, the laryngeal pressure sensor uses a hydrogel pressure sensor and is connected to an electrochemical workstation (PalmSens4). The open circuit potentiometry method is used to collect the laryngeal vibration signal.
[0078] The user wears an EEG cap (Emotive) to record electroencephalogram (EEG) signals, and wears a hydrogel pressure sensor on the larynx to record laryngeal (AT) vibration signals through an electrochemical workstation (PalmSens 4). At the same time, the user gazes at the word pictures displayed on the computer screen for visual stimulation. The EEG signals and laryngeal vibration signals are synchronously collected with a unified timestamp to ensure the time alignment of multi-modal data. In each experimental cycle, the user reads the displayed words within 3 seconds, and the interval between words exceeds 10 seconds for rest. There is an intermission of about 5 minutes after each experimental session. The EEG signals are collected from 32 channels at a sampling rate of 256 Hz (according to the international 10 / 20 system), while the laryngeal vibration signals are collected from the larynx (at the thyroid cartilage) at a sampling rate of 250 Hz. The Fz channel is used as the ground, and the CP2 channel is used as the reference.
[0079] S2. Obtain the laryngeal vibration signals and EEG signals of the user synchronously collected during the speech execution stage, and perform data preprocessing on them respectively to obtain the preprocessed laryngeal vibration signals and preprocessed EEG signals; use the PageRank algorithm to select channels for the preprocessed EEG signals to obtain the EEG signals after channel selection.
[0080] In a specific embodiment, the data preprocessing methods for laryngeal vibration signals include pre-emphasis processing, fade-in and fade-out processing, and smoothing processing; the data preprocessing methods for EEG signals include band-pass filtering processing, processing with a dynamic notch filter, and artifact removal, where the artifacts need to meet the following conditions: the instantaneous fluctuation between adjacent sampling points exceeds ±50 μV, the difference between adjacent peaks exceeds 200 μV and the time span is within 200 ms, the difference between the maximum amplitude and the minimum amplitude exceeds ±100 μV, the fluctuation is lower than 0.5 μV within a continuous 100-ms time interval, or the power spectral density exceeds 3 times the standard deviation of the baseline value within the EEG electrical stimulation frequency bandwidth.
[0081] Specifically, the preprocessing process for multi-modal data is as follows:
[0082] Perform pre-emphasis processing on the laryngeal signals, keep the low-frequency part of the signals unchanged, and enhance the high-frequency part of the signals. Perform fade-in and fade-out on the pre-emphasized laryngeal signals to remove the mutations caused by the pre-emphasis processing of the truncated signals. Then perform smoothing processing on the signals to reduce the noise in the data, enhance the signal quality, and prevent abnormal signal disturbances.
[0083] Perform band-pass filtering on the EEG signals from 0.5 to 50 Hz to remove mains interference, and combine a dynamic notch filter to eliminate the frequency interference of EEG electrical stimulation in real time. Then perform artifact removal (denoising, removing electrooculogram, electromyogram, and residual electrical stimulation). In the processing of EEG signals, any of the following five artifact recognition conditions is marked as an artifact, specifically as follows:
[0084] 1. The instantaneous fluctuation between adjacent sampling points exceeds ±50 μV;
[0085] 2. The difference between adjacent peaks exceeds 200 μV and the time span is within 200 ms;
[0086] 3. The difference between the maximum amplitude and the minimum amplitude exceeds ±100 μV;
[0087] 4. The fluctuation is lower than 0.5 μV within a continuous time interval of 100 ms;
[0088] 5. Within the stimulation frequency bandwidth, the power spectral density exceeds 3 times the standard deviation of the baseline value.
[0089] In a specific embodiment, the preprocessed EEG signals are subjected to channel selection using the PageRank algorithm to obtain the EEG signals after channel selection, which specifically includes:
[0090] Performing dimensionality transformation on the preprocessed EEG signals to obtain the transformed EEG signals;
[0091] Calculating the Pearson correlation coefficient between every two channels of the transformed EEG signals and forming a correlation matrix. Based on the percentile of the correlation matrix, the Pearson correlation coefficient and between channels is mapped to the element in the i-th row and j-th column of the adjacency matrix , as shown in the following formula:
[0092] ;
[0093] where , and respectively represent the 25th percentile, 50th percentile, and 75th percentile of all Pearson correlation coefficients;
[0094] Taking each of the 32 channels of the transformed EEG signals as a node in the graph structure, and the links between the nodes exist in the form of directed edges in the graph structure. The PageRank algorithm is used to calculate the link relationships between the nodes in the graph structure to determine the importance scores of each node; during the calculation of the importance scores, if a clustering algorithm is used, the following formula is used to calculate the clustering coefficient of each node:
[0095] ;
[0096] where represents the -th node, represents the number of triangles formed by passing through the -th node and its neighbor nodes, Denote the degree of the -th node. The degree of the -th node is equal to the sum of the in-degree and the out-degree of the -th node;
[0097] If the degree algorithm is adopted, the degree centrality is calculated by the following formula:
[0098] ;
[0099] where denotes the degree centrality of the -th node, denotes the in-degree of the -th node. The in-degree of the -th node is the sum of the elements in the -th column of the adjacency matrix, denotes the out-degree of the -th node. The out-degree of the -th node is the sum of the elements in the -th row of the adjacency matrix;
[0100] The calculation formula for the importance score of the -th node is as follows:
[0101] ;
[0102] where denotes the importance score of the -th node, denotes the normalization function, denotes the algorithm weight, and N denotes the total number of nodes;
[0103] The adaptive threshold is calculated based on the importance scores of all nodes, as shown in the following formula:
[0104] ;
[0105] where denotes the adaptive threshold, denotes the similarity parameter;
[0106] Select the nodes with importance scores greater than or equal to the adaptive threshold, and use the corresponding channels as important channels. The EEG signals corresponding to the important channels are used as the EEG signals after channel selection.
[0107] Specifically, the PageRank algorithm is used for channel selection on the preprocessed EEG signals. Specifically, the PageRank algorithm is a link analysis algorithm based on a directed graph. It determines the importance of each node by iteratively calculating the link relationships between nodes. This algorithm evaluates their relative importance by analyzing the link (directed edge) relationships between nodes. Specifically, the PageRank algorithm treats links as votes. When a node links to another node, it is equivalent to voting for the linked node. Different nodes have different weights for the votes given to the linked node, and the weights depend on the importance of the voting node, that is, the importance score. In the embodiments of the present application, 32 channels are regarded as nodes in the directed graph. The specific steps of the PageRank algorithm are as follows:
[0108] 1. First, perform dimensionality transformation on the preprocessed EEG signals as shown in the following formula: (trail, channel, seq_len) -> (channel, trail × seq_len), where trail represents the number of data items, channel represents the channel, and seq_len represents the data length of each data item, and calculate the Pearson correlation coefficient between any two channels;
[0109] 2. Divide the elements in the adjacency matrix through the 25th percentile, 50th percentile, and 75th percentile. The elements in the (0, 25) interval are assigned 0, the elements in the (25, 50) interval are assigned 1, the elements in the (50, 75) interval are assigned 2, and the elements in the (75, 100) interval are assigned 3. Finally, obtain a 32×32 adjacency matrix A, where each element represents the connection strength (weight) between channels to . The larger the value, the stronger the correlation.
[0110] 3. Calculate the clustering coefficient and degree centrality respectively. Use the clustering algorithm to calculate the clustering coefficient. The goal of the clustering coefficient is to measure the clustering tightness of a group; use the degree algorithm to calculate the degree centrality. The degree centrality statistics the number of edges directly connected to a node, including the out-degree and in-degree, and can be simply understood as the size of the access opportunity of a node. The out-degree of the th node is the sum of the weights of all directed edges starting from the th node, that is, the sum of the elements in the th row of the adjacency matrix; the in-degree of the th node is the sum of the weights of all directed edges ending at the th node, that is, the sum of the elements in the The sum of the elements in the column. The degree centrality is the sum of the out-degree and in-degree. By harmonizing these two coefficients to obtain the importance score for each channel (node), the larger the value, the more attention is paid to the global network, When it is 1, it is the clustering algorithm, the smaller the value, the more attention is paid to the local network, When it is 0, it is the degree algorithm.
[0111] 4. Calculate the adaptive threshold according to the importance score, and filter out the nodes in the corresponding network through the adaptive threshold. In one embodiment, 16 important channels in the EEG signal are finally filtered out, and the shape of the filtered EEG signal is (channel: 16, seq_len: 750).
[0112] S3. Respectively perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the EEG signal after channel selection to obtain the dimensionality-reduced laryngeal vibration features and the dimensionality-reduced EEG features, and input the dimensionality-reduced laryngeal vibration features and the dimensionality-reduced EEG features into the bimodal SPCA with L1 regularization to obtain sparse features.
[0113] In a specific embodiment, the laryngeal vibration features extracted from the preprocessed laryngeal vibration signal include Mel frequency cepstral coefficients, root mean square energy, and chromaticity features; the EEG features extracted from the EEG signal after channel selection include root mean square energy, energies in multiple frequency bands, and average spectral entropy, where the multiple frequency bands include five frequency bands: delta, theta, alpha, beta, and gamma; the dimensionality reduction method for the laryngeal vibration features and the EEG features adopts the principal component analysis method;
[0114] Apply L1 regularization to the objective function of the bimodal SPCA with L1 regularization to obtain the objective function of the bimodal SPCA with L1 regularization after applying L1 regularization, as shown in the following formula:
[0115] ;
[0116] where, represents the first projection matrix when the objective function takes the maximum value and the second projection matrix , and satisfy =I and =I, I represents the identity matrix, T represents the transpose, ( ) represents the trace of the matrix, and respectively represent and 's L1 norm, C represents the covariance matrix of the dimensionality-reduced laryngeal vibration features and the dimensionality-reduced EEG features, denotes the regularization coefficient;
[0117] By optimizing the objective function of the bimodal SPCA with L1 regularization, the first projection matrix and the second projection matrix are obtained. The reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features are respectively projected using the first projection matrix and the second projection matrix to obtain the projected laryngeal vibration features and the projected EEG features. The projected laryngeal vibration features and the projected EEG features are concatenated to obtain sparse features.
[0118] Specifically, feature extraction and dimensionality reduction processing are performed on the preprocessed laryngeal vibration signals and the EEG signals after channel selection to reduce the computational amount without reducing the accuracy, improve the recognition speed, and thus obtain multi-modal features. The extracted laryngeal vibration features include: Mel-frequency cepstral coefficients (MFCC), root mean square energy (RMS), and chroma features (CHROMA) in the laryngeal vibration signals; the extracted EEG features include the root mean square energy (RMS) of the EEG signals, and the energies and average spectral entropies of the five frequency bands of delta, theta, alpha, beta, and gamma.
[0119] The principal component analysis method is respectively used for the laryngeal vibration features and the EEG features, and the principal components with a variance contribution rate exceeding 95% are retained for dimensionality reduction to obtain the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features; then, the bimodal SPCA (Bimodal-sSPCA) with LI regularization is used for processing, and finally 64-dimensional sparse features with speech specificity are obtained. The bimodal SPCA can simultaneously consider the correlation between the two different modal data of the laryngeal vibration signals and the EEG signals. It constructs a joint objective function to find the projection direction that can make the two modal data both maintain their respective features in the low-dimensional space and can maximize the correlation between the modalities. Applying L1 regularization in the objective function of the bimodal SPCA can obtain a sparse projection matrix and feature representation. The sparse features have better interpretability, can highlight the features that contribute more to speech specificity, and at the same time suppress noise and irrelevant features. This helps to improve the generalization ability and computational efficiency of the model. Ordinary PCA does not have this sparse constraint, and the obtained features may contain more redundant information.
[0120] The objective function of the original bimodal SPCA is:
[0121] ;
[0122] where denotes the covariance matrix of the two-modal features, which describes the second-order statistical relationship between the reduced-dimensional laryngeal vibration features and the reduced-dimensional EEG features and is obtained by calculating the expectation.
[0123] The objective function of the dual-modal SPCA with L1 regularization is as follows:
[0124] ;
[0125] By =I and =I constraint conditions, the first projection matrix and the second projection matrix are orthogonal matrices.
[0126] Use the method combining alternating optimization and Lagrange multiplier method to optimize the objective function of the dual-modal SPCA with L1 regularization. The specific process is as follows:
[0127] Fix the second projection matrix , and take the derivative of the objective function of the dual-modal SPCA with L1 regularization with respect to the first projection matrix (considering the constraint conditions and L1 regularization). Let the Lagrangian function be:
[0128] L( ) = ;
[0129] where is the Lagrange multiplier matrix. Take the derivative of L( ) and set it to 0, and finally obtain the first projection matrix ; Similarly, fix the first projection matrix , and update the second projection matrix .
[0130] Then project the reduced laryngeal vibration features using the first projection matrix to obtain the projected laryngeal vibration features; project the reduced EEG features using the second projection matrix to obtain the projected EEG features, and then concatenate the projected laryngeal vibration features and the projected EEG features to obtain the final 64-dimensional sparse features with speech specificity.
[0131] S4. Construct and train a speech signal classification model based on SVM to obtain a trained speech signal classification model. Input the sparse features into the trained speech signal classification model to obtain the speech classification and recognition results.
[0132] Specifically, the embodiment of the present application uses a support vector machine (SVM) classifier as a speech signal classification model. The sparse features are used as input signals and divided into training data and test data. The support vector machine (SVM) classifier is trained to obtain a trained speech signal classification model. By inputting the sparse features into the trained speech signal classification model, the speech classification and recognition results can be obtained. At the same time, the performance under different test conditions is evaluated. For example, under different environments, different subjects, etc., by calculating indicators such as recognition accuracy, recall rate, and F1 value, the performance of the system is comprehensively measured. As shown in Table 1, it is a comparison of the recognition accuracy, F1 value, recall rate, and precision rate of two different users for ten words of easily confused words "Meet, Meat, Bow, Bough, Hear, Here, Costume, Custom, Implicit, Explicit". It can be seen that the generalization ability of the entire system is relatively strong, and the accuracy rate fluctuates little for different users.
[0133] Table 1
[0134]
[0135] Specifically, when a person is speaking, there are obvious vibrations in the larynx. Different speech contents have different laryngeal vibration signals. For example, when reading the words "airport" and "down", as Figure 3 and 4 shown, the waveforms of the two words are significantly different. However, when it comes to easily confused words, such as Figure 5 and 6 shown, since the pronunciations of "bough" and "bow" are similar, the laryngeal vibration signals are extremely similar and difficult to distinguish. Compared with laryngeal signals, electroencephalogram (EEG) signals have great advantages in identifying easily confused words. It contains a large amount of information. As Figure 7 and 8 shown, the EEG signal waveforms of "bough" and "bow" are significantly different, which is of great significance for identifying and realizing speech signals. Therefore, the speech signal test method based on laryngeal vibration and electroencephalogram electrical stimulation proposed in the embodiment of the present application can overcome the problems of inability to identify easily confused words by a single information source and low recognition accuracy. The accuracy rate in the task of identifying and classifying easily confused words has been greatly improved from 69.33% of a single laryngeal signal to 91.50% of the fused signal, as Figure 9 and 10 shown.
[0136] Similarly, a robust recognition rate of 96.26% (Mult) is obtained for other groups of words or phrases. Tables 2 and 3 are words in other groups, and Table 4 is the recognition accuracy for other groups.
[0137] Table 2:
[0138]
[0139] Table 3:
[0140]
[0141] Table 4:
[0142]
[0143] Further reference Figure 11 As an implementation of the methods shown in the above figures, an embodiment of the present application provides a voice signal testing system based on laryngeal vibration and electroencephalogram (EEG) electrical stimulation. This system embodiment corresponds to Figure 1 the method embodiment shown, and this system can be specifically applied to various electronic devices.
[0144] An embodiment of the present application provides a voice signal testing system based on laryngeal vibration and EEG electrical stimulation, including:
[0145] An electrical stimulation control signal generation module 1, configured to obtain the acquired basic EEG signal of the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signal, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal;
[0146] A data processing module 2, configured to obtain the laryngeal vibration signal and EEG signal synchronously acquired by the user in the speech execution stage and perform data preprocessing on them respectively to obtain the preprocessed laryngeal vibration signal and the preprocessed EEG signal; perform channel selection on the preprocessed EEG signal using the PageRank algorithm to obtain the EEG signal after channel selection;
[0147] A feature extraction and dimensionality reduction module 3, configured to perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the EEG signal after channel selection respectively to obtain the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, and input the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction into a dual-modal SPCA with L1 regularization to obtain sparse features;
[0148] A classification and recognition module 4, configured to construct and train a voice signal classification model based on SVM to obtain a trained voice signal classification model, and input the sparse features into the trained voice signal classification model to obtain a voice classification and recognition result.
[0149] Figure 12 is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. As shown in Figure 12As shown in the figure, the electronic device of this embodiment includes: a processor 1201 and a memory 1202; wherein the memory 1202 is used to store computer-executable instructions; the processor 1201 is used to execute the computer-executable instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. For specific details, please refer to the relevant descriptions in the foregoing method embodiments.
[0150] Optionally, the memory 1202 can be either independent or integrated with the processor 1201.
[0151] When the memory 1202 is independently provided, the electronic device further includes a bus 1203 for connecting the memory 1202 and the processor 1201.
[0152] The embodiment of the present invention also provides a computer storage medium, in which computer-executable instructions are stored. When the processor 1201 executes the computer-executable instructions, the above method is implemented.
[0153] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 1201, the above method is implemented.
[0154] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical or other forms.
[0155] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0156] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module exists physically alone, or two or more modules are integrated in one unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a hardware plus software functional unit.
[0157] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 1201 to execute some steps of the methods according to the various embodiments of the present application.
[0158] It should be understood that the above-mentioned processor 1201 can be a central processing unit (CPU for short), or can also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor 1201 can also be any conventional processor 1201, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor 1201, or can be implemented by the combination of the hardware and software modules in the processor 1201.
[0159] The memory 1202 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disc, etc.
[0160] The bus 1203 can be an Industry Standard Architecture (ISA for short), a Peripheral Component Interconnect (PCI for short) bus, or an Extended Industry Standard Architecture (EISA for short) bus, etc. The bus 1203 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus 1203 in the drawings of the present application is not limited to only one bus 1203 or one type of bus 1203.
[0161] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0162] An exemplary storage medium is coupled to the processor 1201, enabling the processor 1201 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 1201. The processor 1201 and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor 1201 and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0163] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice signal testing method based on laryngeal vibration and electroencephalogram electrical stimulation, characterized in that It includes the following steps: Obtain the basic electroencephalogram (EEG) signals collected from the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signals, and apply EEG electrical stimulation to the user according to the electrical stimulation control signal; Obtain the laryngeal vibration signals and EEG signals synchronously collected in the speech execution stage and perform data preprocessing on them respectively to obtain the preprocessed laryngeal vibration signals and preprocessed EEG signals; Use the PageRank algorithm to perform channel selection on the preprocessed EEG signals to obtain the EEG signals after channel selection; Perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signals and the EEG signals after channel selection respectively to obtain the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction, and input the laryngeal vibration features after dimensionality reduction and the EEG features after dimensionality reduction into the dual-modal SPCA with L1 regularization to obtain sparse features; Construct a speech signal classification model based on SVM and train it to obtain a trained speech signal classification model, and input the sparse features into the trained speech signal classification model to obtain a speech classification and recognition result.
2. The method for testing speech signals based on laryngeal vibration and electroencephalogram electrical stimulation according to claim 1, wherein Determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signals, specifically including: Analyze the basic EEG signals using time-frequency analysis techniques to obtain the average amplitude and spectral energy of the basic EEG signals. The time-frequency analysis techniques include short-time Fourier transform and wavelet transform; In response to determining that the average amplitude of the basic EEG signals in the θ-β frequency band related to language processing is less than the amplitude threshold and the spectral energy is less than the product of the normal reference value and the energy percentage threshold, generate the electrical stimulation control signal in the speech preparation stage. The frequency range of the EEG electrical stimulation corresponding to the electrical stimulation control signal is 20~40Hz, and the intensity range is 0.5~3mA; In response to determining that the average amplitude of the basic EEG signals in the θ-β frequency band related to language processing is greater than or equal to the amplitude threshold or the spectral energy is greater than or equal to the product of the normal reference value and the energy percentage threshold, do not generate the electrical stimulation control signal in the speech preparation stage.
3. The voice signal testing method based on laryngeal vibration and electroencephalogram electrical stimulation according to claim 1, wherein The basic EEG data is collected when the user does not generate laryngeal vibration signals and the user is not subjected to EEG electrical stimulation. In the speech preparation stage, it is the time period corresponding to the time threshold range before the user generates laryngeal vibration signals. In the speech execution stage, it is the time period corresponding to the user generating laryngeal vibration signals and the user not being subjected to EEG electrical stimulation.
4. The voice signal testing method based on laryngeal vibration and electroencephalogram electrical stimulation according to claim 1, wherein The data preprocessing methods for the laryngeal vibration signals include pre-emphasis processing, fade-in and fade-out processing, and smoothing processing; the data preprocessing methods for the EEG signals include band-pass filtering processing, processing using a dynamic notch filter, and artifact removal, where the artifacts need to meet the following conditions: the instantaneous fluctuation between adjacent sampling points exceeds ±50 μV, the difference between adjacent peaks exceeds 200 μV and the time span is within 200 ms, the difference between the maximum amplitude and the minimum amplitude exceeds ±100 μV, the fluctuation is lower than 0.5 μV within a continuous 100-ms time interval, or the power spectral density within the EEG electrical stimulation frequency bandwidth exceeds 3 times the standard deviation of the baseline value.
5. The voice signal testing method based on laryngeal vibration and electroencephalogram electrical stimulation according to claim 1, wherein The PageRank algorithm is used to perform channel selection on the preprocessed EEG signals to obtain the EEG signals after channel selection, specifically including: Performing dimensionality transformation on the preprocessed EEG signals to obtain the transformed EEG signals; Calculate the Pearson correlation coefficient between every two channels of the transformed EEG signals, and form a correlation matrix. Based on the percentile of the correlation matrix, map the Pearson correlation coefficient between to the element in the i-th row and j-th column of the adjacency matrix as follows: as shown in the following formula: ; Among them, , and represent the 25th percentile, 50th percentile, and 75th percentile of all Pearson correlation coefficients, respectively; Regarding each of the 32 channels of the transformed EEG signals as a node in the graph structure, the links between the nodes exist in the form of directed edges in the graph structure, and the PageRank algorithm is used to calculate the link relationships between the nodes in the graph structure to determine the importance scores of each node; during the calculation process of the importance scores, if the clustering algorithm is used, the following formula is used to calculate the clustering coefficient of each node: ; Among them, represents the th node, represents the number of triangles formed by passing through the th node and its neighbor nodes, represents the degree of the th node. The degree of the th node is equal to the sum of the in-degree of the th node and the out-degree of the th node; If the degree algorithm is used, the following formula is used to calculate the degree centrality: ; Among them, represents the degree centrality of the th node, represents the in-degree of the th node. The in-degree of the th node is the sum of the elements in the th column of the adjacency matrix, represents the out-degree of the th node. The out-degree of the th node is the sum of the elements in the th row of the adjacency matrix; The importance score calculation formula for the nth node is as follows: ; Among them, represents the importance score of the th node, represents the normalization function, represents the algorithm weight, and N represents the total number of nodes; The adaptive threshold is calculated based on the importance scores of all nodes, as shown in the following formula: ; Among them, represents the adaptive threshold, represents the similarity parameter; The nodes with importance scores greater than or equal to the adaptive threshold are selected, and the corresponding channels are used as important channels, and the EEG signals corresponding to the important channels are used as the EEG signals after channel selection.
6. The method for testing a speech signal based on laryngeal vibration and electroencephalogram electrical stimulation according to claim 1, characterized in that, The laryngeal vibration features extracted from the preprocessed laryngeal vibration signals include Mel frequency cepstral coefficients, root mean square energy, and chromaticity features; The EEG features extracted from the EEG signals after channel selection include root mean square energy, energies of multiple frequency bands, and average spectral entropy, where the multiple frequency bands include five frequency bands: delta, theta, alpha, beta, and gamma; the dimensionality reduction method for the laryngeal vibration features and EEG features uses the principal component analysis method; The objective function of the dual-modal SPCA with L1 regularization is shown in the following formula: ; Among them, represents the first projection matrix when the objective function takes the maximum value and the second projection matrix , and satisfy = I and = I, where I represents the identity matrix and T represents the transpose, ( ) represents the trace of the matrix, and respectively represent and 's L1 norm, C represents the covariance matrix of the reduced laryngeal vibration characteristics and the reduced EEG characteristics, represents the regularization coefficient; By optimizing the objective function of the dual-modal SPCA with L1 regularization, the first projection matrix and the second projection matrix are obtained, and the first projection matrix and the second projection matrix are used to project the dimensionality-reduced laryngeal vibration features and the dimensionality-reduced EEG features respectively to obtain the projected laryngeal vibration features and the projected EEG features, and the projected laryngeal vibration features and the projected EEG features are concatenated to obtain sparse features.
7. A voice signal testing system based on laryngeal vibration and electroencephalogram electrical stimulation, characterized in that, Including: An electrical stimulation control signal generation module, configured to acquire the basic EEG signals of the user, determine whether to generate an electrical stimulation control signal in the speech preparation stage according to the basic EEG signals, and apply an EEG electrical stimulation to the user according to the electrical stimulation control signal; A data processing module, configured to acquire the laryngeal vibration signal and electroencephalogram signal of a user synchronously collected during the voice execution phase and perform data preprocessing on them respectively, so as to obtain the preprocessed laryngeal vibration signal and the preprocessed electroencephalogram signal; Perform channel selection on the preprocessed electroencephalogram signal by using the PageRank algorithm to obtain the electroencephalogram signal after channel selection; A feature extraction and dimensionality reduction module, configured to perform feature extraction and dimensionality reduction processing on the preprocessed laryngeal vibration signal and the electroencephalogram signal after channel selection respectively, so as to obtain the laryngeal vibration features after dimensionality reduction and the electroencephalogram features after dimensionality reduction, and input the laryngeal vibration features after dimensionality reduction and the electroencephalogram features after dimensionality reduction into the dual-modal SPCA with L1 regularization to obtain sparse features; A classification and recognition module, configured to construct and train a voice signal classification model based on SVM to obtain a trained voice signal classification model, and input the sparse features into the trained voice signal classification model to obtain a voice classification and recognition result.
8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Voice signal processing method and related device therefor
CN114072875A
Voice decoding recognition method and system for multi-mode throat vibration signal and lip moving point data
CN119068870A