Mandarin Chinese speech decoding method based on comparative learning mode matching
By combining the comparative learning mode matching method between SEEG data and audio data, the residual convolutional neural network and information noise comparison estimation function are used to train the feature extractor, which solves the problem of insufficient comprehensive utilization of modals in implantable speech decoding, and improves the decoding accuracy of Mandarin in Chinese.
Patent Information
- Application Number
- CN202510561360.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art lacks a comprehensive utilization of intracranial data and audio-assisted modal decoding methods in the implantable speech decoding brain-computer interface, resulting in insufficient decoding performance.
A method based on contrast learning modal matching is adopted, combining SEEG data and audio data, features are extracted through residual convolutional neural networks, and the SEEG feature extractor is trained using information noise comparison estimation function to establish a modal matching pattern and enhance sample matching in the feature space.
It improves the recognition accuracy of Chinese Mandarin language decoding and improves the decoding accuracy and reliability of the brain-computer interface system.
Smart Images

Figure CN120299450A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of brain-computer interfaces for Mandarin speech decoding, and more particularly, to a Mandarin speech decoding method based on contrastive learning modality matching. Background Art
[0002] Speech communication is a basic ability required by humans. Individuals communicate ideas, convey emotions, and express opinions through speech in interpersonal communication and social environments. Speech production involves a series of complex neural activities: the human brain first generates conceptual language representations, then conducts articulatory motor organization, and finally transmits articulatory motor signals through the brainstem and peripheral nervous system to activate relevant muscles in the face, larynx, tongue, etc. to produce speech. Any interruption in this process may cause aphasia. Neurological diseases such as stroke or amyotrophic lateral sclerosis can damage the motor-related neural pathways, resulting in paralysis and loss of control of the articulatory muscles. Patients often show dysarthria, that is, they still have speech cognitive abilities but have difficulty controlling the articulatory organs to produce intelligible speech. In severe cases, it is manifested as locked-in syndrome, in which almost all other motor functions are lost except for limited eye and head movement abilities, and the quality of life is significantly reduced.
[0003] Brain-computer interfaces (BCIs) are direct interaction pathways between the brain and external devices. In recent years, speech neuroprosthetic brain-computer interfaces that are expected to help aphasic patients restore communication abilities have received increasing attention, and speech neuroprostheses based on implanted electrodes such as electrocorticography (ECoG) or microelectrode arrays (MEA) have made breakthroughs. Using neural signals in speech-related brain regions, researchers have successfully decoded phonemes, words, and even sentences, helping patients restore a certain degree of speech communication.
[0004] Stereo-electroencephalography (SEEG) is a minimally invasive intracranial signal acquisition technology that has currently been widely used for the localization of epileptogenic foci in patients with refractory epilepsy. SEEG can synchronously record neural activities from different depths in multiple brain regions and is a powerful tool for studying the cooperative patterns between brain regions. Currently, there have been works on using SEEG signals for speech synthesis and phoneme classification, reporting performance effects higher than random levels.
[0005] Most traditional speech decoding methods rely on single-modal intracranial data. However, the high cost of implantable neural data acquisition and the scarcity of large-scale available datasets emphasize the importance of using auxiliary modal data to improve decoding performance. As a common auxiliary modality in the speech decoding task of brain-computer interfaces, audio signals can provide additional contextual cues and acoustic feature information, and have the potential to supplement sparse neural recordings and improve speech decoding performance.
[0006] In summary, combining implantable speech neural recordings and audio auxiliary data can fully exploit the inherent information within modalities and the potential correlations between modalities, improving the decoding accuracy and reliability of speech decoding brain-computer interface systems. Summary of the Invention
[0007] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a Mandarin Chinese speech decoding method based on contrastive learning modality matching, aiming to solve the problem of the lack of a decoding method that comprehensively utilizes intracranial data and audio auxiliary modalities in implantable speech decoding brain-computer interfaces.
[0008] To achieve the above purpose, the present invention provides a speech decoding method based on contrastive learning modality matching, including:
[0009] Training stage:
[0010] S1, construct a dataset; the dataset includes original SEEG data and corresponding audio data;
[0011] S2, input the SEEG data into the randomly initialized SEEG feature extractor to obtain SEEG features; input the audio data corresponding to the SEEG data into the pre-trained audio feature extractor to obtain audio features; wherein, the dimensions of the SEEG features and the audio features are the same, and the SEEG feature extractor is a residual convolutional neural network;
[0012] S3, respectively mark each SEEG data as the first anchor point, mark the corresponding audio data as the first positive sample, and mark other audio data as the first negative sample, and calculate the source loss between each audio feature and SEEG feature with the information noise contrast estimation function as the objective function; respectively mark each audio data as the second anchor point, mark the corresponding SEEG data as the second positive sample, and mark other SEEG data as the second negative sample, and calculate the symmetric loss between each audio feature and SEEG feature with the information noise contrast estimation function as the objective function; use the sum of the source loss and the symmetric loss as the total loss to train the SEEG feature extractor;
[0013] Application stage:
[0014] Input the SEEG data to be tested into the trained SEEG feature extractor to obtain SEEG features, and input all candidate audio data into the audio feature extractor to obtain candidate audio features;
[0015] Obtain the matching audio result based on the similarity between the SEEG features and all candidate audio features to complete speech decoding.
[0016] The present invention also provides a method for decoding Mandarin Chinese speech based on contrastive learning modality matching, including:
[0017] Input the SEEG data to be recognized into the SEEG feature extractor trained in the training stage of the above method to obtain the corresponding SEEG features, input all candidate audio data into the audio feature extractor to obtain candidate audio features, and obtain the matching audio result based on the similarity between the SEEG features and all candidate audio features to complete speech decoding; wherein, the audio data is the audio data when the subject reads the Mandarin Chinese target word, and the SEEG data is the stereo electroencephalogram data when the subject reads the Mandarin Chinese target word.
[0018] The present invention also provides an electronic device, including: a computer-readable storage medium and a processor;
[0019] The computer-readable storage medium is used to store executable instructions;
[0020] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the above method.
[0021] The present invention also provides a computer-readable storage medium, which stores computer instructions for causing a processor to execute the above method.
[0022] The present invention also provides a computer program product, including a computer program or instructions, which implement the above method when executed by a processor.
[0023] Through the above technical solutions conceived by the present invention, compared with the prior art, the following beneficial effects can be achieved:
[0024] The present invention proposes a method for SEEG and audio contrastive learning modality matching (SEEG and Audio Contrastive Matching, SACM), which comprehensively utilizes the features of two modalities for Chinese Mandarin speech decoding. The present invention innovatively uses the features of experimentally synchronized audio to assist the training of the SEEG feature extractor. This method extracts the features of the corresponding modality data through a trainable brain signal feature extractor and a pre-trained speech feature extractor respectively, and uses the contrastive learning method to establish a modality matching pattern between relevant samples. Taking the information noise contrast estimation function as the contrastive learning objective function, it iteratively reduces the distance between the anchor feature and the positive sample feature in the feature space, and increases the distance between the anchor feature and the negative sample feature, realizing the modality matching of SEEG and audio corresponding samples in the feature space, improving the recognition accuracy of the speech decoding system, and effectively decoding Chinese Mandarin. Description of the Drawings
[0025] Figure 1 It is a schematic flowchart of SACM proposed by the present invention in the training stage.
[0026] Figure 2 It is a schematic diagram of the method of SACM test stage of the present invention. Detailed Embodiments
[0027] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0028] The present invention provides a speech decoding method based on contrastive learning modality matching, including:
[0029] Training stage:
[0030] S1, construct a data set; the data set includes original SEEG data and corresponding audio data;
[0031] S2, input the SEEG data into the SEEG feature extractor after random initialization to obtain SEEG features; input the audio data corresponding to the SEEG data into the pre-trained audio feature extractor to obtain audio features; wherein, the dimensions of the SEEG features and the audio features are the same, and the SEEG feature extractor is a residual convolutional neural network;
[0032] S3. Respectively mark each SEEG data as the first anchor point, mark the corresponding audio data as the first positive sample, and mark other audio data as the first negative sample. Using the information noise contrast estimation function as the objective function, calculate the source loss between each audio feature and SEEG feature; respectively mark each audio data as the second anchor point, mark the corresponding SEEG data as the second positive sample, and mark other SEEG data as the second negative sample. Using the information noise contrast estimation function as the objective function, calculate the symmetric loss between each audio feature and SEEG feature; use the sum of the source loss and the symmetric loss as the total loss to train the SEEG feature extractor;
[0033] Specifically,
[0034] S31. Divide the SEEG data and the corresponding audio data into batches. Input each batch of SEEG samples into the SEEG feature extractor in turn, and input the audio samples into the audio feature extractor in turn to extract features of the same dimension; the audio feature extractor is an open-source pre-trained model, and this type of model usually combines a convolutional neural network and an attention mechanism to extract the audio features of audio data;
[0035] S32. For each SEEG sample within a batch, mark it as the anchor point, mark the corresponding audio sample as the positive sample, and mark the remaining audio samples in the same batch as the negative samples;
[0036] S33. For the current batch of SEEG samples, use the information noise contrast estimation (InfoNCE) function as the objective function to calculate the source loss;
[0037] S34. For each audio sample within a batch, mark it as the anchor point, mark the corresponding SEEG sample as the positive sample, and mark the remaining SEEG samples in the same batch as the negative samples;
[0038] S35. For this batch of audio samples, use the InfoNCE function as the objective function to calculate the symmetric loss;
[0039] S36. Sum the source loss and the symmetric loss of this batch as the total loss, calculate the gradient and backpropagate to update the computational model of the SEEG feature extractor;
[0040] S37. Repeat steps S32 - S36 between batches for iterative optimization until the computational model of the SEEG feature extractor converges or reaches the maximum number of iteration rounds;
[0041] Application stage:
[0042] Input the SEEG data to be tested into the trained SEEG feature extractor to obtain SEEG features, and input all candidate audio data into the audio feature extractor to obtain candidate audio features;
[0043] Based on the similarity between the SEEG features and all candidate audio features, the matching audio result is obtained to complete speech decoding.
[0044] The present invention also provides a Chinese Mandarin speech decoding method based on contrastive learning modality matching, including:
[0045] Input the SEEG data to be recognized into the SEEG feature extractor trained in the training stage of the above method to obtain the corresponding SEEG features. Input all candidate audio data into the audio feature extractor to obtain candidate audio features. Based on the similarity between the SEEG features and all candidate audio features, the matching audio result is obtained to complete speech decoding. Wherein, the audio data is the audio data when the subject reads the Chinese Mandarin target word, and the SEEG data is the electroencephalogram data when the subject reads the Chinese Mandarin target word.
[0046] It can be understood that the speech decoding method provided by the present invention can be used for speech decoding of any type of language. When training the SEEG feature extractor, the corresponding language type can be adopted for the SEEG features and audio features.
[0047] An embodiment of the present invention provides a Chinese Mandarin speech decoding method based on contrastive learning modality matching. The algorithm training stage process is as Figure 1 shown, including:
[0048] Synchronously collect SEEG and audio data. Given an experimental target word library containing K target words, synchronously collect the SEEG data X ∈ R C×T and the audio data Y ∈ R 2×T , C represents the number of channels of the SEEG acquisition device, T represents the number of sampling points, and Y is a stereo audio;
[0049] Data preprocessing. For the SEEG data X, first manually check and replace the unavailable channels, and then perform a preprocessing process of detrending, rereferencing, 70 - 170Hz bandpass filtering, 50Hz and its harmonic notch filtering, robust scaling, Hilbert transform to extract the envelope, and downsampling to 200Hz. For the audio data Y, first merge it into monophonic speech data, and then perform downsampling. After preprocessing, segment the SEEG and audio according to the time stamps recorded in the experiment to form N pairs of SEEG and audio sample pairs where N = K × m, m represents the number of blocks of data collected in the experiment, K represents the number of vocabulary in the target word library. Since the subject reads each target word once in each block in the experiment, K is also the number of samples collected in each block. X i ∈ R C×T represents the i-th SEEG sample, and Y iDenote the corresponding \(i\)-th audio sample;
[0050] Select an open-source pre-trained large speech model \(f\) audio , such as HuBERT, wav2vec, etc., which can extract good speech features;
[0051] Input the training set audio sample \(Y\) i into \(f\) audio , generating a \(d\)-dimensional feature representation \(A\) i \(\in \mathbb{R}\) d , that is:
[0052] \(A\) i \(= f\) audio (Y i );
[0053] Randomly initialize the convolutional architecture SEEG feature extractor \(f\) seeg ;
[0054] Input the training set SEEG sample \(X\) i into \(f\) seeg , generating a \(d\)-dimensional feature representation \(B\) i \(\in \mathbb{R}\) d , that is:
[0055] \(B\) i \(= f\) audio (Y i );
[0056] In the training process, use mini-batch gradient descent. Each batch of data contains \(n\) pairs of samples. For each SEEG sample \(X\) i , mark it as the contrastive learning anchor sample, and the corresponding audio sample \(Y\) i is marked as the positive sample, and the remaining \(n - 1\) audio samples in the same batch are marked as negative samples. Calculate the contrastive learning loss for all SEEG samples in this batch, reduce the distance between the positive sample feature and the anchor feature in the \(d\)-dimensional feature space, increase the distance between the negative sample feature and the anchor feature, and obtain the average loss within the batch, denoted as the source loss:
[0057]
[0058] where \(\tau\) represents the contrastive learning temperature coefficient hyperparameter, and \(\text{sim}(B\) i , \(A\) i ) represents the cosine similarity between the SEEG feature \(B\) i and the audio feature \(A\) i , that is:
[0059]
[0060] Change the anchor samples in the source loss to the audio samples within the same batch, mark the corresponding SEEG samples as positive samples, and mark the remaining SEEG samples within the same batch as negative samples. Perform symmetric operations to enhance the modality matching effect, that is, calculate the symmetric loss within the batch:
[0061]
[0062] Calculate the total contrastive learning loss for this batch, that is, sum the source loss and the symmetric loss:
[0063]
[0064] Calculate the gradient based on the total loss to update the computational model of the SEEG feature extractor.
[0065] Repeat the training process iteratively between batches until the computational model of the SEEG feature extractor converges or reaches the maximum number of iterations to obtain the SEEG feature extractor f seeg 。
[0066] In the testing phase, as Figure 2 shown, for the given SEEG sample X to be decoded, use the trained f seeg to extract features and perform modality matching based on the similarity between features:
[0067] Input the audio sample Y i of the test set into the pre-trained audio feature extractor f audio to generate the feature representation f audio (Y i );
[0068] Input the SEEG sample X to be tested into f seeg to generate the feature representation f seeg (X);
[0069] Perform modality matching based on the similarity between features, that is, select the speech segment corresponding to the audio feature with the highest similarity to the SEEG feature as the decoding result through the retrieval function f retrieval :
[0070]
[0071] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A speech decoding method based on contrastive learning modality matching, characterized in that, Including: Training stage: S1. Construct a data set; the data set includes original SEEG data and corresponding audio data; S2. Input the SEEG data into the SEEG feature extractor after random initialization to obtain SEEG features; Input the audio data corresponding to the SEEG data into the pre-trained audio feature extractor to obtain audio features; wherein, the dimensions of the SEEG features and the audio features are the same, and the SEEG feature extractor is a residual convolutional neural network; S3. Mark each SEEG data as the first anchor point, mark the corresponding audio data as the first positive sample, and mark other audio data as the first negative sample, and calculate the source loss between each SEEG feature and audio feature with the information noise contrast estimation function as the objective function; mark each audio data as the second anchor point, mark the corresponding SEEG data as the second positive sample, and mark other SEEG data as the second negative sample, and calculate the symmetric loss between each audio feature and SEEG feature with the information noise contrast estimation function as the objective function; use the sum of the source loss and the symmetric loss as the total loss to train the SEEG feature extractor; Application stage: Input the SEEG data to be tested into the trained SEEG feature extractor to obtain SEEG features, and input all candidate audio data into the audio feature extractor to obtain candidate audio features; Obtain the matching audio result according to the similarity between the SEEG features and all candidate audio features, and complete speech decoding.
2. The method according to claim 1, characterized in that Before step S2, it also includes preprocessing the original SEEG data and the corresponding audio data, and the preprocessing includes: detrending the SEEG data, re-referencing, band-pass filtering at 70 - 170 Hz, notch filtering at 50 Hz and its harmonics, robust scaling, extracting the envelope by Hilbert transform, and downsampling to 200 Hz; merging the audio data into monophonic speech data and then performing downsampling.
3. The method according to claim 1, wherein The source loss is: where sim(B i , A i ) represents the cosine similarity between the SEEG feature B i and the audio feature A i , that is: The symmetric loss is: where sim(A i , B i ) represents the cosine similarity between audio feature A i and SEEG feature B i , that is: Where τ represents the contrast learning temperature coefficient hyperparameter, and n is the number of sample pairs per batch.
4. The method according to claim 1, wherein In the application stage, the speech segment corresponding to the audio feature with the highest similarity to the SEEG feature among all candidate audio features is used as the decoding result: Among them, Y represents the speech segment as the decoding result, X represents the SEEG data to be tested, and Y i represents the i-th candidate audio data, and f seeg represents the SEEG feature extractor, and f audio represents the audio feature extractor, sim represents the calculation of cosine similarity, and τ represents the contrastive learning temperature coefficient hyperparameter.
5. The method according to claim 1, characterized in that, The audio feature extractor is a HuBERT or wav2vec model.
6. A method for decoding Mandarin Chinese speech based on contrastive learning modality matching, characterized in that, Including: Input the SEEG data to be recognized into the SEEG feature extractor trained in the training stage of the method according to any one of claims 1 - 5 to obtain the corresponding SEEG features, input all candidate audio data into the audio feature extractor to obtain candidate audio features, and obtain the matching audio result according to the similarity between the SEEG features and all candidate audio features, and complete speech decoding; wherein, the audio data is the audio data when the subject reads the Chinese Mandarin target word, and the SEEG data is the stereo electroencephalogram data when the subject reads the Chinese Mandarin target word.
7. An electronic device, characterized in that, Including: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the method according to any one of claims 1-5 or claim 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to execute the method according to any one of claims 1-5 or claim 6.
9. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, the method according to any one of claims 1-5 or claim 6 is implemented.