A brain-computer interface system for recognizing the intention of Chinese oral language based on a sound-meaning integration double model
By collecting neural signals from different areas of the brain and constructing speech and semantic decoders, combined with sEEG technology, the problems of homophones and similar-sounding words in spoken Chinese were solved, achieving accurate recognition and reconstruction of spoken Chinese intentions and improving the communication ability of ALS patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV
- Filing Date
- 2025-11-21
- Publication Date
- 2026-07-03
Smart Images

Figure CN121560160B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical engineering technology, and relates to language brain-computer interface technology, and in particular to a Chinese spoken intention recognition brain-computer interface system based on a sound-meaning integrated dual model. Background Technology
[0002] The sixth national census revealed that there are 1.3 million people with speech disabilities in my country. Major central nervous system diseases such as stroke, traumatic brain injury, and amyotrophic lateral sclerosis (ALS) can lead to locked-in syndrome (LIS), a severe disorder affecting speech and movement. A survey of LIS in my country showed that regardless of whether they have family caregivers, over half of the patients wish to continue living. There are approximately 200,000 ALS patients in my country, and the 10-year survival rate after diagnosis exceeds 80%. Currently, there is no specific and effective treatment for ALS, meaning that even with preserved consciousness and cognitive activity, ALS patients face the gradual loss of speech and motor abilities. Caregivers struggle to identify the needs of ALS patients, leading to a loss of social identity and a state of isolation.
[0003] Language-based brain-computer interfaces (BCIs) decode neural signals to form commands, bypassing paralyzed muscles and establishing electronic neural bypasses to help restore language interaction, significantly improving the quality of life for ALS patients. Currently, several studies have decoded low-order processes in language production—the articulation and phonation processes—using neural signals from motor and speech-related cortices. The basic unit of an English word is a phoneme; directly decoding phonemes to form words significantly improves communication efficiency compared to directly decoding the entire word's pronunciation. In 2023, teams from the University of California, San Francisco, and Braingate achieved clinical trial results of 78 words / minute and 62 words / minute, respectively, by decoding phonemes to form words (normal adult spoken English communication is approximately 150 words / minute). In the aforementioned BCI systems, achieving rapid speech synthesis efficiency relies on breaking down the decoding of English words into decoding at the lower-order phoneme level.
[0004] However, Chinese has significant differences from English in many attributes. Typical thorny problems include a large number of homophones, which makes it difficult to simply decompose Chinese into "phonemes" for efficient and fast speech synthesis. At present, directly migrating the most advanced English language brain-computer interface system (for decoding and predicting oral articulation movements) to speech synthesis for Chinese oral expression is expected to encounter difficulties due to the unique language phenomena in Chinese. Moreover, in the population with motor disorders, there are many unknown problems regarding the significant, robust, and accurate representation of the cerebral cortex for model recognition.
[0005] To address the above problems, a large amount of research work has been carried out in China for the recognition of Chinese oral intentions. In 2023, Huashan Hospital affiliated to Fudan University, together with Shanghai University of Science and Technology and Tianjin University, conducted neural decoding through high-density ECoG to explore the possibility of "tone + syllable" synthesis. By separately decoding tone information and syllable information, Chinese speech was generated through combination. The average classification accuracy of this model for "tone - syllable" of a single subject reached 75.6%, and the highest accuracy could reach 91.4%. However, there are a large number of Chinese words with exactly the same tones and syllables. For example, for the exactly the same speech pronunciation " / ting2 / ", it can represent multiple Chinese characters with different semantics, including "停 (stop)", "庭 (courtyard)", "亭 (pavilion)"; and the adjustment of tones and the detailed changes of initials and finals will also produce Chinese characters with pronunciations similar to " / ting2 / ", such as "听 / ting1 / ", "盯 / ding1 / ", "顶 / ding3 / " and so on. There are also a large number of similar pronunciation phenomena for disyllabic words, such as "好看 / hao3kan4 / " and "好暗 / hao3an4 / ". The occurrence frequency of these Chinese words with similar pronunciations in expressions is not small. When faced with these words, the combination of "tone + syllable" is difficult to work completely. Summary of the Invention
[0006] Aiming at the problems existing in the above-mentioned existing language brain-computer interfaces, the present invention aims to propose a brain-computer interface system for recognizing Chinese oral intentions based on a dual model of sound-meaning integration, providing an effective interaction tool for patients with severe speech and motor disorders such as ALS.
[0007] Taking the spoken expression of Chinese as the core, different Chinese words can exhibit the following characteristics in terms of pronunciation and semantics: (1) Homophonic or near-homophonic words have different semantic contents. For example, "妈 (mā)" is homophonous with "抹 (mǒ)", and "姨 (yí)" is homophonous with "移 (yí)"; and (2) Words with the same or similar semantics have different pronunciations. For example, "抹 (mǒ)" and "移 (yí)", both representing the execution of a certain action, and "妈 (mā)" and "姨 (yí)", both representing kinship appellations. In the neural signal pattern, the neural representations in the brain's articulatory motor coding area retain similar neural signal characteristics for the above situation (1); the neural representations in the brain's semantic concept organization coding area retain similar neural signal characteristics for the above situation (2), which makes it possible to generate Chinese speech based on the combination of two types of information, "pronunciation" and "semantics".
[0008] In the brain-computer interface system for Chinese spoken language intention recognition described in this invention, with the "sound-meaning" integration as the central idea, through the language phenomena of "same pronunciation but different meanings" and "same meaning but different pronunciations", electrodes are covered in the brain's articulatory motor coding area and the brain's semantic concept organization coding area respectively to obtain the neural signal records of the same spoken expression in the above two different functional areas, and "speech decoders" and "semantic decoders" are constructed respectively. Based on the "speech decoders" and "semantic decoders", subsequently, a "Chinese word and phrase sound-semantic fusion synthesizer" is completed. In addition, based on the actual process of an individual completing spontaneous language organization and pronunciation, a "vocalization start decoder" also needs to be formed.
[0009] Based on the similarities and differences of different spoken words and phrases in the above speech decoders and semantic decoders, the problems of sparse and similar-confused speech information of Chinese spoken words and phrases are solved. In the recognition of spoken information - such as "妈 (mā)", the neural pattern of " / ma1 / " is obtained in the speech decoder, and the neural pattern of "kinship appellation" is obtained in the semantic decoder; in the recognition of "抹 (mǒ)", although there is the same neural pattern as "妈 (mā)" (" / ma1 / ") in the speech decoder, in the semantic decoder, it belongs to a semantic concept different from "kinship appellation" ("limb movement"). Through the combination of the above two models of "speech decoder" and "semantic decoder", the accurate locking and reconstruction of Chinese spoken language intention can be completed.
[0010] This invention proposes a "sound-meaning" integration scheme for spoken Chinese, enabling a highly efficient Chinese speech synthesis framework. However, there are currently gaps in our understanding of the neural processing mechanisms involved in the convergence and processing of Chinese speech and semantic information to ultimately form articulation output. Therefore, this invention employs stereotactic electroencephalography (sEEG) instead of traditional signal acquisition methods such as EEG and ECoG. sEEG electrodes can be implanted inside the brain, directly contacting brain tissue, thus providing more precise spatial resolution, higher signal amplitude, and a signal-to-noise ratio—approximately 10 times higher than EEG—and eliminating motion artifacts caused by speech and movement. sEEG can also acquire signals from multiple brain regions, capturing spatiotemporal transitions within the network at a more refined neuroanatomical (millimeter) and temporal (millisecond) scale. In other words, sEEG technology can accurately analyze and map the interaction between speech and semantic information, as well as the spatiotemporal dynamics of autonomous articulation in a semantic context. This reveals the key role of higher-order lexical semantic representation in the convergence of sound-meaning information and its driving of near-synonymous autonomous expression. It is beneficial to the speech synthesis practice of language brain-computer interfaces, improves the degree of speech discrimination in language systems such as Chinese, Japanese, and Korean, which have a large number of homophones and polyphones, and enhances the understanding of information interaction between different brain regions in the "sound-meaning-articulation" network.
[0011] The present invention provides the following specific solutions:
[0012] In a first aspect, the present invention provides a brain-computer interface system for recognizing spoken Chinese intentions based on a sound-meaning integrated dual model, comprising:
[0013] Electrodes were placed in the articulation motor coding area and the semantic concept organization coding area of the brain, respectively, to collect neural signal records of the same spoken expression in the two different functional areas.
[0014] A vocal initiation decoder is used to decode vocal initiation information from the brain's articulation and semantic concept organization response area before the actual vocal output;
[0015] A speech decoder is used to decode Chinese speech information from the articulation motor coding area of the brain;
[0016] A semantic decoder is used to decode semantic information from the coding areas of the brain's semantic concept organization.
[0017] The Chinese character-word-speech-semantic fusion synthesizer is used to integrate the above-mentioned Chinese speech and semantic information to accurately locate and reconstruct the intention of spoken Chinese.
[0018] In some embodiments, the Chinese spoken intention recognition brain-computer interface system based on the sound-meaning integration dual model includes the acquisition of neural signals distributed in different anatomical spatial functional coding regions of the brain via stereotactic intracranial electroencephalography (sEEG).
[0019] In some embodiments, the Chinese spoken intention recognition brain-computer interface system based on the sound-meaning integration dual model includes a vocalization initiation decoder to analyze whether an individual intends to pronounce words, analyzes the pronunciation intention information before the actual pronunciation output is generated, extracts the sEEG channels of the brain that have different responses to whether pronunciation is generated within a predictable number of milliseconds, and performs feature extraction.
[0020] In some embodiments, the Chinese spoken intention recognition brain-computer interface system based on the sound-meaning integration dual model includes a speech decoder to analyze the speech information in an individual's pronunciation, extracting sEEG channels in which the brain responds differently to different articulation motor structures, and performing feature extraction.
[0021] In some embodiments, the Chinese spoken intention recognition brain-computer interface system based on the sound-meaning integration dual model includes a semantic decoder to analyze the semantic information in an individual's pronunciation, extracting sEEG channels in which the brain responds differently to different semantic categories in spoken expression, and performing feature extraction.
[0022] In some embodiments, the Chinese spoken intention recognition brain-computer interface system based on the sound-meaning integration dual model includes a Chinese word-speech-semantic fusion synthesizer for integrating features from the speech decoder and the semantic decoder in a unified representation space to achieve more stable multi-level decoding results, and finally generating accurate Chinese words and outputting them in speech or text form.
[0023] In some embodiments, the Chinese spoken intention recognition brain-computer interface system based on the sound-meaning integration dual model includes acquiring signals in the key regions of the frontal lobe and the key regions of the temporal lobe, wherein the key regions of the frontal lobe include: the left inferior frontal gyrus, which includes the orbital part and the triangular part, the tegmentum and the premotor cortex and the lower part of the primary motor cortex;
[0024] The key areas of the temporal lobe are divided into three main parts: the anterior part, namely the anterior temporal lobe; the lateral part, namely the dorsolateral temporal lobe, which includes the posterior part of the left superior temporal gyrus, the anterior part of the superior temporal gyrus, and the middle temporal gyrus; and the basotemporal language area, which includes the anterior and middle parts of the inferior temporal gyrus, the anterior part of the occipital-temporal gyrus, and the anterior part of the parahippocampal gyrus.
[0025] Feature extraction was performed on the high-γ frequency bands of electrode channels in key regions of the frontal lobe and key regions of the temporal lobe, and a target decoder was constructed. The target decoder includes a vocal priming decoder, a speech decoder, and a semantic decoder.
[0026] Secondly, the present invention provides a training paradigm for autonomously organizing spoken vocabulary expression in an audiovisual-induced scenario, the steps of which include:
[0027] (1) The subject reads the task instructions on the screen and presses the space bar to start the experiment. At the same time, the sound of pressing the space bar will be recorded by the microphone to re-align the audio recording time with the system and eliminate time errors within 500ms.
[0028] (2) Then the formal training begins. A cross is displayed on the screen to guide the visual focus. The subjects receive different forms of prompts. Then the red light is on for 2 seconds to give the subjects time to think so that they can complete the spontaneous word formation when the green light is on. Then the green light is on for 3.5 seconds, and the subjects can complete the expression after the green light is on.
[0029] (3) After the green light disappears, rest for 1.5 seconds, and then start the next cycle.
[0030] In some embodiments, the training paradigm for word association and generation through audiovisual induction also includes training involving three tasks, wherein the first task is listening and forming words. Each individual receives cues through auditory means, i.e., through sound playback in headphones. The neural signal record of the individual's word search, extraction and speech generation process under the guidance of the target speech cues is established, and the desktop microphone records the individual's vocal information.
[0031] Task 2 involves forming words from images. Each individual receives visual cues by viewing text on a screen. The task involves recording neural signals of the individual's process of word search, retrieval, and speech generation under the guidance of the target word. A desktop microphone records the individual's vocal information.
[0032] Task 3 involves word association. Each participant receives visual cues for a specific semantic category. Within that category, the participant spontaneously associates and generates words. The entire process is recorded by neural signals and a desktop microphone.
[0033] In some embodiments, the training paradigm for word association and generation through audiovisual induction also includes forming a speech decoder in Task 1 and Task 2, and forming a semantic decoder in Task 3, since the focus is on word generation after the individual makes associations based on semantic categories.
[0034] Task 1, Task 2, and Task 3 all involve vocabulary generation. The neural signals generated during the vocabulary generation process in each of Task 1, Task 2, and Task 3 are used to form a vocal priming decoder. The vocal priming decoder detects and tests the individual's intention to make a sound.
[0035] The neural signals generated during the vocabulary generation process in Task 1, Task 2, and Task 3 are also used to form a Chinese word-speech-semantic fusion synthesizer. The Chinese word-speech-semantic fusion synthesizer integrates the output information of the trained speech decoder and semantic decoder, and then performs spoken expression to achieve accurate recognition and output of spoken Chinese intentions.
[0036] The Chinese spoken intention recognition brain-computer interface system described in this invention is based on a dual model of Chinese "sound-meaning integration," which can solve the problems of sparse word and sound information and confusion of similar words in spoken Chinese, thereby achieving accurate locking and reconstruction of spoken intention. The training paradigm described in this invention has two different forms of input: visual and auditory. After the subject sees the Chinese characters displayed on the screen or hears the speech in the headphones, the individual is inspired to autonomously select appropriate words for spoken expression based on the characters or pronunciation. It also has a third form of semantic association; by setting an expression scenario, the individual is inspired to autonomously select all suitable words that can be associated with the semantic content based on that scenario for spoken expression, closely resembling real-life language use. Attached Figure Description
[0037] Figure 1 This is a flowchart of the training paradigm of the present invention;
[0038] Figure 2 This invention is a speech initiation decoder based on convolutional neural networks and bidirectional long short-term memory networks;
[0039] Figure 3 This invention is a speech decoder based on convolutional neural networks and Transformer encoders;
[0040] Figure 4 This invention provides a semantic decoder based on a layered Transformer and a multimodal attention network.
[0041] Figure 5 This invention relates to a Chinese character-speech-semantic fusion synthesizer based on a cross-modal feature fusion network. Detailed Implementation
[0042] The specific embodiments of the present invention will be described below with reference to the accompanying drawings, and the technical solutions in the specific embodiments of the present invention will be clearly and completely explained. Obviously, the specific embodiments described below are only a part of the multiple preferred embodiments of the present invention. Based on the specific embodiments listed in the present invention, all other embodiments obtained by those skilled in the art without creative effort should also fall within the scope of protection of the present invention.
[0043] To address the issues of similar pronunciations in spoken Chinese and the ambiguity in spoken expression among people with illnesses, this invention designs a training paradigm for autonomously organizing spoken vocabulary in an audiovisually induced scenario. Previous studies on speech decoding often pre-specified the specific content spoken by the subjects, while this invention allows subjects to autonomously associate, select, and generate vocabulary within a certain range, thus more closely resembling real-life language use.
[0044] (a) The inclusion criteria for the subjects are as follows:
[0045] 1) Age 21-45;
[0046] 2) Undergo stereotactic intracranial electrode implantation;
[0047] 3) Native Chinese speaker, fluent and proficient in Mandarin;
[0048] 4) Six years or more of education (primary school graduation);
[0049] 5) Hearing and vision, or corrected hearing and vision, are normal;
[0050] 6) Intracranial electrode implantation sites include the frontal and temporal lobes;
[0051] 7) Participate in this study voluntarily and sign informed consent.
[0052] (II) Training Paradigm Process as follows Figure 1 As shown.
[0053] First, before the formal training begins, participants read the task instructions on the screen and press the space bar to start the experiment. Then, the formal experiment begins, with the "cross" on the screen serving as a visual focus guide. Next, in the "cue" phase, participants receive prompts either visually or aurally, depending on the task. A red light illuminates for 2 seconds, during which participants think about the prompts received visually or aurally, so that they can quickly and effectively generate a word for verbal expression when the green light illuminates. Participants complete their expression within the 3.5 seconds the green light is on, then rest for 1.5 seconds before starting the next cycle.
[0054] The training involved three tasks, which could be presented with visual or auditory cues. All three tasks required participants to choose appropriate spoken Chinese vocabulary and express it.
[0055] The main difference between the three tasks lies in the content of the prompts provided.
[0056] Specifically, Task 1 and Task 2 are tasks of forming words according to prompt clues. Among them, Task 1 is a task of forming words by listening (auditory input). That is to say, the prompts received by the subject in each trial are played through noise-canceling headphones. Then, the subject makes oral expressions according to the received voice information. For example, for the prompt voice / yi1 / , the subject can spontaneously choose to express words such as "clothes", "doctor", or "one".
[0057] Task 2 is a task of forming words by looking at characters (visual input). That is to say, the prompts received by the subject in each trial are received through the visual Chinese characters displayed on the computer screen. Then, the subject makes oral expressions according to the received visual text information. For example, for the prompt character "医", the subject can spontaneously choose to express words such as "doctor", "hospital", or "heal".
[0058] Task 3 is a vocabulary retrieval task, which requires the subject to retrieve corresponding vocabulary according to the prompted semantic attributes, inspiring the individual to autonomously select all suitable vocabulary that can be联想到 based on this semantic scenario and make oral expressions. After each green light亮起 after the prompt, the subject selects an oral vocabulary that meets the requirements of the prompt according to the requirements. For example, if the prompted semantic content is "fruit", each time the green light appears, the subject spontaneously selects a vocabulary belonging to the fruit content, such as "apple", "banana", "grape", etc. for expression.
[0059] During the process of the individual's autonomous oral expression, neural signal recordings are simultaneously performed in multiple functionally specific target areas of the brain, including identifying speech-dependent neural patterns and constructing a speech decoder: The functional areas of the brain that are sensitive to articulation processing should have similar neural response patterns to oral vocabulary with similar articulation movements, that is, similar pronunciations, or be sensitive to articulation movement differences, that is, have significantly different neural responses to oral vocabulary with large pronunciation differences.
[0060] And identifying semantic concept-dependent neural patterns and constructing a semantic decoder: The functional areas of the brain that are sensitive to semantic processing should respond to semantic similarity, that is, have similar neural response patterns to oral vocabulary with similar semantic information in oral vocabulary, or be sensitive to semantic content differences, that is, have significantly different neural responses to oral vocabulary with large semantic content differences.
[0061] The neural signal acquisition method selected by the present invention can be achieved by means of stereotactic intracranial electroencephalogram (sEEG) in order to obtain neural signals distributed in different anatomical space functional coding regions of the brain.
[0062] This invention preprocesses sEEG neural signals and targets key brain regions that exhibit significant responses to different functional modes, including specific regions in the frontal and temporal lobes. The key frontal lobe regions include the left inferior frontal gyrus, encompassing the orbital and triangular portions, the tegmentum, and the premotor cortex and lower primary motor cortex. The key temporal lobe regions are divided into three parts: the anterior part (anterior temporal lobe); the lateral part (dorsolateral temporal lobe), including the posterior part of the left superior temporal gyrus, the anterior part of the superior temporal gyrus, and the middle temporal gyrus; and the basotemporal language region, including the anterior and middle parts of the inferior temporal gyrus, the anterior part of the occipitotemporal gyrus, and the anterior part of the parahippocampal gyrus. Features are extracted from the high-γ frequency bands of the electrode channels in the key frontal and temporal lobe regions, and a target decoder is constructed. This target decoder includes a vocal priming decoder, a speech decoder, and a semantic decoder.
[0063] Regarding the sound-initiating decoder described in this invention:
[0064] In embodiments of the present invention, a speech priming decoder is used to detect speech-related feature patterns from neural activity signals and to distinguish between speech generation-related activities and non-speech activities. For example... Figure 2 As shown, the vocal priming decoder employs a CNN-BiLSTM structure, balancing the modeling capabilities of local spatial features and temporal dynamic information. The vocal priming decoder consists of the following modules: a one-dimensional convolutional neural network (1D-CNN), group normalization, rectified linear units (ReLU), max pooling, a bidirectional long short-term memory neural network (BiLSTM), attention pooling, layer normalization, and a fully connected layer (FC). Specifically, the vocal priming decoder first uses a one-dimensional convolutional network (1D-CNN) to extract short-time frequency domain features of neural signals, and enhances its robustness through group normalization and max pooling. Subsequently, a Squeeze-and-Excitation (SE) channel attention mechanism is introduced to adaptively weight the importance of different channels, thereby improving the discriminative power of acoustic features. In the temporal modeling part, a bidirectional long short-term memory network (BiLSTM) is used to capture temporal dependencies, and attention-weighted pooling is combined to achieve weighted feature aggregation. Finally, a fully connected layer is used for binary classification to detect speech segments and distinguish between non-speech segments.
[0065] In one instance, this invention performed real-time sliding window-based speech detection on five subjects using the aforementioned vocal activation decoder. The test results for the five subjects are as follows:
[0066] Table 1 shows the vocal priming recognition performance of the subjects in this example.
[0067]
[0068] In Table 1, we use AUC, ACC, and F1 to evaluate the performance of the decoding algorithm. AUC refers to the area under the ROC curve, which is used to calculate the true positive rate and false positive rate. Generally, an AUC > 0.8 indicates that the sound-starting decoder has good decoding ability. ACC refers to (number of trials correctly predicted by the sound-starting decoder) / (total number of trials), reflecting the overall accuracy of the sound-starting decoder. The F1 score is the harmonic mean of precision and recall. Precision is "how many of the trials predicted as targets are true targets"; recall is "how many of the trials with true targets are successfully predicted", referring to the ability of the sound-starting decoder to identify true targets.
[0069] Regarding the speech decoder described in this invention:
[0070] A speech decoder is used to map detected neural speech activity into specific phoneme sequences or word pronunciation sequences. In one implementation, such as... Figure 3 As shown, the speech decoder employs a multi-layer two-dimensional convolutional neural network (2D-CNN) combined with a Transformer decoding structure to capture temporal-frequency distribution features and global contextual dependencies. The speech decoder consists of the following modules: a 2D convolutional neural network (2D-CNN), batch normalization, rectified linear units (ReLU), dropout, layer normalization, a Transformer encoder, a Transformer decoder, and a fully connected layer (FC) + Softmax function. Specifically, the input neural signal features are first processed by the 2D-CNN to extract local acoustic texture information, and then layer normalization and residual connections are used to maintain feature stability. Subsequently, the Transformer encoder uses a multi-head self-attention mechanism to achieve global modeling across time steps to capture the dynamic evolution patterns of speech generation. Finally, the Transformer decoder outputs the phoneme probability distribution through a fully connected layer, achieving fine-grained reconstruction of continuous speech. In certain instances, the speech decoder can recognize speech pronunciation sequences (such as “yī wù”), but cannot distinguish their semantic categories; that is, it cannot distinguish between “clothing” and “medical” in “yī wù”.
[0071] Regarding the semantic decoder described in this invention:
[0072] A semantic decoder is used to further identify the semantic category or conceptual attribute of the tested word formation based on the speech decoding results, thereby achieving semantic-level differentiation. In one implementation, such as... Figure 4 As shown, the semantic decoder employs a layered Transformer + multimodal attention network structure. The semantic decoder consists of the following modules: a linear projection layer (LinearProjector), a Transformer encoder, a multi-head attention mechanism, a fully connected layer (FC), and a Softmax function. Specifically, the feature embeddings (phoneme-level or word-level embeddings) output by the speech decoder are first mapped to the semantic embedding space, and then input to the Transformer encoder layer through linear projection and positional encoding. Subsequently, the semantic decoder utilizes the multi-head attention mechanism to capture the association between speech features and potential semantic categories. For classification output, a Softmax multi-class classification layer can be used to output semantic category probabilities (such as "clothing" or "medical"). In an optional embodiment, to enhance the recognition of fine-grained semantics, the semantic decoder can also introduce pre-trained semantic decoder embeddings (such as BERT or ERNIE) to re-encode the semantic context of candidate words.
[0073] Regarding the Chinese word-speech-semantic fusion synthesizer described in this invention:
[0074] The Chinese character-word-speech-semantic fusion synthesizer is used to integrate features from the speech decoder and semantic decoder in a unified representation space to achieve more stable multi-level decoding results. For example... Figure 5 As shown, the Chinese character-word-speech-semantic fusion synthesizer can employ a cross-modal fusion network (CFFN), which consists of a multi-layer interactive attention structure. The Chinese character-word-speech-semantic fusion synthesizer comprises the following modules: a bi-directional cross-attention layer, a multi-layer perceptron fusion layer (MLP fusion layer), layer normalization, a fully connected layer (FC), and a softmax function.
[0075] Specifically, the Chinese character-word speech-semantic fusion synthesizer first normalizes and linearly transforms the speech feature vector and semantic embedding vector respectively to unify the feature dimensions. Then, the speech stream and semantic stream are fused through a bidirectional cross-attention layer, where speech features are used as queries to calculate semantic attention distributions, and semantic features are used as queries to calculate speech attention distributions, achieving complementary alignment of information. The fused representation is then integrated through a multilayer perceptron (MLP) and residual connections, and can be used as input for downstream tasks (such as word sense recognition, concept classification, or semantic consistency verification). In some instances, to further enhance intermodal collaboration, the Chinese character-word speech-semantic fusion synthesizer can employ a gated fusion strategy or a feature weighting strategy. By learning variable fusion weights, the contribution ratio of speech and semantic features is dynamically adjusted, thereby achieving accurate joint decoding at the speech-semantic level.
[0076] In the training paradigm of this invention, Tasks 1 and 2 primarily involve autonomous word formation based on auditory and visual cues, which can be used to identify dominant brain regions and information in speech processing, thus benefiting the construction of speech decoders. Task 3 primarily involves word retrieval and expression based on semantic attributes, which can be used to identify dominant brain regions and information in semantic processing, thus benefiting the construction of semantic decoders. Furthermore, the neural signals generated during the word generation process in Tasks 1, 2, and 3 can also be used to construct a Chinese character-speech-semantic fusion synthesizer.
[0077] This invention conducted a simple word formation classification based on different pronunciations of single characters on 6 subjects. In this example, each subject formed words based on three pronunciations ("yī", "yào", "wò"). The number of trials collected is shown in Table 2.
[0078] Table 2 shows the dataset size for the three pronunciation combinations of the subjects in this example.
[0079]
[0080] In this example, a simplified speech decoder is used to classify and decode speech information and perform five-fold cross-validation. The preliminary results are shown in Table 3.
[0081] Table 3 shows the three decoding effects of the subjects on the speech words in this example.
[0082]
[0083] In some cases, by screening the significant response channels of the subjects during training, the following 10 dominant response brain regions were identified (thalamus, posterior middle temporal gyrus, insula, central tectal cortex, posterior inferior temporal gyrus, superior frontal gyrus, supplementary motor area, anterior middle temporal gyrus, putamen, and insular region of inferior frontal gyrus). Feature extraction was then performed on the brain region information.
[0084] Table 4 demonstrates that, among the six subjects, speech processing features and semantic processing features are distinguishable under the training paradigm of this invention. Sub-2 and Sub-6 exhibit the best classification performance, with average accuracies of 76.61% and 74.34%, respectively, significantly higher than the random level of 50%. This indicates that the "sound-meaning" integrated dual-model Chinese spoken intention recognition brain-computer interface system described in this invention is effective and efficient.
[0085] Table 4 shows the classification performance of the subjects on the processing forms of the speech and semantic advantage tasks in this example.
[0086]
Claims
1. A brain-computer interface system for recognizing spoken Chinese intent based on a sound-meaning integrated dual model, characterized in that... include: Electrodes were placed in the articulation motor coding area and the semantic concept organization coding area of the brain, respectively. The neural signals of the same spoken expression were recorded in the articulation motor coding area and the semantic concept organization coding area of the brain through listening to words and seeing words. Signals were acquired in key regions of the frontal lobe and the temporal lobe. The key regions of the frontal lobe include the left inferior frontal gyrus, which includes the orbital and triangular parts, the tegmentum, and the premotor cortex and the lower part of the primary motor cortex. The key areas of the temporal lobe are divided into three main parts: the anterior part, namely the anterior temporal lobe; the lateral part, namely the dorsolateral temporal lobe, which includes the posterior part of the left superior temporal gyrus, the anterior part of the superior temporal gyrus, and the middle temporal gyrus; and the basotemporal language area, which includes the anterior and middle parts of the inferior temporal gyrus, the anterior part of the occipital-temporal gyrus, and the anterior part of the parahippocampal gyrus. Feature extraction was performed on the high-γ frequency band of the electrode channels in key regions of the frontal lobe and key regions of the temporal lobe, and a target decoder was constructed. The target decoder includes a vocalization initiation decoder, a speech decoder, and a semantic decoder. The vocal priming decoder is used to pre-decode vocal priming information from neural signals generated by auditory word formation and visual word formation tasks in the brain's articulation and semantic concept organization response area before actual vocal output. Specifically, the vocal priming decoder utilizes a one-dimensional convolutional network to extract short-time frequency domain features of neural signals and enhances their robustness through group normalization and max pooling. A Squeeze-and-Excitation channel attention mechanism is introduced to adaptively weight the importance of different channels, thereby improving the discriminative power of acoustic features. In the temporal modeling part, a bidirectional long short-term memory network is used to capture temporal dependencies, and attention-weighted pooling is combined to achieve feature weighting and aggregation, followed by binary classification through a fully connected layer. A speech decoder is used to decode Chinese speech information from neural signals generated by tasks of listening to words and seeing characters, based on the brain's articulation motor coding area. A semantic decoder is used to decode semantic information from the coding areas of the brain's semantic concept organization. The Chinese character-word-speech-semantic fusion synthesizer is used to integrate the above-mentioned Chinese speech and semantic information to accurately locate and reconstruct the intention of spoken Chinese.
2. The Chinese spoken intention recognition brain-computer interface system based on a sound-meaning integrated dual model as described in claim 1, characterized in that, Data acquisition was achieved through stereotactic electroencephalography (sEEG) to obtain neural signals distributed in functional coding regions of different anatomical spaces in the brain.
3. The Chinese spoken intention recognition brain-computer interface system based on a sound-meaning integrated dual model as described in claim 1, characterized in that, The vocalization initiation decoder analyzes whether an individual intends to produce a sound and analyzes the intention to produce sound before the actual sound output occurs. It extracts the sEEG channels in the brain that show differential responses to whether a sound will be produced within a predictable number of milliseconds and performs feature extraction.
4. The Chinese spoken intention recognition brain-computer interface system based on a sound-meaning integrated dual model as described in claim 1, characterized in that, The speech decoder analyzes the speech information in an individual's pronunciation, extracts the sEEG channels in which the brain responds differently to different articulation motor structures, and performs feature extraction.
5. The Chinese spoken intention recognition brain-computer interface system based on a sound-meaning integrated dual model as described in claim 1, characterized in that, The semantic decoder analyzes the semantic information in an individual's pronunciation, extracts the sEEG channels that the brain responds differently to different semantic categories in spoken expression, and performs feature extraction.
6. The Chinese spoken intention recognition brain-computer interface system based on a sound-meaning integrated dual model as described in claim 1, characterized in that, The Chinese character-word speech-semantic fusion synthesizer is used to integrate features from the speech decoder and the semantic decoder in a unified representation space to achieve more stable multi-level decoding results, and finally generate accurate Chinese characters and words and output them in speech or text form.
7. A training method for word association and generation through audiovisual induction, in which individuals spontaneously process language and complete spoken word expression, applied to the Chinese spoken intention recognition brain-computer interface system based on the sound-meaning integration dual model as described in any one of claims 1 to 6, characterized in that the steps... include: (1) The subject reads the task instructions on the screen and presses the space bar to start the experiment. At the same time, the sound of pressing the space bar will be recorded by the microphone to re-align the audio recording time with the system and eliminate time errors within 500ms. (2) Next, the formal training begins. A cross is displayed on the screen to guide the visual focus. The subjects then receive different forms of prompts. A red light illuminates for 2 seconds to give the subjects time to think, so that they can complete spontaneous word formation when the green light illuminates. The green light then illuminates for 3.5 seconds, after which the subjects can complete the expression. The training involves three tasks. Task 1 is listening to words. Each individual receives prompts through the sound played in the headphones. The neural signal recording of the individual's word search, retrieval, and speech generation process under the guidance of the target speech prompts is established. The desktop microphone records the individual's vocal information. Task 2 is seeing words. Each individual sees words through the visual form. The first task involves receiving textual cues on a screen and recording neural signals during the individual's vocabulary search, retrieval, and speech generation process guided by a target word. A desktop microphone records the individual's vocalizations. The second task involves vocabulary association, where each participant receives visual cues of a specific semantic category. Within this category, the individual spontaneously associates and generates words, with neural signals and the desktop microphone recording the entire process. The third task involves word formation based on cues, while the fourth task is a vocabulary retrieval task. Participants are required to retrieve words based on the semantic attributes of the cues, encouraging them to autonomously select all suitable words that correspond to the semantic content for oral expression within a semantic context. (3) After the green light disappears, rest for 1.5 seconds, and then start the next cycle.
8. The training method for word association and generation through audiovisual induction as described in claim 7, characterized in that, In Task 1 and Task 2, it is used to train the speech decoder. In Task 3, since the focus is on word generation after individuals make associations based on semantic categories, it is used to train the semantic decoder. Task 1, Task 2, and Task 3 all involve vocabulary generation. The neural signals generated during the vocabulary generation process in each of Task 1, Task 2, and Task 3 are used to form a vocal priming decoder. The vocal priming decoder detects and tests the individual's intention to make a sound. The neural signals generated during the vocabulary generation process in Task 1, Task 2 and Task 3 are also used to form a Chinese word-speech-semantic fusion synthesizer. The Chinese word-speech-semantic fusion synthesizer integrates the output information of the trained speech decoder and semantic decoder, and then performs spoken expression to achieve accurate recognition and output of spoken Chinese intentions.
Citation Information
Patent Citations
Electroencephalogram signal sounding detection method based on convolutional sequential network
CN117688372A