Electroencephalogram signal text decoding method based on joint loss framework and post-editing optimization

By adopting a hybrid decoding framework of CTC and Attention and post-editing optimization method in EEG text decoding, the problems of difficulty in alignment and insufficient high-level semantic capture in the prior art are solved, and a more efficient and accurate text decoding effect is achieved.

CN119939343AActive Publication Date: 2025-05-06HARBIN INST OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510021200.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The existing EEG text decoding methods have inadequacy, making it difficult to effectively capture high-level semantic information of thinking content, and the difficulty of overfitting and alignment limits the generalization ability and practical application effect of the model.

Method used

The EEG signal text decoding method based on joint loss framework and post-editing optimization is adopted, and a hybrid framework of CTC and Attention decoder is combined to dynamically align through CTC losses, and post-editing optimization is used to inject semantic information.

Benefits of technology

It significantly improves the accuracy and fluency of text decoding, enhances the ability to capture high-level semantic information, reduces dependence on externally aligned information, and improves the universality and practical application effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939343A_ABST
    Figure CN119939343A_ABST
Patent Text Reader

Abstract

The invention discloses an electroencephalogram signal text decoding method based on a joint loss framework and post-editing optimization, which belongs to the field of brain-computer interface and natural language processing, and comprises the following steps: acquiring an electroencephalogram signal, constructing a classification model based on an electroencephalogram feature extraction network, and inputting the electroencephalogram signal into the classification model, obtaining a predetermined semantic category; constructing a decoding model based on a hybrid decoding framework, and inputting the electroencephalogram signal into the decoding model to obtain an output feature text; and constructing a post-editing model based on a pre-training language model, and performing post-editing on the predetermined semantic category and the output feature text based on the post-editing model to obtain a text decoding result. According to the method, the CTC / Attention hybrid framework is introduced, so that the decoding process can be automatically completed without external alignment information, the dependence on external signals such as eye movement is reduced, and the universality of the method is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of brain-computer interface and natural language processing, and in particular to an electroencephalogram signal text decoding method based on a joint loss framework and post-editing optimization. Background Art

[0002] With the development of pre-trained models, some studies have begun to try to model the EEG text decoding task as a translation task, using the encoder to extract high-level features of the EEG signal, and then connecting the pre-trained model to achieve end-to-end training. This type of method has improved the decoding performance to a certain extent with the help of the powerful generation ability of the pre-trained model. The model takes the word-level feature sequence of EEG as input, first inputs it into a multi-layer Transformer encoder, and then obtains the mapped embedded representation through a single-layer feedforward network. Next, the mapped embedded representation is input into the pre-trained model BART, and finally the decoded text is generated through the BART model.

[0003] However, there are obvious limitations in directly treating the EEG text decoding task as a translation task. EEG signals have significant time series characteristics, and the alignment relationship between their input and output sequences is closer to sequence prediction tasks such as speech recognition (ASR) rather than traditional natural language translation tasks. Therefore, there is a certain incompatibility in directly treating the EEG text decoding task as a translation task. To solve this problem, methods such as EEG-TO-TEXT rely on additional information such as eye movement signals to intercept the EEG sequence corresponding to each word and convert it into word-level features. However, in actual usage scenarios, this type of additional information is often difficult to obtain, which makes the deployment and practical application of the model more complicated and difficult. On the other hand, since EEG data is usually limited, the powerful fitting ability of large pre-trained models is prone to overfitting, which affects the generalization ability of the model and limits its effect in practical applications. In addition, existing EEG text decoding methods often have the problem of over-emphasizing low-granularity information (such as word level), resulting in the quality of generated text being less than ideal, especially in capturing high-level semantic information of thought content. This is mainly because the loss function usually only includes word-level cross entropy loss during model training, and the loss function does not explicitly consider high-level semantic information. This also leads to poor decoding results to a certain extent. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides an EEG signal text decoding method based on a joint loss framework and post-editing optimization, comprising:

[0005] Acquire an EEG signal, construct a classification model based on an EEG feature extraction network, input the EEG signal into the classification model, and obtain a predetermined semantic category;

[0006] Building a decoding model based on a hybrid decoding framework, inputting the EEG signal into the decoding model, and obtaining output feature text;

[0007] A post-editing model is constructed based on the pre-trained language model, and the predetermined semantic category and the output feature text are post-edited based on the post-editing model to obtain a text decoding result.

[0008] Preferably, the process of obtaining the predetermined semantic category includes: extracting the time domain and frequency domain features of the EEG signal based on the EEG feature extraction network, and classifying the time domain and frequency domain features of the EEG signal into predetermined semantic categories.

[0009] Preferably, the hybrid decoding framework includes an encoder, a classification module and an attention mechanism module.

[0010] Preferably, the process of obtaining the output feature text includes:

[0011] Performing feature extraction and downsampling on the EEG signal based on the encoder to obtain an extracted signal;

[0012] Performing text transcription on the extracted signal based on the attention mechanism module to obtain a target text sequence;

[0013] Based on the classification module, the extraction signal and the target text sequence are aligned to obtain the output feature text.

[0014] Preferably, the encoder includes a front-end feature extractor and an attention encoder; the attention mechanism module is a Transformer decoder; and the classification module is a connection temporal classification module.

[0015] Preferably, the process of obtaining the text decoding result includes:

[0016] The predetermined semantic category is used as a prompt, the output feature text is input into the post-editing model for calculation, the decoded text is grammatically corrected and semantically optimized, and the text decoding result is obtained.

[0017] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method is implemented when the processor executes the computer program.

[0018] On the other hand, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and wherein the method is implemented when the computer program is executed by a processor.

[0019] Compared with the prior art, the present invention has the following advantages and technical effects:

[0020] This paper applies the joint decoding framework of connectionist temporal classification (CTC) and attention-based decoder to the task of electroencephalogram (EEG) text decoding for the first time. The CTC decoding branch solves the problem of inconsistent lengths of input EEG signals and output text sequences by establishing temporal alignment between EEG signals and output text; the Attention module models the contextual dependencies of EEG signals through the attention mechanism, thereby achieving efficient and accurate text generation. The combination of CTC and Attention is groundbreaking in this field. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0022] Figure 1 A schematic diagram of a decoding model structure according to an embodiment of the present invention;

[0023] Figure 2 The figure is an overall flow chart of the method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0025] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0026] Embodiment 1

[0027] like Figure 1-2 As shown, this embodiment provides an EEG signal text decoding method based on a joint loss framework and post-editing optimization, including:

[0028] Acquire an EEG signal, construct a classification model based on an EEG feature extraction network, input the EEG signal into the classification model, and obtain a predetermined semantic category;

[0029] Building a decoding model based on a hybrid decoding framework, inputting the EEG signal into the decoding model, and obtaining output feature text;

[0030] A post-editing model is constructed based on the pre-trained language model, and the predetermined semantic category and the output feature text are post-edited based on the post-editing model to obtain a text decoding result.

[0031] The specific implementation method is:

[0032] A post-editing optimization method based on semantic category prediction and pre-trained language model is proposed, which includes a classifier module for predicting the semantic category of EEG signals, and inputting the semantic category into the pre-trained language model as a prompt, which is combined with the decoded text for grammatical and semantic optimization, thereby improving the accuracy and fluency of the decoding results.

[0033] The classification model is used to extract semantic category information from EEG signals. The model extracts the time domain and frequency domain features of EEG signals through the EEG feature extraction network EEGNet and classifies them into predetermined semantic categories. This semantic category information will be input into the post-editing model to guide the post-editing process.

[0034] The decoding model uses a hybrid decoding framework of CTC (connected temporal classification) and Attention. The CTC decoding branch dynamically aligns the input EEG signal with the output text sequence, allowing sequence generation without explicit alignment. The Attention decoding branch further strengthens context modeling and semantic capture through the attention mechanism, improving the flexibility and accuracy of decoding. The two decoding branches are trained end-to-end together to improve the feature extraction capability of the encoder, effectively handle the temporal dependency of EEG signals, and ultimately improve the accuracy and fluency of the text generated by the two decoding branches. The decoding model specifically includes:

[0035] Encoder: The encoder consists of a front-end feature extractor and an attention encoder. The front-end feature extractor is usually composed of several layers of convolutional networks, which are designed to extract features and downsample the input signal to reduce computational overhead; in this invention, EEGNet is used as a front-end feature extractor to process the original EEG signal to adapt to the decoding requirements of the EEG task. The attention encoder is composed of several stacked Conformer blocks to capture long-term and short-term dependencies and perform deep modeling of the input signal.

[0036] Attention decoding branch: This branch uses the Transformer decoder for decoding and is trained using cross-entropy loss. The Transformer attention mechanism allows the model to accurately focus on a specific part of the entire input sequence at each step, extracting only the information at the key position, thereby achieving accurate text transcription; while the cross-entropy loss function models the mapping relationship between the input and output sequences in a non-explicit alignment manner without introducing any conditional independence assumptions, ensuring effective end-to-end learning.

[0037] CTC decoding branch: This branch converts the output features of the encoder into the probability distribution of each symbol through a fully connected layer and constrains it through CTC loss. CTC loss uses frame-level annotation information and the conditional independence assumption to effectively solve the sequence alignment problem between input and output.

[0038] The feature sequence of the EEG signal is temporally aligned with the target text sequence through the CTC loss function, and the intermediate state of the output provides alignment constraints.

[0039] The encoder-decoder structure based on the attention mechanism is used to capture the contextual information of the EEG signal. The encoder extracts the signal features and combines it with the decoder to generate the target text sequence.

[0040] The CTC module and the Attention module are integrated into a hybrid decoding framework, and the CTC output is used to guide the attention alignment process of Attention. The two are jointly trained to achieve complementary advantages of the two decoding methods.

[0041] In order to further optimize the decoding results, this embodiment introduces a post-editing model, which uses the pre-trained language model BART to post-edit the output of the decoding model. The post-editing model uses the text output by the decoding model and the semantic category output by the classification model as input to perform targeted modification and polishing on the text output by the decoding model. This step effectively improves the quality of the decoding results and ensures the semantic consistency and fluency of the decoding results.

[0042] This embodiment adopts a CTC / Attention hybrid decoding framework in the decoding part, and dynamically aligns the input EEG signal and the output text through CTC loss, avoiding the limitations of the fixed alignment mode in the traditional method. At the same time, the Attention part enhances the modeling of the context through the attention mechanism, improving the flexibility and accuracy of decoding. In the post-editing part, the decoding results of the decoding model are post-edited according to the predicted semantic category, and semantic information is injected into it, further improving the accuracy of decoding. As shown in Table 1, this embodiment achieves a BLEU1 value of 25.31% and a ROUGE1 value of 25.82% on the ZUCO dataset, which is significantly higher than the existing methods.

[0043] Table 1

[0044]

[0045] By introducing a pre-trained language model as a post-editing model and combining it with the semantic category hint information provided by the classification model, the post-editing process can optimize the decoding results in a targeted manner, correct grammatical errors, eliminate ambiguity, and enhance semantic consistency. This post-editing optimization significantly improves the fluency and semantic consistency of the generated text. Compared with the original decoding model, BLEU-N (N = 2, 3, 4) is improved by 1.94%, 2.93%, and 2.90%, respectively.

[0046] On the other hand, this embodiment further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method is implemented when the processor executes the computer program.

[0047] On the other hand, this embodiment further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and wherein the method is implemented when the computer program is executed by a processor.

[0048] The technical effects of this embodiment are:

[0049] This embodiment adopts a CTC / Attention hybrid decoding framework in the decoding part, and dynamically aligns the input EEG signal and the output text through CTC loss, avoiding the limitations of the fixed alignment mode in the traditional method. At the same time, the Attention part enhances the modeling of the context through the attention mechanism, improving the flexibility and accuracy of decoding. In the post-editing part, the decoding results of the decoding model are post-edited according to the predicted semantic category, and semantic information is injected into it, further improving the accuracy of decoding. The present invention achieved a BLEU1 value of 25.08% and a ROUGE1 value of 25.35% on the ZUCO dataset, which is significantly higher than the existing methods.

[0050] By introducing a pre-trained language model as a post-editing model and combining it with the semantic category hint information provided by the classification model, the post-editing process can optimize the decoding results in a targeted manner, correct grammatical errors, eliminate ambiguity, and enhance semantic consistency. This post-editing optimization significantly improves the fluency and semantic consistency of the generated text, and BLEU4 is improved by 1.05% compared to the EEG-TO-TEXT method.

[0051] Traditional EEG text decoding methods often rely on external auxiliary information, such as eye movement signals, to align the decoding process. However, this information is often difficult to obtain in practical applications, limiting the applicability of the model. This embodiment introduces a CTC / Attention hybrid framework so that the decoding process can be automatically completed without external alignment information, reducing the dependence on external signals such as eye movement and enhancing the versatility of the method.

[0052] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for EEG signal text decoding based on a joint loss framework and post-editing optimization, characterized in that: include: Acquire an EEG signal, construct a classification model based on an EEG feature extraction network, input the EEG signal into the classification model, and obtain a predetermined semantic category; Building a decoding model based on a hybrid decoding framework, inputting the EEG signal into the decoding model, and obtaining output feature text; A post-editing model is constructed based on the pre-trained language model, and the predetermined semantic category and the output feature text are post-edited based on the post-editing model to obtain a text decoding result.

2. The method according to claim 1, characterized in that: The process of obtaining the predetermined semantic category includes: extracting the time domain and frequency domain features of the EEG signal based on the EEG feature extraction network, and classifying the time domain and frequency domain features of the EEG signal into predetermined semantic categories.

3. The method according to claim 1, characterized in that The hybrid decoding framework includes an encoder, a classification module and an attention mechanism module.

4. The method according to claim 3, characterized in that: The process of obtaining the output feature text includes: Performing feature extraction and downsampling on the EEG signal based on the encoder to obtain an extracted signal; Performing text transcription on the extracted signal based on the attention mechanism module to obtain a target text sequence; Based on the classification module, the extraction signal and the target text sequence are aligned to obtain the output feature text.

5. The method according to claim 3, characterized in that: The encoder includes a front-end feature extractor and an attention encoder; the attention mechanism module is a Transformer decoder; and the classification module is a connection timing classification module.

6. The method according to claim 1, characterized in that The process of obtaining the text decoding result includes: The predetermined semantic category is used as a prompt, the output feature text is input into the post-editing model for calculation, the decoded text is grammatically corrected and semantically optimized, and the text decoding result is obtained.

7. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1 to 6 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Language model stability optimization method based on mixed prompt learning

    CN118364050A

  • Voice brain-computer interface decoding method based on multi-modal feature fusion

    CN119002706A

  • Pre-Training With Alignments For Recurrent Neural Network Transducer Based End-To-End Speech Recognition

    US20210312905A1

  • Target speaker separation system, device and storage medium

    US20240005941A1