Method and device for electroencephalogram-voice-text three-mode alignment
Through deep learning algorithms and feature matching technology, the features of three-modal data of EEG, speech and text are extracted and aligned, and the problems of low alignment efficiency and poor accuracy in the existing technology are solved, and efficient and accurate three-modal alignment is achieved.
Patent Information
- Application Number
- CN202510164781.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has problems of low efficiency, poor accuracy and high computational complexity when processing three-modal alignment of EEG signals, text information and voice information, making it difficult to achieve real-time alignment.
Deep learning algorithms and feature matching technology are used to extract the global and local features of the three-modal data respectively, and the global alignment features are optimized through the similarity function to enhance the consistency of the three-modal modes.
It realizes efficient and accurate three-modal data alignment, improves the accuracy and consistency of data alignment, and provides a solid foundation for subsequent data analysis and processing.
Smart Images

Figure CN120011826A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and more specifically to a method and device for electroencephalogram-speech-text tri-modal alignment. Background Art
[0002] As artificial intelligence and neuroscience are increasingly integrated, multimodal data analysis has become a hot topic in the research field. As three key modalities, electroencephalogram (EEG), text information and voice information each contain rich physiological, semantic and emotional information, which are of great significance to emotion recognition, cognitive science research, human-computer interaction and other fields. However, the asynchrony and misalignment of these modalities in time, space and semantic levels have brought great challenges to data fusion.
[0003] As a direct reflection of brain activity, EEG signals have a high physiological basis and time resolution, but their susceptibility to noise interference makes data processing difficult. Text information provides a direct semantic description and is an important window for understanding human thinking and emotions. However, its acquisition often relies on subjective input, which is subjective and delayed. Voice information contains rich emotional information such as intonation, speech speed, and timbre, and is an important means of emotional expression, but it is also easily affected by environmental noise and has a relatively low time resolution.
[0004] In order to achieve effective fusion of trimodal data, accurate alignment is the key. The alignment method needs to be able to handle the asynchrony and misalignment of different modal data at the temporal, spatial and semantic levels. Temporal alignment requires that different modal data remain consistent in time, which is crucial for real-time emotion recognition and cognitive science research. Spatial alignment is the basis for ensuring that different modal data correspond correctly in spatial position, which is of great significance for tasks such as image and speech processing. Semantic alignment is to match data of different modalities at the semantic level, which usually requires a deep understanding of the meaning and context of the data, which is particularly critical for tasks such as natural language processing and emotion recognition.
[0005] However, existing technologies still have limitations when processing the trimodal alignment of EEG signals, text information, and speech information. Some methods may only be applicable to specific types of multimodal data, and are not effective for EEG data with high temporal resolution and susceptible to noise. Other methods may have high computational complexity and are difficult to achieve real-time alignment in practical applications.
[0006] Therefore, providing an efficient and accurate method and device for EEG-speech-text tri-modal alignment is an issue that needs to be urgently addressed by those skilled in the art. Summary of the invention
[0007] In view of this, the present invention provides a method and device for EEG-speech-text trimodal alignment, which accurately extracts and aligns the global and local features of the trimodal data through deep learning algorithms and feature matching technology, optimizes the global alignment features using a similarity function, and enhances the EEG-speech-text trimodal consistency.
[0008] In order to achieve the above object, the present invention adopts the following technical solution:
[0009] A method for EEG-speech-text trimodal alignment, comprising:
[0010] Acquire the three-modal data of the tester respectively; the three-modal data includes electroencephalogram signal, speech signal and text information;
[0011] Preprocess the collected three-modal data respectively;
[0012] Based on the preprocessed EEG signals, speech signals and text information, deep learning algorithms are used to extract global features and local features.
[0013] Using a feature matching algorithm to align the global features and the local features respectively, to obtain a global alignment feature and a local alignment feature;
[0014] The similarity between the global alignment feature and the local alignment feature is calculated using a similarity function, and the global alignment feature is optimized based on the similarity to obtain final trimodal alignment data.
[0015] Preferably, the collected three-modal data are preprocessed respectively, including:
[0016] The preprocessing of the EEG signal includes: filtering the collected EEG signal to obtain a filtered EEG signal; performing noise reduction processing on the filtered EEG signal to remove the interference of the biological electrical signal to obtain a noise-reduced EEG signal;
[0017] The preprocessing of speech signals includes:
[0018] Acquire a first spectrum of the collected speech signal;
[0019] Using a preset noise spectrum as a first noise spectrum of the first spectrum, and filtering the first spectrum based on the first noise spectrum to obtain a second spectrum;
[0020] Determine whether there is voice information in the second spectrum; if there is no voice information in the second spectrum, discard the voice signal; if there is a voice signal in the second spectrum, separate the second spectrum to obtain a valid voice signal;
[0021] The preprocessing of text includes:
[0022] Performing denoising processing on the text information to obtain denoised text, wherein the denoising processing includes format standardization and removal of special symbols and punctuation marks;
[0023] Perform multi-dimensional vectorization on the denoised text according to preset dimensions to obtain vectorized text;
[0024] Acquire text information in the vectorized text that complies with a preset state transfer rule;
[0025] The text information is calculated using a dynamic programming algorithm, and optimal text information that complies with a preset format is determined, and the optimal text information is output as a preprocessing result of the text information.
[0026] Preferably, the global features and local features of the trimodal data are extracted using a deep learning algorithm based on the preprocessed EEG signal, speech signal and text information, respectively, including:
[0027] Use Conv1-Conv4 in ResNet50 to extract local features of preprocessed EEG signals, speech signals, and text information respectively;
[0028] Locally encoding the local features respectively to obtain local encoding signals;
[0029] The local coding signal is input into a trained multi-layer perceptron feature extraction model for feature extraction to obtain global features; the multi-layer perceptron feature extraction model includes at least one multi-layer perceptron module, and the multi-layer perceptron module includes: a global channel correlation feature extraction layer and a global temporal feature extraction layer, and the global channel correlation feature of the EEG local coding signal is extracted through the global channel correlation feature extraction layer; the global temporal feature of the global channel correlation feature is extracted through the global temporal feature extraction layer to obtain global features.
[0030] Preferably, the global features and the local features are aligned respectively using a feature matching algorithm to obtain a global alignment feature and a local alignment feature, including:
[0031] The global features and local features are respectively input into a hybrid attention mechanism that supports online learning. The global features and local features of the EEG signal, speech signal, and text information are linearly transformed to generate corresponding key, value, and query pairs. The dot product attention mechanism is used to extract the correlation information between multimodal signals. The residual operator is used to fuse single modal features with multimodal correlation information to obtain global alignment features and local alignment features.
[0032] Preferably, the similarity between the global alignment feature and the local alignment feature is calculated using a similarity function, and the global alignment feature is optimized based on the similarity to obtain the final trimodal alignment data, including:
[0033] The similarity of the global alignment feature with respect to each local alignment feature is calculated, and the similarity calculation formula is:
[0034] S i =α i V
[0035] Among them, α i is the attention coefficient of the global alignment feature with respect to the i-th local alignment feature, and V is the global alignment feature representation;
[0036] Determine, based on the similarity, an area where the similarity between the global alignment feature and the local feature is lower than a threshold;
[0037] For the areas where the similarity between the global alignment features and the local features is lower than the threshold, the global alignment features and the local alignment features are weightedly fused to form an optimized global alignment feature.
[0038] On the other hand, the present invention provides a device for EEG-speech-text trimodal alignment, comprising:
[0039] A data acquisition device, used to obtain three-modal data of the tester respectively; the three-modal data includes electroencephalogram signals, voice signals and text information;
[0040] A preprocessing module, used to preprocess the collected three-modal data respectively;
[0041] A feature extraction module, for extracting global features and local features of the trimodal data using a deep learning algorithm based on the preprocessed EEG signal, speech signal and text information;
[0042] An alignment module, used to align the global features and the local features respectively by using a feature matching algorithm to obtain a global alignment feature and a local alignment feature;
[0043] The optimization module is used to calculate the similarity between the global alignment feature and the local alignment feature by using a similarity function, and optimize the global alignment feature based on the similarity to obtain final trimodal alignment data.
[0044] Preferably, the preprocessing module comprises:
[0045] The EEG signal preprocessing unit is used to filter the collected EEG signals to obtain filtered EEG signals; and to perform noise reduction processing on the filtered EEG signals to remove the interference of biological electrical signals to obtain noise-reduced EEG signals;
[0046] A speech signal preprocessing unit is used to obtain a first spectrum of a collected speech signal; use a preset noise spectrum as a first noise spectrum of the first spectrum, and based on the first noise spectrum, filter the first spectrum to obtain a second spectrum; determine whether there is speech information in the second spectrum; if there is no speech information in the second spectrum, discard the speech signal; if there is a speech signal in the second spectrum, separate the second spectrum to obtain a valid speech signal;
[0047] The text preprocessing unit is used to perform denoising on the text information to obtain denoised text, wherein the denoising includes format standardization and removal of special symbols and punctuation marks; multi-dimensional vectorization is performed on the denoised text according to preset dimensions to obtain vectorized text; text information in the vectorized text that complies with preset state transfer rules is obtained; the text information is calculated using a dynamic programming algorithm, and the optimal text information that complies with the preset format is determined, and the optimal text information is output as the preprocessing result of the text information.
[0048] Preferably, the feature extraction module includes:
[0049] The local feature extraction unit uses Conv1-Conv4 in ResNet50 to extract the local features of the preprocessed EEG signal, speech signal and text information respectively;
[0050] A local encoding unit, used for locally encoding the local features respectively to obtain local encoding signals;
[0051] The global feature extraction unit is used to input the local coding signal into a trained multi-layer perceptron feature extraction model for feature extraction to obtain global features; the multi-layer perceptron feature extraction model includes at least one multi-layer perceptron module, and the multi-layer perceptron module includes: a global channel correlation feature extraction layer and a global temporal feature extraction layer, and the global channel correlation feature of the EEG local coding signal is extracted by the global channel correlation feature extraction layer; the global temporal feature of the global channel correlation feature is extracted by the global temporal feature extraction layer to obtain global features.
[0052] Preferably, the feature alignment module is configured as follows:
[0053] The global features and local features are respectively input into a hybrid attention mechanism that supports online learning. The global features and local features of the EEG signal, speech signal, and text information are linearly transformed to generate corresponding key, value, and query pairs. The dot product attention mechanism is used to extract the correlation information between multimodal signals. The residual operator is used to fuse single modal features with multimodal correlation information to obtain global alignment features and local alignment features.
[0054] Preferably, the optimization module is configured as follows:
[0055] The similarity of the global alignment feature with respect to each local alignment feature is calculated, and the similarity calculation formula is:
[0056] S i =α i V
[0057] Among them, α i is the attention coefficient of the global alignment feature with respect to the i-th local alignment feature, and V is the global alignment feature representation;
[0058] Determine, based on the similarity, an area where the similarity between the global alignment feature and the local feature is lower than a threshold;
[0059] For the areas where the similarity between the global alignment features and the local features is lower than the threshold, the global alignment features and the local alignment features are weightedly fused to form an optimized global alignment feature.
[0060] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a method and device for EEG-speech-text trimodal alignment. First, the present invention can accurately extract and align the global and local features of trimodal data through deep learning algorithms and feature matching technology, effectively improving the accuracy and consistency of data alignment. This precise alignment provides a solid foundation for subsequent data analysis and processing, and provides reliable data support for research in related fields. Secondly, the present invention provides a complete and efficient preprocessing process, including filtering, noise reduction, denoising, vectorization and other steps, to ensure the accuracy and reliability of input data. This helps to reduce data noise and interference, improve data quality, and provide a strong guarantee for subsequent alignment and analysis work. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0062] Figure 1 The present invention provides a flowchart for an EEG-speech-text tri-modal alignment method.
[0063] Figure 2 A framework diagram for an EEG-speech-text tri-modal alignment system is provided for the present invention. DETAILED DESCRIPTION
[0064] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] The embodiment of the present invention discloses a method for alignment of EEG-speech-text trimodality, such as Figure 1 As shown, including:
[0066] The three-modal data of the test subject are obtained respectively; the three-modal data include EEG signals, speech signals and text information.
[0067] First, using sophisticated EEG acquisition equipment, the weak electrical signals generated by the tester's brain activity, namely EEG signals, can be recorded in real time. These signals can reflect the tester's thinking state, emotional changes and other neural activities. Secondly, through the high-sensitivity microphone equipment, the tester's voice signals are captured. These signals contain the tester's verbal expression, intonation, speaking speed and other language characteristics. Finally, the tester's text information is also collected, including written text, input electronic documents or information entered through the keyboard. These text information can directly reflect the tester's writing characteristics, ideological content and expression methods. The tester's emotional state can be judged by writing characteristics such as the neatness of the handwriting, the weight of the strokes, and the size and shape of the font. For example, when the author feels nervous or anxious, the handwriting may become sloppy or unstable; when the author feels calm or happy, the handwriting may be more neat and fluent. These writing characteristics can serve as an auxiliary basis for judging emotions. By comprehensively using the data of these three modalities, we can more comprehensively understand the tester's cognitive process, language behavior, and the relationship between thinking and language, thereby providing rich data support for in-depth research. More importantly, when the writing features of the written text need to be analyzed, image information needs to be added, that is, the alignment of the four modal data is similar to the three-modal alignment method provided in the embodiment of the present invention.
[0068] Preprocess the collected three-modal data respectively;
[0069] Based on the preprocessed EEG signals, speech signals and text information, deep learning algorithms are used to extract global features and local features.
[0070] Using feature matching algorithm to align global features and local features respectively, to obtain global alignment features and local alignment features;
[0071] The similarity function is used to calculate the similarity between the global alignment features and the local alignment features, and the global alignment features are optimized based on the similarity to obtain the final trimodal alignment data.
[0072] Specifically, the collected three-modal data are preprocessed respectively, including:
[0073] The preprocessing of the EEG signal includes: filtering the collected EEG signal to obtain a filtered EEG signal; and performing noise reduction processing on the filtered EEG signal to remove the interference of the biological electrical signal to obtain a noise-reduced EEG signal.
[0074] The raw EEG signals collected from the tester's scalp using high-quality EEG acquisition equipment (such as an EEG cap) usually contain multiple frequency components, including useful EEG activity signals and interference signals from the outside and inside. Therefore, it is necessary to perform bandpass filtering on the collected raw EEG signals to remove signal components that are not within the frequency range of interest. Usually, the frequency range of EEG signals is between 0.5Hz and 100Hz, so a suitable bandpass filter can be selected, such as 0.5Hz to 70Hz, to retain the signal within this range. This step helps to reduce the influence of high-frequency noise (such as electromyographic interference) and low-frequency noise (such as baseline drift). At the same time, in order to remove interference of specific frequencies (such as power frequency interference, usually 50Hz or 60Hz), a notch filter is used to set narrowband suppression at these frequency points. This helps to reduce the impact of power line interference on EEG signals.
[0075] Furthermore, it is necessary to remove the interference of biological electrical signals, including:
[0076] Electrooculogram (EOG) signal removal (EOG correction): The electrooculogram (EOG) signal generated by eye movement is one of the most common interferences in EEG signals. By simultaneously recording the electrooculogram signal and using regression algorithms (such as independent component analysis (ICA) or principal component analysis (PCA)) to separate and remove the electrooculogram component from the EEG signal, a purer EEG signal can be obtained.
[0077] ECG signal removal (ECG correction): Although the ECG signal has relatively little interference with the EEG signal, it may also need to be removed in some cases. Similarly, regression algorithms or template-based removal methods can be used to reduce the impact of the ECG signal.
[0078] EMG signal removal: EMG signals usually appear as high-frequency noise, which can be reduced by bandpass filtering or adaptive filtering techniques. In addition, for obvious EMG artifacts, manual marking and interpolation methods can be used to remove them.
[0079] Other noise reduction processing: baseline correction, artifact removal.
[0080] The preprocessing of speech signals includes:
[0081] Acquire a first spectrum of the collected speech signal; perform Fourier transform (such as short-time Fourier transform STFT) on the collected original speech signal, convert it from the time domain to the frequency domain, and obtain a first spectrum, so as to facilitate observation and analysis of the distribution characteristics of the speech signal at different frequencies.
[0082] The preset noise spectrum is used as the first noise spectrum of the first spectrum. Based on the first noise spectrum, the first spectrum is filtered to obtain the second spectrum. The preset noise spectrum is a noise signal (such as environmental noise) without speech collected in advance under ideal conditions, and the spectrum is analyzed to obtain the second spectrum, which is used as a reference in subsequent filtering. The first spectrum is compared with the preset noise spectrum, and the first spectrum is filtered using algorithms such as Kalman filtering to remove or weaken the noise component.
[0083] Determine whether there is voice information in the second spectrum; if there is no voice information in the second spectrum, discard the voice signal; if there is a voice signal in the second spectrum, separate the second spectrum to obtain a valid voice signal.
[0084] The preprocessing of text includes:
[0085] De-noising the text information to obtain de-noised text, where the de-noising process includes format standardization and removal of special symbols and punctuation marks;
[0086] Perform multi-dimensional vectorization on the denoised text according to the preset dimensions to obtain vectorized text; perform multi-dimensional vectorization on the denoised text according to the preset dimensions. Vectorization is the process of converting text into numerical form so that the computer can understand and process it. The preset dimensions may include word frequency, bag of words model, TF-IDF (term frequency-inverse document frequency), etc. These dimensions can reflect different characteristics of the text. Through multi-dimensional vectorization, a vectorized text containing multiple numerical features can be obtained.
[0087] Obtain text information in the vectorized text that complies with the preset state transition rules; Obtain text information in the vectorized text that complies with the preset state transition rules. The state transition rules may be based on the grammatical structure, semantic relationship, or specific pattern in the text. For example, rules can be set to extract noun phrases, verb phrases, or sentences in a specific format from the text. This step helps to filter out useful information from the text and facilitates subsequent processing.
[0088] The text information is calculated using a dynamic programming algorithm, and the optimal text information that conforms to a preset format is determined, and the optimal text information is output as a preprocessing result of the text information. The acquired text information is calculated using a dynamic programming algorithm, and the optimal text information that conforms to a preset format is determined. Dynamic programming is an algorithm for solving optimization problems, which can find the global optimal solution while ensuring efficiency. In this step, the optimal text information combination can be calculated by a dynamic programming algorithm according to preset format requirements (such as length restrictions, keyword inclusion, etc.). Finally, the optimal text information is output as a preprocessing result of the text information.
[0089] Furthermore, based on the preprocessed EEG signals, speech signals and text information, deep learning algorithms are used to extract global and local features of the trimodal data, including:
[0090] Use Conv1-Conv4 in ResNet50 to extract local features of preprocessed EEG signals, speech signals, and text information respectively;
[0091] Local features are locally encoded to obtain local encoding signals; specifically, the following are included:
[0092] Multiple EEG data are input into the trained local encoding module, each channel is locally encoded, and the local spatial feature encoding data of each channel is obtained.
[0093] The EEG local coding signal is obtained according to the local spatial feature coding data of each channel.
[0094] In one embodiment, the local coding module is a one-dimensional convolutional coding module, including three one-dimensional convolutional layers and three batch normalization layers. In this embodiment, since the distribution of EEG data of multiple channels is different, the output local spatial feature coding data of each layer is batch-converted into a normal distribution through the batch normalization layer, so that the subsequent network layer does not need to learn the data distribution of the previous layer, which reduces the dependence of the subsequent network on the previous network, improves the feedback ability of the local coding module, reduces the problem of gradient disappearance during the training process of the local coding module, and makes the convergence speed of the local coding module faster.
[0095] The local coding signal is input into a trained multi-layer perceptron feature extraction model for feature extraction to obtain global features; the multi-layer perceptron feature extraction model includes at least one multi-layer perceptron module, and the multi-layer perceptron module includes: a global channel correlation feature extraction layer and a global temporal feature extraction layer, and the global channel correlation feature of the EEG local coding signal is extracted through the global channel correlation feature extraction layer; the global temporal feature of the global channel correlation feature is extracted through the global temporal feature extraction layer to obtain the global feature.
[0096] In a specific implementation, the multi-layer perceptron (MLP) feature extraction model architecture includes at least one multi-layer perceptron module. The MLP consists of an input layer, a hidden layer, and an output layer, and a fully connected design is adopted between each layer. The multi-layer perceptron module in this embodiment is specially designed to include two key layers: a global channel correlation feature extraction layer and a global temporal feature extraction layer. The former focuses on learning fine spatial information, that is, the correlation between channels, and it consists of the first batch of normalization layers and channel attention layers; the latter aims to capture temporal information and consists of a second batch of normalization layers and activation layers.
[0097] Specifically, the input signal of the multilayer perceptron module is denoted as X. Signal X is first processed by the first batch of normalization layers to generate an output signal X1. Subsequently, X1 is sent to the channel attention layer, which outputs a signal X2 containing the weight values of each channel (i.e., attention information). Next, the original input signal X is superimposed with the output signal X2 of the channel attention layer to obtain a new output signal X3. X3 then enters the second batch of normalization layers to generate an output signal X4. X4 is then processed by the activation layer to obtain an output signal X5. Finally, X5 is superimposed with the previous X3 to obtain the output signal Y of the multilayer perceptron module.
[0098] In addition, in one implementation case, the multilayer perceptron feature extraction model also integrates a data preprocessing module, a local encoding module, N sequentially connected multilayer perceptron modules, and a fully connected classification layer. The model takes the original EEG signal data as input and outputs the classification result. Among them, N multilayer perceptron modules work in series to extract global channel correlation features and global timing features from the EEG local encoding signal, and N is a positive integer not less than 1. When N is greater than 1, the first multilayer perceptron module receives the EEG local encoding signal as input, and the last module outputs the global feature signal.
[0099] In another embodiment, the global features and the local features are aligned respectively using a feature matching algorithm to obtain global alignment features and local alignment features, including:
[0100] The global features and local features are respectively input into the hybrid attention mechanism that supports online learning. The global features and local features of EEG signals, speech signals, and text information are linearly transformed to generate corresponding key, value, and query pairs. The dot product attention mechanism is used to extract the correlation information between multimodal signals, and the residual operator is used to fuse the single modality features with the multimodal correlation information to obtain global alignment features and local alignment features. Specifically, the global features and local features are firstly sent into the hybrid attention mechanism respectively. Inside the mechanism, these features will undergo a series of linear transformation operations to generate the corresponding key (Key), value (Value), and query (Query) pairs. This step lays a solid foundation for the subsequent attention calculation and can more accurately capture the correlation and importance between different features. The efficient dot product attention mechanism is used to further explore the deep correlation information between multimodal signals. The dot product attention mechanism calculates the dot product between the query and the key and normalizes it through the softmax function to obtain the attention weight corresponding to each key. These weights are then used to weighted sum the values to generate feature representations that integrate multimodal correlation information. In order to fully integrate the single modality features with the multimodal correlation information, the residual operator is introduced. The residual operator generates the final global alignment features and local alignment features by adding the feature representation of the single modality with the feature representation that integrates the multimodal correlation information, and adding a learnable bias term. This step not only retains the unique information of the single modality features, but also cleverly integrates the correlation information between the multimodal features, thereby significantly improving the alignment effect of the features and the processing performance of subsequent tasks.
[0101] In another embodiment, the similarity between the global alignment feature and the local alignment feature is calculated using a similarity function, and the global alignment feature is optimized based on the similarity to obtain the final trimodal alignment data, including:
[0102] Calculate the similarity of the global alignment feature with respect to each local alignment feature. The similarity calculation formula is:
[0103] S i =α i V
[0104] Among them, α i is the attention coefficient of the global alignment feature with respect to the i-th local alignment feature, and V is the global alignment feature representation;
[0105] Based on the similarity, determine the area where the similarity between the global alignment feature and the local feature is lower than a threshold;
[0106] For the areas where the similarity between the global alignment features and the local features is lower than the threshold, the global alignment features and the local alignment features are weightedly fused to form an optimized global alignment feature.
[0107] On the other hand, the present invention provides a device for EEG-speech-text trimodal alignment, such as Figure 2 As shown, including:
[0108] A data acquisition device is used to obtain the three-modal data of the test subject respectively; the three-modal data includes EEG signals, voice signals and text information;
[0109] A preprocessing module, used to preprocess the collected three-modal data respectively;
[0110] A feature extraction module is used to extract global features and local features of the trimodal data using a deep learning algorithm based on the preprocessed EEG signal, speech signal and text information;
[0111] An alignment module is used to align the global features and the local features respectively by using a feature matching algorithm to obtain a global alignment feature and a local alignment feature;
[0112] The optimization module is used to calculate the similarity between the global alignment feature and the local alignment feature using a similarity function, optimize the global alignment feature based on the similarity, and obtain the final trimodal alignment data.
[0113] Preferably, the preprocessing module comprises:
[0114] The EEG signal preprocessing unit is used to filter the collected EEG signals to obtain filtered EEG signals; and to perform noise reduction processing on the filtered EEG signals to remove the interference of biological electrical signals to obtain noise-reduced EEG signals;
[0115] A speech signal preprocessing unit is used to obtain a first spectrum of the collected speech signal; use a preset noise spectrum as a first noise spectrum of the first spectrum, and based on the first noise spectrum, filter the first spectrum to obtain a second spectrum; determine whether there is speech information in the second spectrum; if there is no speech information in the second spectrum, discard the speech signal; if there is a speech signal in the second spectrum, separate the second spectrum to obtain a valid speech signal;
[0116] The text preprocessing unit is used to perform denoising on the text information to obtain denoised text, wherein the denoising process includes format standardization and removal of special symbols and punctuation marks; multi-dimensional vectorization is performed on the denoised text according to preset dimensions to obtain vectorized text; text information in the vectorized text that complies with preset state transfer rules is obtained; the text information is calculated using a dynamic programming algorithm, and the optimal text information that complies with the preset format is determined, and the optimal text information is output as the preprocessing result of the text information.
[0117] Preferably, the feature extraction module includes:
[0118] The local feature extraction unit uses Conv1-Conv4 in ResNet50 to extract the local features of the preprocessed EEG signal, speech signal and text information respectively;
[0119] A local coding unit, used for locally coding local features to obtain local coding signals;
[0120] The global feature extraction unit is used to input the local coding signal into the trained multi-layer perceptron feature extraction model for feature extraction to obtain global features; the multi-layer perceptron feature extraction model includes at least one multi-layer perceptron module, and the multi-layer perceptron module includes: a global channel correlation feature extraction layer and a global temporal feature extraction layer, and the global channel correlation feature of the EEG local coding signal is extracted through the global channel correlation feature extraction layer; the global temporal feature of the global channel correlation feature is extracted through the global temporal feature extraction layer to obtain the global feature.
[0121] Preferably, the feature alignment module is configured as follows:
[0122] The global features and local features are respectively input into the hybrid attention mechanism that supports online learning. The global features and local features of EEG signals, speech signals, and text information are linearly transformed to generate corresponding key, value, and query pairs. The dot product attention mechanism is used to extract the correlation information between multimodal signals, and the residual operator is used to fuse single modal features with multimodal correlation information to obtain global alignment features and local alignment features.
[0123] Preferably, the optimization module is configured as follows:
[0124] Calculate the similarity of the global alignment feature with respect to each local alignment feature. The similarity calculation formula is:
[0125] S i =α i V
[0126] Among them, α i is the attention coefficient of the global alignment feature with respect to the i-th local alignment feature, and V is the global alignment feature representation;
[0127] Based on the similarity, determine the area where the similarity between the global alignment feature and the local feature is lower than a threshold;
[0128] For the areas where the similarity between the global alignment feature and the local feature is lower than the threshold, the global alignment feature and the local alignment feature are weightedly fused to form an optimized global alignment feature.
[0129] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0130] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for EEG-speech-text trimodal alignment, characterized in that: include: Acquire the three-modal data of the tester respectively; the three-modal data includes electroencephalogram signal, speech signal and text information; Preprocess the collected three-modal data respectively; Based on the preprocessed EEG signals, speech signals and text information, deep learning algorithms are used to extract global features and local features. Using a feature matching algorithm to align the global features and the local features respectively, to obtain a global alignment feature and a local alignment feature; The similarity between the global alignment feature and the local alignment feature is calculated using a similarity function, and the global alignment feature is optimized based on the similarity to obtain final trimodal alignment data.
2. The method for trimodal alignment of EEG, speech and text according to claim 1, characterized in that: The collected three-modal data are preprocessed separately, including: The preprocessing of the EEG signal includes: filtering the collected EEG signal to obtain a filtered EEG signal; performing noise reduction processing on the filtered EEG signal to remove the interference of the biological electrical signal to obtain a noise-reduced EEG signal; The preprocessing of speech signals includes: Acquire a first spectrum of the collected speech signal; Using a preset noise spectrum as a first noise spectrum of the first spectrum, and filtering the first spectrum based on the first noise spectrum to obtain a second spectrum; Determine whether there is voice information in the second spectrum; if there is no voice information in the second spectrum, discard the voice signal; if there is a voice signal in the second spectrum, separate the second spectrum to obtain a valid voice signal; The preprocessing of text includes: Performing denoising processing on the text information to obtain denoised text, wherein the denoising processing includes format standardization and removal of special symbols and punctuation marks; Perform multi-dimensional vectorization on the denoised text according to preset dimensions to obtain vectorized text; Acquire text information in the vectorized text that complies with a preset state transfer rule; The text information is calculated using a dynamic programming algorithm, and optimal text information that complies with a preset format is determined, and the optimal text information is output as a preprocessing result of the text information.
3. The method for trimodal alignment of EEG, speech and text according to claim 1, characterized in that: Based on the preprocessed EEG signal, speech signal and text information, the global features and local features of the three modal data are extracted using deep learning algorithms, including: Use Conv1-Conv4 in ResNet50 to extract local features of preprocessed EEG signals, speech signals, and text information respectively; Locally encoding the local features respectively to obtain local encoding signals; The local coding signal is input into a trained multi-layer perceptron feature extraction model for feature extraction to obtain global features; the multi-layer perceptron feature extraction model includes at least one multi-layer perceptron module, and the multi-layer perceptron module includes: a global channel correlation feature extraction layer and a global temporal feature extraction layer, and the global channel correlation feature of the EEG local coding signal is extracted through the global channel correlation feature extraction layer; the global temporal feature of the global channel correlation feature is extracted through the global temporal feature extraction layer to obtain global features.
4. The method for trimodal alignment of EEG, speech and text according to claim 1, characterized in that: The global features and the local features are respectively aligned using a feature matching algorithm to obtain global alignment features and local alignment features, including: The global features and local features are respectively input into a hybrid attention mechanism that supports online learning. The global features and local features of the EEG signal, speech signal, and text information are linearly transformed to generate corresponding key, value, and query pairs. The dot product attention mechanism is used to extract the correlation information between multimodal signals. The residual operator is used to fuse single modal features with multimodal correlation information to obtain global alignment features and local alignment features.
5. The method for trimodal alignment of EEG, speech and text according to claim 1, characterized in that: Calculating the similarity between the global alignment feature and the local alignment feature using a similarity function, optimizing the global alignment feature based on the similarity, and obtaining final trimodal alignment data, including: The similarity of the global alignment feature with respect to each local alignment feature is calculated, and the similarity calculation formula is: S i =a i V Among them, α i is the attention coefficient of the global alignment feature with respect to the i-th local alignment feature, and V is the global alignment feature representation; Determine, based on the similarity, an area where the similarity between the global alignment feature and the local feature is lower than a threshold; For the areas where the similarity between the global alignment features and the local features is lower than the threshold, the global alignment features and the local alignment features are weightedly fused to form an optimized global alignment feature.
6. A device for EEG-speech-text trimodal alignment, characterized in that: include: A data acquisition device, used to respectively acquire the three-modal data of the tester; The trimodal data includes EEG signals, speech signals and text information; A preprocessing module, used to preprocess the collected three-modal data respectively; A feature extraction module, for extracting global features and local features of the trimodal data using a deep learning algorithm based on the preprocessed EEG signal, speech signal and text information; An alignment module, used to align the global features and the local features respectively by using a feature matching algorithm to obtain a global alignment feature and a local alignment feature; The optimization module is used to calculate the similarity between the global alignment feature and the local alignment feature by using a similarity function, and optimize the global alignment feature based on the similarity to obtain final trimodal alignment data.
7. The device for EEG-speech-text trimodal alignment according to claim 6, characterized in that: The preprocessing module comprises: The EEG signal preprocessing unit is used to filter the collected EEG signals to obtain filtered EEG signals; and to perform noise reduction processing on the filtered EEG signals to remove the interference of biological electrical signals to obtain noise-reduced EEG signals; A speech signal preprocessing unit is used to obtain a first spectrum of a collected speech signal; use a preset noise spectrum as a first noise spectrum of the first spectrum, and based on the first noise spectrum, filter the first spectrum to obtain a second spectrum; determine whether there is speech information in the second spectrum; if there is no speech information in the second spectrum, discard the speech signal; if there is a speech signal in the second spectrum, separate the second spectrum to obtain a valid speech signal; The text preprocessing unit is used to perform denoising on the text information to obtain denoised text, wherein the denoising includes format standardization and removal of special symbols and punctuation marks; multi-dimensional vectorization is performed on the denoised text according to preset dimensions to obtain vectorized text; text information in the vectorized text that complies with preset state transfer rules is obtained; the text information is calculated using a dynamic programming algorithm, and the optimal text information that complies with the preset format is determined, and the optimal text information is output as the preprocessing result of the text information.
8. The device for EEG-speech-text trimodal alignment according to claim 6, characterized in that: The feature extraction module comprises: The local feature extraction unit uses Conv1-Conv4 in ResNet50 to extract the local features of the preprocessed EEG signal, speech signal and text information respectively; A local encoding unit, used for locally encoding the local features respectively to obtain local encoding signals; The global feature extraction unit is used to input the local coding signal into a trained multi-layer perceptron feature extraction model for feature extraction to obtain global features; the multi-layer perceptron feature extraction model includes at least one multi-layer perceptron module, and the multi-layer perceptron module includes: a global channel correlation feature extraction layer and a global temporal feature extraction layer, and the global channel correlation feature of the EEG local coding signal is extracted by the global channel correlation feature extraction layer; the global temporal feature of the global channel correlation feature is extracted by the global temporal feature extraction layer to obtain global features.
9. The device for EEG-speech-text trimodal alignment according to claim 6, characterized in that: The feature alignment module is configured to: The global features and local features are respectively input into a hybrid attention mechanism that supports online learning. The global features and local features of the EEG signal, speech signal, and text information are linearly transformed to generate corresponding key, value, and query pairs. The dot product attention mechanism is used to extract the correlation information between multimodal signals. The residual operator is used to fuse single modal features with multimodal correlation information to obtain global alignment features and local alignment features.
10. The device for EEG-speech-text trimodal alignment according to claim 6, characterized in that: The optimization module is configured to: The similarity of the global alignment feature with respect to each local alignment feature is calculated, and the similarity calculation formula is: S i =a i V Among them, α i is the attention coefficient of the global alignment feature with respect to the i-th local alignment feature, and V is the global alignment feature representation; Determine, based on the similarity, an area where the similarity between the global alignment feature and the local feature is lower than a threshold; For the areas where the similarity between the global alignment features and the local features is lower than the threshold, the global alignment features and the local alignment features are weightedly fused to form an optimized global alignment feature.