A method for predicting brain electrical signals based on speaker voice induction

By using a speech-evoked EEG signal prediction method, and by employing signal preprocessing, mapping relationship modeling, and generative adversarial networks, the problem of difficulty in collecting listener EEG signals in existing technologies has been solved, achieving more accurate EEG signal prediction and brain cognitive feature extraction.

CN115620751BActive Publication Date: 2026-03-20SHANXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies make it difficult to collect listeners' EEG signals in real-world situations, and linear regression models cannot effectively model the complexity and dynamics of the brain, limiting the predictive applicability of speech-evoked EEG signals.

Method used

A method for predicting EEG signals based on speaker speech is adopted. Through signal preprocessing, mapping relationship modeling and EEG signal prediction steps, the mapping relationship between speaker speech and EEG signals is established using random Gaussian matrices and generative adversarial networks. EEG signal prediction is then performed by combining generative adversarial networks and OMP signal reconstruction technology.

Benefits of technology

It achieves more realistic prediction of speech signals induced by EEG signals, is applicable to nonlinear signal analysis, solves the problem of training with small sample data, and can predict the characteristics of listener's EEG signals, providing brain cognitive information for single-modal speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620751B_ABST
    Figure CN115620751B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of speaker voice-induced electroencephalogram prediction method based on brain signal.The main problem to be solved is that the existing voice-induced method cannot collect listener brain electrical signals in actual situation.The present application includes three steps of signal preprocessing, mapping relationship modeling and electroencephalogram prediction, wherein the mapping relationship modeling includes two parts of signal coding and signal generation, signal coding is to use Gaussian random matrix to carry out perceptual observation on speaker voice signal and listener electroencephalogram, to obtain corresponding observation value, signal generation is to use generative adversarial network as basic model, the observation value of voice signal is used as input, the observation value of electroencephalogram is used as target, and generative adversarial network is trained, in the prediction process of electroencephalogram, speaker voice is used as the input of mapping relationship model, the observation value of electroencephalogram induced by the speaker voice is generated, and it is reconstructed, to obtain reconstructed electroencephalogram, and the electroencephalogram is the predicted listener electroencephalogram.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of electroencephalogram, and particularly relates to a method for predicting electroencephalogram signals based on speaker voice induction. BACKGROUND

[0002] The electroencephalogram (EEG) signal is the overall reflection of the electrical physiological activity of the internal nerve cells of the brain, and can reflect the thinking activity of the human brain. It is not affected by subjective factors, can effectively shield irrelevant interference induced by tasks, and has strong stability and anti-interference of the recognition result. The electroencephalogram signal contains all the information of human brain cognition, and the recognition result of the electroencephalogram signal is the direct result of human brain cognition. With the significant progress of AI, human-computer interaction and brain-computer interface (BCI) technology, the coding mechanism of the brain to external stimulation is also constantly explored. As the most important way of communication between people, it is particularly important for the machine to understand and understand the semantics and emotions in the voice. Therefore, if the analysis results of the electroencephalogram signal and the voice signal are fused, the result will be more consistent with the cognitive analysis result of the human brain.

[0003] However, in the current human-computer interaction application, the AI intelligent agent can easily collect the voice of the speaker, but it is difficult to obtain the electroencephalogram signal of the listener in time. In addition, in the research on voice-induced electroencephalogram signals, most researchers focus on how to decode the voice-induced signal from the electroencephalogram signal, and it is not clear about using what method to predict the voice-induced electroencephalogram signal. The current technology uses a linear regression model, but the linear regression model cannot model the complexity and dynamics of the brain, so the correlation between the actual signal and the predicted signal is very small, thus limiting its applicability. SUMMARY

[0004] The purpose of the application is to solve the technical problem that the existing voice induction method cannot collect the electroencephalogram signal of the listener in actual situations, and to provide a method for predicting electroencephalogram signals based on speaker voice induction. Through this method, not only the electroencephalogram signal of the listener induced by the speaker voice can be predicted, but also the brain cognition features of the single-modal voice can be provided, which helps to solve the emotional recognition problem of the single-modal voice.

[0005] To solve the above technical problems, the technical scheme adopted by the application is:

[0006] A method for predicting electroencephalogram signals based on speaker voice induction, the specific steps of which are:

[0007] 1) Signal preprocessing

[0008] The brain electric signals of a listener induced by a speaker's voice are collected, and the speaker's voice signal and the brain electric signals of the listener induced by the speaker's voice are preprocessed, the preprocessing of the speaker's voice signal refers to time length regularization and frame processing of the speaker's voice signal, to obtain the processed real voice signal X1, the time length regularization is to set voice data of different time lengths into voice data of fixed time length to facilitate frame processing of the signal; the preprocessing of the brain electric signals of the listener refers to filtering, de-artifact, baseline correction and data integration of the brain electric signals of the listener, to obtain the processed real brain electric signal X2, the filtering is to perform 0.5-30Hz filtering processing of the brain electric signals by Fourier transform, the de-artifact is to remove electrooculogram, electrocardiogram and bad lead signals and invalid brain electric signals, the data integration is to firstly extract the brain electric signals of the listener corresponding to the speaker's voice signal, then select effective electrode lead data as target brain electric signals from the brain electric signals of the listener, and finally perform time length regularization and frame processing on the target brain electric signals, which is the same as the time length regularization and frame processing of the speaker's voice signal;

[0009] 2) Mapping relationship modeling

[0010] The mapping relationship modeling refers to simulating the encoding process of the brain to the speaker's voice signal, establishing a mapping relationship model between the speaker's voice signal and the brain electric signal, the mapping relationship model includes two parts of signal encoding and signal generation, the signal encoding is to perform perceptual observation on the preprocessed real voice signal X1 and the real brain electric signal X2 by a random Gaussian matrix Φ, that is, to perform perceptual observation on the N×1 dimensional signal X1 and the N×1 dimensional signal X2 by the M×N dimensional random Gaussian matrix Φ, to obtain the M×1 dimensional observation value Y The perceptual observation is performed on the M×N dimensional and M<<N random Gaussian matrix Φ, to obtain the M×1 dimensional observation value Y and the dimension M of the observation value Y is much smaller than the dimension N of the original signal X, obtaining the observation value Y1 of the speech signal with a compression ratio of M / N which is much smaller than the dimension of the original signal, and the observation value Y2 of the electroencephalogram, wherein the observation value Y2 of the electroencephalogram is obtained by setting multiple random Gaussian matrices Φ to observe and sample the electroencephalogram multiple times, obtaining multiple observation values Y2 of the same electroencephalogram; that is, by setting a different random Gaussian matrix Φ, a ≥ 2, so that one electroencephalogram has a observation value, by observing and sensing b identical electroencephalograms, b ≥ 2, obtaining b × a observation values, wherein the identical electroencephalograms include the electroencephalograms induced by different listeners when hearing the same speech; the signal generation is to generate the observation value of the electroencephalogram according to the observation value of the speech signal, using a generative adversarial network as a basic model for signal generation, training the mapping relationship between the observation value Y1 of the speech signal and the observation value Y2 of the electroencephalogram, taking the observation value Y1 of the speech signal as the input of the generative adversarial network, taking the observation value Y2 of the electroencephalogram as the target data of the generative adversarial network, training the generative adversarial network until a new observation value with the same distribution as the real electroencephalogram observation value can be generated, combining the trained generative adversarial network with the signal coding as a mapping relationship model of the speech signal and the electroencephalogram, and establishing the mapping relationship between the speech signal and the electroencephalogram.

[0011] 3) Electroencephalogram prediction

[0012] The electroencephalogram prediction refers to the prediction of the listener's electroencephalogram induced by the speaker's speech signal. The process first inputs the preprocessed speaker's speech signal into the mapping relationship model constructed in step 2), and then goes through the signal coding and signal generation processes of the mapping relationship model to obtain the observation value Y2' of the generated electroencephalogram of the speech signal. Then the observation value Y2' of the generated electroencephalogram is reconstructed to obtain the reconstructed electroencephalogram, which is the predicted signal X2' of the listener's electroencephalogram induced by the speaker's speech signal.

[0013] Further, the selection of effective electrode leads refers to selecting 2 leads in each of the 6 brain regions on the surface of the brain scalp, i.e. left frontal central region: FC1, FC3; left central region: C1, C3; left central parietal region: CP1, CP3; right frontal central region: FC2, FC4; right central region: C2, C4; right central parietal region: CP2, CP4, a total of 12 leads as effective electrodes.

[0014] Further, the generative adversarial network is composed of a generative network G and a discriminative network D. In the training process, the input of the generative network G is the observation value Y1 of the speech signal, and the target data of the discriminative network D is the observation value Y2 of the real listener's electroencephalogram. By continuously rejecting the discriminative network D, the training samples of the speech signal are trained into new electroencephalograms with the same distribution as the real electroencephalogram.

[0015] Further, the generated observation value Y2' of the brain electrical signal is generated according to the constructed mapping relationship model, and the constructed generation network G is used to generate the observation value of the brain electrical signal of the listener induced by the speaker voice signal, that is, the generated observation value Y2' of the brain electrical signal.

[0016] Further, the reconstruction of the generated observation value Y2' of the brain electrical signal refers to the recovery of the brain electrical signal by adopting the OMP signal reconstruction method, and the brain electrical prediction signal X2' of the listener induced by the speaker voice signal is obtained.

[0017] The beneficial effects of the present application are:

[0018] 1. Unlike the brain electrical prediction method of the linear regression model, the mapping relationship modeling method adopted by the present application is more suitable for the analysis and processing of the brain electrical signal which is a nonlinear signal. Through the method, the predicted brain electrical signal is closer to the brain electrical signal induced by the voice signal in the real situation.

[0019] 2. The present application can obtain multiple observation values of brain electrical signals of different listeners when hearing the same induced voice by setting multiple Gaussian random matrices. Since the observation value is the target data of the generative adversarial network in the present application, the model training problem of small sample data can be solved by adopting the present application.

[0020] 3. By predicting the brain electrical signal of the listener induced by the speaker voice, the brain electrical signal characteristics of the listener can be further predicted, the cognitive results of the human brain are simulated, the brain cognitive information is provided for the single-mode voice signal, and the method proposed in the present application can not only predict the brain electrical signal induced by the voice, but also change the input data for predicting the brain electrical signal induced by vision. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is the effective electrode schematic diagram of the brain region of the present application;

[0022] Figure 2 is the experimental method and process diagram of the present application. DETAILED DESCRIPTION

[0023] The present application will be described in detail below in combination with the drawings and examples.

[0024] The specific steps of the brain electrical signal prediction method based on the speaker voice induction in the embodiment are as follows:

[0025] 1) Signal preprocessing

[0026] The 64-lead electroencephalogram device is used to collect the listener's electroencephalogram signals induced by the speaker's voice, and the speaker's voice signals and the listener's electroencephalogram signals induced by the speaker's voice signals are preprocessed, the preprocessing of the speaker's voice signals refers to the time length normalization and frame processing of the speaker's voice signals, to obtain the processed real voice signal X1, the time length normalization is to set the voice data of different time lengths into fixed time length voice data to facilitate the frame processing of the signals, the preprocessing of the listener's electroencephalogram signals refers to filtering, de-artifact, baseline correction and data integration of the listener's electroencephalogram signals, to obtain the processed real electroencephalogram signal X2, first, the electroencephalogram signal is filtered at 0.5-30Hz by Fourier transform, then the electrooculogram, electrocardiogram and bad lead signals and invalid electroencephalogram signals are removed, and the listener's electroencephalogram signals corresponding to the speaker's voice signals are extracted, then the effective electrode lead data of the listener's electroencephalogram signals are selected as the target electroencephalogram signals, and finally the target electroencephalogram signals are time length normalized and frame processed, which is the same as the time length normalization and frame processing of the speaker's voice signals.

[0027] In the integration of electroencephalogram data, since one voice can map the electroencephalogram data of multiple participants, and the electroencephalogram signal of each participant is a multi-channel electroencephalogram signal, it is necessary to integrate the multi-subject and multi-channel electroencephalogram data, in addition, since the multi-channel electroencephalogram signals have similarity, and some lead channel signals have large interference and cannot be used in the collection process, therefore, in the frontal region (left / right), central region (left / right), parietal region (left / right) 6 brain regions, 2 lead channels are selected respectively, that is: left frontal central region (LFC): FC1, FC3; left central region (LC): C1, C3; left central parietal region (LP): CP1, CP3; right frontal central region (RF): FC2, FC4; right central region (RC): C2, C4; right central parietal region (RP): CP2, CP4, as shown in Figure 1 , a total of 12 leads are used as effective electrodes.

[0028] 2) Mapping relationship modeling

[0029] The mapping relationship modeling refers to simulating the encoding process of the brain to the speaker's voice signal, establishing the mapping relationship model between the speaker's voice signal and the electroencephalogram signal, and the mapping relationship model includes signal encoding and signal generation.

[0030] The signal encoding is to perform perceptual observation on the preprocessed real voice signal X1 and the real electroencephalogram signal X2 by using a random Gaussian matrix Φ, that is, the N×1 dimensional signal X is observed by the M×N dimensional random Gaussian matrix Φ, so that Y=ΦX, and the M×1 dimensional observation value Y is obtained. The perceptual observation is performed on the M×N dimensional and M<<N random Gaussian matrix Φ, so that Y=ΦX, and the M×1 dimensional observation value Y is obtained. And the dimension M of the observation value Y is much smaller than the dimension N of the original signal X, and the observation value Y1 of the compressed speech signal and the observation value Y2 of the electroencephalogram signal with a compression ratio of M / N much smaller than the dimension of the original signal are obtained, wherein the observation value Y2 of the electroencephalogram signal is obtained by setting a plurality of random Gaussian matrices Φ to observe and sample the electroencephalogram signal multiple times to obtain a plurality of observation values Y2 of the same electroencephalogram signal, that is, by setting a(a≥2) different random Gaussian matrices Φ, so that one electroencephalogram signal has a observation value, and by observing the perception of b(b≥2) same electroencephalogram signals, b×a observation values are obtained, wherein the same electroencephalogram signals include the electroencephalogram signals induced by different listeners when listening to the same speech.

[0031] The signal generation is to generate the observation value of the electroencephalogram signal according to the observation value of the speech signal, to adopt a generative adversarial network as a basic model of signal generation, to train a mapping relationship between the observation value Y1 of the speech signal and the observation value Y2 of the electroencephalogram signal, to take the observation value Y1 of the speech signal as an input of the generative adversarial network, to take the observation value Y2 of the electroencephalogram signal as target data of the generative adversarial network, to train the generative adversarial network until a new observation value with the same distribution as the real electroencephalogram signal observation value can be generated, to set a(a≥2) different random Gaussian matrices Φ according to Y2=ΦX2, so that the number of observation values Y2 is expanded by a times, and the number of target data of the generative adversarial network is expanded, thereby solving the problem of small number of generative adversarial network samples. The generative adversarial network is composed of a generative network G and a discriminative network D, and in the training process, the input of the generative network G is the observation value Y1 of the speech signal, and the target data of the discriminative network D is the observation value Y2 of the real listener electroencephalogram signal, and through the continuous rejection of the discriminative network D, the training sample of the speech signal is trained into a new electroencephalogram signal with the same distribution as the real electroencephalogram signal.

[0032] The trained generative adversarial network combined with signal coding is a mapping relationship model of the speech signal and the electroencephalogram signal.

[0033] 3) Electroencephalogram signal prediction

[0034] The brain electrical signal prediction refers to the prediction of the listener brain electrical signal induced by the speaker voice signal, and the process is first inputting the preprocessed speaker voice signal according to the mapping relationship model constructed in step 2), and then obtaining the observation value Y2' of the generated brain electrical signal of the voice signal through the signal coding and signal generation of the mapping relationship model, and then reconstructing the observation value Y2' of the generated brain electrical signal to obtain the reconstructed brain electrical signal, and the reconstructed brain electrical signal is the listener brain electrical prediction signal X2' induced by the speaker voice signal. The specific steps are: the voice signal of the speaker is observed by using a random Gaussian matrix Ф, and the observation value Y1 of the voice signal of the speaker is obtained, and according to the constructed mapping relationship model, the generated speaker voice induced listener brain electrical signal observation value is output by giving the generation network G, and the observation value is the observation value Y2' of the generated brain electrical signal, and when the observation value Y2' of the generated brain electrical signal is reconstructed, if the sparse value of the listener brain electrical prediction signal X2' under the sparse matrix Ψ is θ, then according to Y2'=ФX2'=ФΨθ=A CS θ, since the perception matrix A CS =ФΨ is a known matrix, and the sparse matrix Ψ is a known DCT matrix, according to X2'=Ψθ, the OMP signal reconstruction method in the compressed sensing theory can be used to solve the sparse value θ, and the speaker voice induced listener brain electrical prediction signal X2' can be reconstructed, and the OMP algorithm used in the application is an orthogonal matching pursuit algorithm, which selects each element in θ in a greedy iterative manner, and each selected element is the element with the maximum correlation degree with the current A CS , and the least square method is used to obtain the coefficient approximation of the original signal, and finally the listener brain electrical prediction signal X2' under the observation value Y2' of the generated brain electrical signal is solved, that is, the speaker voice induced listener brain electrical signal is predicted.

Claims

1. A method for predicting EEG signals based on speaker speech evoked by an individual, characterized in that, The specific steps are as follows: 1) Signal preprocessing The speaker's speech was collected and the listener's EEG signal was evoked. The speaker's speech signal and the listener's EEG signal evoked by the speaker's speech signal were preprocessed. The preprocessing of the speaker's speech signal refers to the duration normalization and framing of the speaker's speech signal to obtain the processed real speech signal X1. The duration normalization is to set the speech data of different durations into speech data of fixed duration to facilitate the framing of the signal. The preprocessing of the listener's EEG signal refers to filtering, artifact removal, baseline correction, and data integration of the listener's EEG signal to obtain the processed real EEG signal X2. The filtering is performed by using Fourier transform to filter the EEG signal from 0.5 to 30 Hz. The artifact removal is to remove electrooculogram (EOG), electrocardiogram (ECG), bad lead signals, and invalid EEG signals. The data integration first extracts the listener's EEG signal corresponding to the speaker's speech signal, then selects valid electrode lead data from the listener's EEG signal as the target EEG signal, and finally performs duration normalization and framing processing on the target EEG signal. This operation is the same as the duration normalization and framing processing of the speaker's speech signal. 2) Mapping relationship modeling The mapping relationship modeling refers to simulating the encoding process of the speaker's speech signal by the brain and establishing a mapping relationship model between the speaker's speech signal and the electroencephalogram (EEG) signal. The mapping relationship model includes two parts: signal encoding and signal generation. The signal encoding is to perform perceptual observation on the preprocessed real speech signal X1 and real EEG signal X2 using a random Gaussian matrix Φ, that is, to perform perceptual observation on the N×1-dimensional signal on a random Gaussian matrix Φ of M×N dimensions with M << N, so that Y = ΦX, and obtain an M×1-dimensional observation value and the dimension M of the observation value Y is much smaller than the dimension N of the original signal X, obtaining the observation value Y1 of the speech signal and the observation value Y2 of the EEG signal with a compression ratio of M / N that is much smaller than the dimension of the original signal. Among them, the observation value Y2 of the EEG signal is obtained by setting multiple random Gaussian matrices Φ to perform multiple observation samplings on the EEG signal, obtaining multiple observation values Y2 of the same EEG signal; that is, by setting a different random Gaussian matrices Φ, a ≥ 2, so that an EEG signal has a observation values, and by performing perceptual observation on b identical EEG signals, b ≥ 2, obtaining b×a observation values, where these identical EEG signals include the EEG signals induced by different listeners when hearing the same speech; The signal generation is to generate the observation value of the EEG signal according to the observation value of the speech signal. Using the generative adversarial network as the basic model of signal generation, training the mapping relationship between the observation value Y1 of the speech signal and the observation value Y2 of the EEG signal, taking the observation value Y1 of the speech signal as the input of the generative adversarial network, taking the observation value Y2 of the EEG signal as the target data of the generative adversarial network, training the generative adversarial network until new observation values with the same distribution as the real EEG signal observation values can be generated, combining the trained generative adversarial network with the signal encoding as the mapping relationship model of the speech signal and the EEG signal, and establishing the mapping relationship between the speech signal and the EEG signal; 3) EEG signal prediction The EEG signal prediction refers to the prediction of the EEG signal induced by the speaker's speech signal in the listener. The process is as follows: First, based on the mapping relationship model constructed in step 2), the preprocessed speaker's speech signal is input. After two processes, signal encoding and signal generation, the observed value Y2' of the generated EEG signal of the speech signal is obtained. Then, the observed value Y2' of the generated EEG signal is reconstructed to obtain the reconstructed EEG signal. This reconstructed EEG signal is the predicted EEG signal X2' of the listener induced by the speaker's speech signal in the listener.

2. The method for predicting EEG signals based on speaker speech evoked according to claim 1, characterized in that: The selection of effective electrode leads refers to selecting two lead channels from each of the six brain regions on the surface of the scalp: left frontal central region: FC1, FC3; left central region: C1, C3; left central parietal region: CP1, CP3; right frontal central region: FC2, FC4; right central region: C2, C4; right central parietal region: CP2, CP4, for a total of 12 leads as effective electrodes.

3. The method for predicting EEG signals based on speaker speech evoked according to claim 1, characterized in that: The generative adversarial network consists of a generator network G and a discriminator network D. During training, the input of the generator network G is the observed value Y1 of the speech signal, and the target data of the discriminator network D is the observed value Y2 of the real listener's EEG signal. Through the continuous rejection of the discriminator network D, the training samples of the speech signal are trained into new EEG signals with the same distribution as the real EEG signals.

4. The method for predicting EEG signals based on speaker speech evoked according to claim 1, characterized in that: The observed value Y2' of the generated EEG signal is generated by using the constructed generative network G to generate the observed value of the EEG signal induced by the speaker's speech signal, based on the constructed mapping relationship model.

5. The method for predicting EEG signals based on speaker speech evoked according to claim 1, characterized in that: The reconstruction of the observed value Y2' of the generated EEG signal refers to the use of the OMP signal reconstruction method to recover the EEG signal and obtain the listener's EEG prediction signal X2' induced by the speaker's speech signal.

Citation Information

Patent Citations

  • Speech conversion method based on Gaussian process output post-filtering

    CN106782599A

  • Emotion electroencephalogram signal inducing method based on dialogues

    CN113208635A