A weak electrical signal analysis method and system based on multi-level multi-modal fusion
By constructing a multi-level, multi-modal fusion method for weak electrical signal analysis on a three-layer spatial structure of scalp-periauria-inner ear, the hardware and methodological deficiencies in ear region EEG signal analysis have been solved. This method enables high-precision decoding of speech-related neural activities and the revelation of dynamic causal relationships, thereby improving the effectiveness of speech-brain-controlled communication and rehabilitation training.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2026-03-24
- Publication Date
- 2026-07-03
AI Technical Summary
Existing methods for analyzing EEG signals in the ear region are insufficient in terms of hardware integration, wearing stability, dynamic acquisition capabilities, and multimodal synchronous acquisition. Furthermore, they lack systematic modeling of multi-level structural features of the scalp, periauricular region, and inner ear, making it difficult to reveal the nonlinear transmission and dynamic causal relationships of speech-related neural activities. The experimental paradigm and data analysis model have low coupling, making it difficult to support complex speech tasks and high-dimensional data analysis.
We employ a multi-level, multi-modal fusion approach to analyze weak electrical signals. By constructing a unified data analysis model across three spatial layers—scalp, periauricular region, and inner ear—and combining time-frequency feature extraction, kernel principal component analysis, and deep network fusion, we build a multi-level joint model of scalp, periauricular region, and inner ear. This model enables nonlinear dynamic modeling and causal relationship analysis, and we design an ear region device and a multi-condition speech imagination experiment paradigm to achieve synchronous acquisition and analysis of multi-region, multi-modal signals.
It achieves highly robust decoding of weak electrical signals, improves the intelligibility and prediction accuracy of speech-image EEG decoding, provides quantitative analysis tools, and supports speech-brain-controlled communication and rehabilitation training.
Smart Images

Figure CN122332776A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of weak signal processing and human-computer interaction, and particularly relates to a method and system for weak electrical signal analysis based on multi-level multimodal fusion. Background Technology
[0002] The acquisition and analysis of weak electrical signals (such as bioelectrical signals and weak sensor signals) are of great value in fields such as biomedical engineering, human-computer interaction, and precision measurement. However, weak electrical signals are typically characterized by extremely low signal-to-noise ratios, susceptibility to environmental interference, and often accompanied by complex spatial attenuation and transmission paths, making accurate decoding extremely challenging. Electroencephalography (EEG), as a typical example of weak biological electrical signals, plays a central role in the field of brain-computer interfaces (BCI). BCI is a technology that establishes a direct information pathway between the brain and external devices by acquiring and analyzing neurophysiological signals such as EEG signals. Speech-related BCIs, focusing on processes such as "speech perception, speech imagination, and speech generation," hold promise for providing new auxiliary communication methods for patients with severe speech disorders and can also be used in scenarios such as human-computer interaction and smart home control. Existing non-invasive BCI mainly relies on scalp EEG acquisition, which requires shaving the hair and applying conductive gel to the multi-channel electrodes. This is not only cumbersome to operate, but also easily affected by artifacts such as hair, sweat and head movements. It is also uncomfortable to wear for a long time, which is not conducive to daily use and wearable applications.
[0003] To improve wearing comfort and ease of use, researchers are increasingly focusing on EEG acquisition in hairless areas such as the frontal or auricular regions. The auricular region (including the inner and periauricular areas), with its fixed location, thin skin, and proximity to auditory-related brain regions, is considered an important candidate for long-term monitoring and daily wear. Existing research on the auricular region has explored various hardware structures, including in-ear, behind-the-ear, and surround-the-ear designs. Some systems integrate a few acquisition channels or simple audio playback modules, which can be used for sleep monitoring, basic auditory tasks, or simplified brain-controlled command output. However, these systems still have shortcomings in functional integration, wearing stability, long-term dynamic acquisition capabilities, mode switching, and multimodal simultaneous acquisition, making it difficult to provide a stable data foundation for complex speech tasks and high-dimensional data analysis.
[0004] Beyond hardware limitations, the problems with data analysis methods are even more pronounced. Existing speech-related EEG research largely follows two paths: one focuses on speech acoustic signal processing, examining the correlation between acoustic features such as spectrograms and Mel-frequency cepstral coefficients and scalp EEG for speech perception or reconstruction; the other focuses on motor imagery or speech intent classification, often using word / phrase labels to train classifiers or deep neural networks. The former type of work has revealed the time-frequency coupling relationship of speech processing to some extent, but it usually only uses linear correlation or simplified mapping models with limited scalp channels, making it difficult to characterize the nonlinear transmission and dynamic causal relationships of speech-related neural activity between different brain regions; the latter type of work relies heavily on discrete labels and large-scale labeled data, resulting in poor model interpretability, limited generalization ability, and rarely considering the multi-level structural features of the ear region as a special observation window.
[0005] Existing methods for ear region data analysis still have the following shortcomings: First, most studies only treat intra-ear or periauricular signals as a "single alternative channel," simply performing correlation analysis or feature splicing with scalp EEG, lacking systematic modeling that treats the scalp-periauricular-intra-ear region as different levels of nodes in the same speech processing network; Second, commonly used feature extraction and classification methods are still mainly based on linear features, shallow machine learning models, or single-structure convolutional neural networks, making it difficult to fully utilize the nonlinear and non-stationary characteristics of EEG signals and the complementary information of multiple modalities (such as bone conduction and acceleration); Third, traditional analyses are mostly based on static correlation or average response, lacking dynamic modeling of the temporal ordered information flow between the frontal lobe, temporal lobe, and auditory cortex of the ear region, which is not conducive to revealing the causal path and network mechanism of the speech imagery process.
[0006] In terms of experimental design, existing speech-related EEG experimental paradigms mostly focus on single or a few conditions, such as examining only auditory speech reception or simple speech imagery tasks, with few systematic comparisons of typical cognitive conditions such as "auditory reception," "speech imagery," "silent reading," and "auditory reading." This lack of multi-condition, hierarchical task designs makes it difficult for the obtained data to support a unified modeling framework: on the one hand, the neural activity characteristics of different speech processing stages are difficult to organically compare, limiting the separation and analysis of speech imagery-specific components; on the other hand, the low coupling between the experimental paradigm and subsequent data analysis models reduces the efficiency and reliability of model training and validation.
[0007] Currently, in the field of ear-region speech-related brain-computer interfaces, while wearable ear-region devices have been initially explored at the hardware level, a systematic data analysis method is still lacking at the methodological level. This method is designed for weak electrical signals (especially ear-region speech-imagination EEG), capable of simultaneously characterizing multi-level spatial structures of the scalp-periauricular-intraauricular region, fusing multimodal signals, and analyzing dynamic causal relationships. Furthermore, a multi-condition speech-imagination experimental paradigm and corresponding data acquisition system closely matched to this method are also lacking. To address these issues, it is necessary to propose a new multi-level joint modeling and multimodal coupling data analysis technology system to perform structured, interpretable, and high-precision decoding of weak electrical signals (especially ear-region speech-imagination EEG). Based on this, a corresponding system and experimental paradigm should be constructed, laying a methodological foundation for speech-based brain-controlled communication and rehabilitation training. Summary of the Invention
[0008] To address the problems of low signal-to-noise ratio, lack of dynamic modeling of transmission paths, and limitations of single-modal information in the acquisition of weak electrical signals (especially weak biological signals such as ear region EEG), this invention proposes a method and system for weak electrical signal analysis based on multi-level multimodal fusion. Taking ear region speech-image EEG as a typical weak electrical signal as a specific breakthrough, this invention constructs a unified data analysis model across a three-layer spatial structure of scalp-periauricular-inner ear, and introduces methods such as time-frequency feature extraction, principal component analysis, deep network fusion, and nonlinear dynamic modeling. This invention can achieve highly robust decoding of weak electrical signals, providing new analytical tools and technical pathways for brain-controlled communication and rehabilitation training based on weak signals.
[0009] To achieve the above objectives, the technical solution adopted by this invention is as follows: a method for analyzing weak electrical signals based on multi-level multimodal fusion, wherein the weak electrical signals are specifically multi-regional electroencephalogram (EEG) signals, and the method includes the following steps: S1. Under task conditions, weak electroencephalogram (EEG) signals from the scalp, periauricular, and intraauricular regions are collected simultaneously, along with bone conduction vibration and acceleration signals, to obtain multi-regional, multimodal signals. S2. Preprocess the acquired multi-region multimodal signals, obtain the time-frequency spectrum representation of each region by performing time-frequency joint analysis, and perform frequency band division and normalization processing; S3. Construct a multi-level joint model of scalp-periauria-inner ear; S4. Based on the normalized time-frequency representation, using the scalp-periauria-intraauria multi-level joint model, nonlinear features are extracted to obtain scalp feature vectors, periauria feature vectors, and intraauria feature vectors. S5. Input the scalp feature vector, periauricular feature vector, intraauricular feature vector, bone conduction vibration signal and acceleration signal into the deep neural network to obtain the fused latent space representation; S6. Construct a state-space model and / or a dynamic causal model to model the system state evolution and information flow between the frontal lobe, temporal lobe and auditory cortex of the ear during speech stimulation and speech imagination, and estimate the causal connection parameters between regions; wherein, the hierarchical transfer topology of the scalp-periauria-intraauria multi-level joint model provides a priori structural constraints for the state-space model and / or dynamic causal model. S7. Based on the fused latent space representation and estimated dynamic causal connection parameters, train the speech imagery decoder to output speech category labels, semantic information and / or reconstructed speech features for speech imagery-related EEG signals, and complete the analysis of the data.
[0010] Furthermore, the joint time-frequency analysis is performed using at least one of the following methods: Short-time Fourier transform of EEG signals using overlapping time windows yields a time-spectrum representation;
[0011] in, Indicates the time of the original EEG signal t and frequency f The time-spectral representation of the short-time Fourier transform at a given point. Represents raw brain electrical signals. Indicated by t The time window function centered on j Represents the imaginary unit. f Represents frequency components, t Indicates the center position of the time window. T Indicates the total duration. d Represents an integral infinitesimal operator; Continuous or discrete wavelet transforms are used to construct time-spectrum representations, enhancing the ability to represent transient speech-related neural oscillations.
[0012] in, The time spectrum is represented. Indicates the scale parameter. Indicates the translation parameter. Indicates complex conjugation. This represents the original brain electrical signals.
[0013] Furthermore, step S3 includes the following steps: A hierarchical transfer function is constructed from the scalp region to the periauricular region and from the periauricular region to the inner ear region. The periauricular region is explicitly introduced as an intermediate transfer link. The signals from the scalp, periauricular region and inner ear are modeled as multi-level outputs driven by the same speech-related potential neural activity. In the multi-level joint model of scalp-periauricular region-inner ear, the transfer noise at each level is modeled and suppressed separately to obtain a multi-level joint model describing the complete transfer path of scalp-periauricular region-inner ear. The construction of the multi-level joint model of scalp-periauricular region-inner ear is completed.
[0014] Furthermore, the expression for the scalp-periauricular-intraauricular multi-level joint model is as follows:
[0015] in, Indicates the time of the inner ear region EEG signals, Represents the combined error term. This indicates the electroencephalogram (EEG) signals in the periauricular region. This represents the electroencephalogram (EEG) signals in the scalp region. This represents a composite function.
[0016] Furthermore, step S4 includes the following steps: The time-spectral representations of each region after normalization are mapped to construct a high-dimensional feature space; Based on the high-dimensional feature space, the principal component features of each region are extracted by solving the eigenvalues and eigenvectors of the kernel matrix; The extracted principal component features of each region are denoised to obtain scalp feature vectors, periauricular feature vectors, and intraauricular feature vectors.
[0017] Furthermore, the deep neural network includes: The first feature extraction subnetwork, consisting of several convolutional and pooling layers connected in series, is used to extract local patterns based on scalp feature vectors, periauricular feature vectors, intraauricular feature vectors, bone conduction vibration signals, and acceleration signals, and outputs the extracted local pattern feature sequences to the second feature modeling subnetwork.
[0018] The second feature modeling subnetwork is used to process time series dependencies based on local pattern feature sequences and weight the contributions of different time steps to obtain time-weighted features. The time-weighted features are then used as EEG features and input into the feature fusion subnetwork. The feature fusion subnetwork, including a multi-head attention module and a fully connected layer, maps scalp feature vectors, periauricular feature vectors, intraauricular feature vectors, bone conduction vibration signals, and acceleration signals to a unified space. Based on the mapping results, temporally weighted features, and the multi-head attention mechanism, it performs nonlinear fusion of multi-region and multi-modal features to obtain the fused latent space representation.
[0019] in, This represents the latent space representation after merging. This indicates multi-head attention computation. Representing EEG characteristics, i.e., time-weighted features. This represents multimodal splicing features, including bone conduction features. and acceleration characteristics , Indicates adding These are residual connections used to preserve the original EEG features.
[0020] The decoding output layer is used to output speech category labels, semantic information, and / or reconstructed speech features based on the fused latent space representation.
[0021] Furthermore, the state equations and observation equations of the state-space model are as follows:
[0022]
[0023] in, express The hidden state vector at time t. express The hidden state vector at time t. Represents the state transition matrix. Represents the input matrix, Indicates external input. Indicates process noise. Represents the observation vector. Represents the observation matrix. Indicates a direct transmission matrix. Indicates observation noise; The continuous-time form of the dynamic causal model is as follows:
[0024] in, This represents the derivative of the neural activity state vector of each brain region with respect to time. Let m represent the intrinsic connectivity matrix, and m represent the number of experimental conditions. Indicates the first j Switching variables for each experimental condition. Indicates the first j Each experimental condition corresponds to a modulation matrix connected to the connection. This represents a vector representing the neural activity state of each brain region. Represents the input matrix, Indicates external input; The expressions for the causal connection parameters between the regions are as follows:
[0025] in, Indicates time t Time zone j To the area i Causal connection parameters, Indicates the area j To the area i The inherent connection strength, k An index indicating experimental conditions. Indicates time t Time k Switching variables for each experimental condition. Indicates the first k Experimental conditions for the region j To the area i Modulation parameters for the connection strength between them.
[0026] This invention also provides a weak electrical signal analysis system based on multi-level multi-modal fusion, comprising: Ear area device, used to be fixed in the periauricular region of the subject's ear and to carry various sensing and actuation modules; An intra-ear electrode assembly, placed inside the ear canal, is used to collect electroencephalogram (EEG) signals from the intra-ear region and / or to play audio. Periauricular electrode assembly, distributed along the circumference of the auricle, is used to collect electroencephalogram (EEG) signals from the periauricular region; The bone conduction and acceleration module, fixed at the mastoid process, is used to output bone conduction speech stimulation and collect bone conduction vibration and acceleration signals; The system consists of two microcontroller modules, one on the left and one on the right, serving as the master and the other on the right. The master module is used for audio playback control, data management, and wireless transmission, while the slave module is used for multi-channel physiological signal acquisition and preprocessing. The master and slave modules are connected via a high-speed data cable to achieve timing synchronization. The clock synchronization module is used to synchronize the time of electroencephalogram (EEG) signals, bone conduction vibration signals, and acceleration signals in the scalp region, periauricular region, and intraauricular region. The communication module is used to transmit synchronized multi-region multimodal signals to external computing devices; An external computing device is configured to execute the aforementioned weak electrical signal analysis method based on multi-level multimodal fusion and output the speech imagery decoding results to complete the analysis of the data.
[0027] Furthermore, the main body of the ear area device is a support structure that bypasses the ear, with a connecting rod at the upper end for connecting the inner ear components and an insulating ear clip at the lower end for holding the earlobe. A bone conduction and acceleration module is fixedly installed at a preset position below the support, and left and right dual microcontroller modules are detachably installed at a preset position above the support. The in-ear electrode assembly is a hollow cylindrical structure with electrodes arranged in four rings from the inside to the outside along the outer surface of the earplug, and is made of conductive material; The left and right dual microcontroller modules control the switching between data acquisition mode and audio playback mode through a mode switching module. During switching, the opening and closing of the in-ear electrode acquisition channel is automatically controlled and the timestamp of the audio stimulus is recorded to support the time stamping of multi-condition tasks.
[0028] Furthermore, the experimental paradigm for speech imagination based on the weak electrical signal analysis method using multi-level multimodal fusion includes: A1. Semantic Paradigm Design: Select several independently presented keywords and construct sentences with complete semantics through fixed arrangements to balance phonetic features, semantic complexity and cognitive load, and divide them into randomly arranged keyword sequences and sentence sequences that constitute complete sentence meanings according to experimental needs; A2. Standardized speech perception acquisition: Subjects wore ear area devices set to audio playback mode and listened to speech stimuli consisting of randomly arranged keywords and complete sentences. Rest periods of no less than a preset duration were set between each group of stimuli. The speech stimuli and corresponding multi-region multimodal physiological signals were recorded and labeled through a clock synchronization module. A3. Random sequence speech acquisition: Subjects wore an in-ear component set to acquisition mode, and performed at least two different cognitive depth processing on randomly arranged keywords, including speech imagination, silent reading and / or spoken reading, while simultaneously acquiring EEG signals, bone conduction vibration signals and acceleration signals from the scalp area, periauricular area and in-ear area. A4. Complete Sentence Speech Acquisition: Using the same equipment configuration as the random sequence speech generation stage, the stimuli are upgraded to sentences with complete semantics. The speech imagination, silent reading and / or spoken reading tasks are repeatedly performed to obtain multi-region multimodal signals containing contextual information. A5. Data Quality and Labeling: Pre-experiment training is conducted before the formal experiment. In the formal experiment, real-time data quality monitoring is used to identify and remove severe artifacts. Data from each time period is bound to the corresponding task conditions, semantic labels, and behavioral responses to form a multi-region multimodal EEG dataset adapted to the weak electrical signal analysis method based on multi-level multimodal fusion.
[0029] The beneficial effects of this invention are as follows: This invention uses speech-based EEG, a typical weak electrical signal, as the processing object, and the ear region as the core observation channel. It treats signals from the inner ear, periauricular region, and scalp region as different levels of nodes on the same speech processing network. On the one hand, by constructing a multi-level spatial transmission model of scalp-periauricular region-inner ear, and introducing an intermediate periauricular link, it quantitatively describes the propagation, attenuation, and transformation of speech-related neural activities in space. On the other hand, through time-frequency joint analysis, principal component analysis, and deep feature fusion methods, it compresses and represents multi-regional, multi-modal signals and extracts nonlinear features, effectively improving the intelligibility and prediction accuracy of speech-based EEG decoding. This invention has at least the following effects: (1) A weak signal analysis method combining scalp-periauria-inner ear levels is proposed. This invention introduces a three-layer spatial structure of "scalp-periauricular-inner ear" for the first time in the decoding of weak EEG signals. By constructing a hierarchical transmission model and a multi-level noise processing mechanism, it explicitly models the transmission relationship of speech-related neural activities between different regions. This invention overcomes the prediction bias caused by using only the scalp and inner ear as the two ends and ignoring the intermediate links, and achieves a more complete and refined structural modeling of weak electrical signals.
[0030] (2) Construct a spatiotemporal multimodal fusion data analysis framework This invention combines short-time Fourier transform / wavelet transform and other time-frequency analysis, kernel principal component analysis, and deep neural networks to propose a spatiotemporal feature extraction and multimodal coupling analysis method for weak electrical signals. This invention can fuse weak EEG signals from the scalp, periauricular region, and inner ear, as well as bone conduction and acceleration signals within a unified framework, significantly improving the robustness and intelligibility of intent decoding of weak electrical signals in complex environments.
[0031] (3) Construction of nonlinear dynamic systems and causal path reconstruction methods To address the nonlinear and non-stationary characteristics of EEG signals, this invention introduces a state-space model and dynamic causal modeling to quantitatively describe the information flow between the frontal lobe, temporal lobe, and auditory cortex of the ear during speech imagery, reconstructing the dynamic causal paths of the speech processing network. This invention provides a quantitative and interpretable analytical tool for understanding the neural mechanisms of speech imagery.
[0032] (4) System and experimental paradigms that are tightly coupled with design and data analysis methods Based on the aforementioned data analysis methods, this invention designs an ear-region brain-computer interface system that supports simultaneous acquisition and playback mode switching of multimodal data from the inner ear, periauricular region, and bone conduction. It also proposes a speech imagination experiment paradigm with a progressive task structure of "listening-thinking-silent reading-audio reading," which ensures that the acquired data is highly matched with the analysis methods in terms of time stamping, spatial distribution, and cognitive conditions, thereby guaranteeing the efficiency and reliability of model training and validation from the source. Attached Figure Description
[0033] Figure 1 This is a flowchart of the method of the present invention.
[0034] Figure 2 This is a schematic diagram of a multi-level joint modeling method for scalp-periauria-inner ear provided in an embodiment of the present invention.
[0035] Figure 3 This is a schematic diagram of a multimodal deep fusion network structure provided in an embodiment of the present invention.
[0036] Figure 4 This is a schematic diagram of the composition of an ear-area speech imagination brain-computer interface system provided in an embodiment of the present invention.
[0037] Figure 5 A flowchart illustrating the speech imagination experiment paradigm provided in an embodiment of the present invention. Detailed Implementation
[0038] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0039] Example 1 This embodiment uses the ear region's speech EEG, a typical weak electrical signal, as an example to explain in detail the weak electrical signal analysis method based on multi-level multimodal fusion of the present invention. The data analysis method of the present invention uses the ear region as the core observation channel, treating signals from the inner ear, periauricular region, and scalp region as different levels of nodes on the same speech processing network: On the one hand, by constructing a multi-level spatial transmission model of scalp-periauricular region-inner ear, introducing an intermediate periauricular link, it quantitatively describes the propagation, attenuation, and transformation of speech-related neural activities in space; on the other hand, through time-frequency joint analysis, nuclear principal component analysis, and deep feature fusion methods, it compresses and represents multi-regional, multimodal signals and extracts nonlinear features, effectively improving the intelligibility and prediction accuracy of speech imagery EEG decoding.
[0040] like Figure 1 As shown, this invention provides a method for analyzing weak electrical signals based on multi-level multi-modal fusion, the implementation of which is as follows: S1. Under task conditions, simultaneously collect EEG signals from the scalp region, periauricular region, and inner ear region, and simultaneously collect bone conduction vibration signals and acceleration signals to obtain multi-regional multimodal signals; In this embodiment, multi-regional multimodal EEG data acquisition is performed as follows: under tasks such as speech perception, speech imagination, silent reading, and spoken reading, EEG signals from the scalp, periauricular, and intraauricular regions are simultaneously acquired using scalp electrodes, periauricular electrodes, and intraauricular electrodes. Bone conduction vibration signals and acceleration signals are also acquired simultaneously. Data from all channels are time-aligned using a high-precision clock synchronization module. Specifically, the multi-source information data includes: EEG signals and auxiliary modal signals; the EEG signals include: scalp EEG signals, periauricular EEG signals, and intraauricular EEG signals; the auxiliary modal signals include: bone conduction vibration signals and acceleration signals.
[0041] In this embodiment, the designed task conditions include sequentially executing two or more of the following three states on the same semantic material: a) Auditory reception state of passively listening to randomly arranged or complete sentences of speech stimuli; b) Speech imagination state of internally reciting speech content without vocalization; c) Silent reading state of participating articulatory organs but without vocalization; d) Vocal reading state of accompanied by actual vocal output; and a multi-condition contrastive dataset was constructed by repeatedly measuring the above multiple states on the same subject to train and validate the scalp-periauricular-intraauricular multi-level joint modeling and speech imagination decoder.
[0042] S2. Preprocess the acquired multi-region multimodal signals, obtain the time-frequency spectrum representation of each region by performing time-frequency joint analysis, and perform frequency band division and normalization processing; In this embodiment, signal preprocessing and time-frequency transformation are performed as follows: preprocessing operations such as detrending, bandpass filtering, artifact detection and removal, and rereference are performed on the original multi-region multimodal signals; time-frequency joint analysis is performed on the preprocessed signals, and time-frequency representations of the signals in each region and multimodal mode are obtained by using short-time Fourier transform and / or wavelet transform.
[0043] Preprocessing methods include, but are not limited to, any one or more of the following: detrending, bandpass filtering, artifact detection and removal, and rereference. Time-frequency transformation methods include, but are not limited to, any one or more of the following: short-time Fourier transform and wavelet transform.
[0044] In this embodiment, time-frequency joint analysis is performed using at least one of the following methods: Short-time Fourier transform of EEG signals using overlapping time windows yields a time-spectrum representation;
[0045] in, Indicates the time of the original EEG signal t and frequency f The time-spectral representation of the short-time Fourier transform at a given point. Represents raw brain electrical signals. Indicated by t The time window function centered on j Represents the imaginary unit. f Represents frequency components, t Indicates the center position of the time window. T Indicates the total duration. d This represents the integral differential operator.
[0046] Continuous or discrete wavelet transforms are used to construct time-spectrum representations, enhancing the ability to represent transient speech-related neural oscillations.
[0047] in, Represents multi-scale time-frequency representation (wavelet coefficients). Indicates the scale parameter. Indicates the translation parameter. Indicates complex conjugation. This represents the original brain electrical signals.
[0048] The obtained time spectrum is then divided into frequency bands and normalized to facilitate feature alignment and comparison among different subjects.
[0049] S3. Construct a multi-level joint model of scalp-periauria-inner ear, the implementation method of which is as follows: A hierarchical transfer function is constructed from the scalp region to the periauricular region and from the periauricular region to the inner ear region. The periauricular region is explicitly introduced as an intermediate transfer link. The signals from the scalp, periauricular region and inner ear are modeled as multi-level outputs driven by the same speech-related potential neural activity. In the multi-level joint model of scalp-periauricular region-inner ear, the transfer noise at each level is modeled and suppressed separately to obtain a multi-level joint model describing the complete transfer path of scalp-periauricular region-inner ear. The construction of the multi-level joint model of scalp-periauricular region-inner ear is completed.
[0050] In this embodiment, a hierarchical transfer function is constructed from the scalp to the periauricular region to the inner ear, and a multi-level joint model of scalp-periauricular region-inner ear is established to model and suppress the transmitted noise at each level separately.
[0051] like Figure 2 As shown, the multi-level joint model models EEG signals from scalp regions (especially the frontal and temporal lobes), periauricular regions, and intraauricular regions as multi-level outputs driven by the same underlying speech-related neural activity. Definition , , These represent the scalp area, periauricular area, and inner ear area respectively, representing the time... t EEG signals, The mathematical expression representing the potential neural activity related to speech is:
[0052]
[0053] in, This represents the signal transfer function from the scalp to the periauricular region. This represents the signal transfer function from the periphery of the ear to the inside of the ear. , and This indicates the intensity of the influence of potential speech-related neural activity on each region. and External noise introduced for signal transmission.
[0054] Because neural activity is complex and difficult to quantify, it is omitted and combined to obtain a complete multi-level joint model:
[0055] in, To represent a composite function, Represents the combined error term. This indicates the electroencephalogram (EEG) signals in the periauricular region. It represents the electroencephalogram (EEG) signals in the scalp region. Its key feature is the introduction of a multi-level noise processing mechanism, enabling the model to process noise at each transmission stage, thus more accurately describing the signal transmission path.
[0056] S4. Based on the normalized time-spectrum representation, using a multi-level joint model of scalp-periauricular-intraauricular region, nonlinear features are extracted to obtain scalp feature vectors, periauricular feature vectors, and intraauricular feature vectors. The implementation method is as follows: The time-spectral representations of each region after normalization are mapped to construct a high-dimensional feature space; Based on the high-dimensional feature space, the principal component features of each region are extracted by solving the eigenvalues and eigenvectors of the kernel matrix; The extracted principal component features of each region are denoised to obtain scalp feature vectors, periauricular feature vectors, and intraauricular feature vectors.
[0057] In this embodiment, nonlinear feature extraction specifically includes: mapping the time-spectrum representation of each region to a high-dimensional feature space constructed by kernel principal component analysis (Kernel PCA) or an equivalent kernel method, extracting the principal component features of each region using nonlinear kernel functions such as radial basis function kernels, removing redundant information and noise, and forming scalp feature vectors, periauricular feature vectors, and intraauricular feature vectors.
[0058] In this embodiment, the kernel function used in kernel principal component analysis includes, but is not limited to, any of the following: radial basis function kernel, polynomial kernel.
[0059] In this embodiment, kernel principal component analysis uses radial basis function kernels or polynomial kernels to map the time spectrum to a high-dimensional feature space, and then obtains the principal components by solving the eigenvalues and eigenvectors of the kernel matrix. The radial basis function (RBF) kernel is defined as follows:
[0060] in, and Represents any two time-spectrum sample vectors. This represents the kernel width parameter, which controls the rate at which sample similarity decays.
[0061] The polynomial kernel is defined as:
[0062] in, Indicates the scaling factor. Represents a constant term. d This indicates the order of the polynomial.
[0063] kernel matrix elements After centering the kernel matrix, the eigenvalue problem is solved. Obtain eigenvalues and corresponding feature vectors For new samples , No. k The projection values of each principal component are:
[0064] Preferably, the principal components with a cumulative energy contribution rate of not less than 90% are retained as the nonlinear eigenvectors of each region, i.e., those satisfying the following conditions are selected. The former Principal components are used to reduce the computational complexity of subsequent deep networks while ensuring sufficient information.
[0065] S5. Input the scalp feature vector, periauricular feature vector, intraauricular feature vector, bone conduction vibration signal, and acceleration signal into the deep neural network to obtain the fused latent space representation; such as Figure 3 As shown, a deep neural network includes: The first feature extraction subnetwork, consisting of several convolutional and pooling layers connected in series, is used to extract local patterns based on scalp feature vectors, periauricular feature vectors, intraauricular feature vectors, bone conduction vibration signals, and acceleration signals, and outputs the extracted local pattern feature sequences to the second feature modeling subnetwork.
[0066] The second feature modeling subnetwork is used to process time series dependencies based on local pattern feature sequences and weight the contributions of different time steps to obtain time-weighted features. The time-weighted features are then used as EEG features and input into the feature fusion subnetwork. The feature fusion subnetwork, including a multi-head attention module and a fully connected layer, is used to map scalp feature vectors, periauricular feature vectors, intraauricular feature vectors, bone conduction vibration signals and acceleration signals to a unified space. Based on the mapping results, temporal weighted features and the multi-head attention mechanism, nonlinear fusion of multi-region and multi-modal features is performed to obtain the fused latent space representation.
[0067] The decoding output layer is used to output speech category labels, semantic information, and / or reconstructed speech features.
[0068] In this embodiment, multimodal deep fusion modeling specifically includes: inputting scalp feature vectors, periauricular feature vectors, and intraauricular feature vectors, along with auxiliary modal features such as bone conduction and acceleration, into a deep neural network. Specifically, a convolutional neural network is used to extract time-frequency local structural features, while a recurrent neural network and / or attention mechanism are used to model long-range dependencies across time steps and cross-regional correlations. Nonlinear fusion of multi-regional and multimodal features is achieved through feature projection and a multi-head attention mechanism to obtain the fused latent space representation.
[0069] In this embodiment, the multi-head attention module adopts a scaled dot product attention mechanism, the calculation formula of which is:
[0070] in, , and These represent the query matrix, key matrix, and value matrix, respectively. The dimension of the key vector, divided by Used to prevent excessively high electrode values from causing... Gradient vanishing.
[0071] Multi-head attention projects the input to h Attention is computed in parallel in three different subspaces, and the formula is as follows:
[0072]
[0073] in, , , and Both represent learnable projection matrices. h This indicates the number of attention heads.
[0074] In multimodal fusion, EEG features As a query, multimodal concatenation features (Including bone conduction characteristics) and acceleration characteristics Using these as keys and values, cross-modal feature fusion is achieved:
[0075] in, This represents the latent space representation after merging. This indicates multi-head attention computation. Representing EEG characteristics, i.e., time-weighted features. This represents multimodal splicing features, including bone conduction features. and acceleration characteristics , Indicates adding These are residual connections used to preserve the original EEG features.
[0076] S6. Construct a state-space model and / or a dynamic causal model to model the system state evolution during speech stimulation and speech imagination, as well as the information flow between the frontal lobe, temporal lobe, and auditory cortex of the ear, and estimate the causal connection parameters between regions; wherein, the hierarchical transmission topology of the scalp-periauricular-intraauricular multi-level joint model provides prior structural constraints for the state-space model and / or the dynamic causal model.
[0077] In this embodiment, dynamic system modeling and parameter estimation specifically include: constructing a state-space model and / or a dynamic causal model, modeling the system state evolution and information flow between the frontal lobe, temporal lobe and auditory cortex of the ear region during speech stimulation and speech imagination, and estimating the intensity of causal influence between regions.
[0078] In this embodiment, the dynamic causal model quantitatively describes the information flow between the frontal lobe, temporal lobe, and auditory cortex of the ear region during speech imagination, and reconstructs the dynamic causal path of the speech processing network.
[0079] In this embodiment, the state equation and observation equation of the state-space model are as follows:
[0080]
[0081] in, express The hidden state vector at time t. for The hidden state vector at time t represents the internal state of the speech processing network. Represents the state transition matrix. Represents the input matrix, For external input, representing verbal stimuli or task instructions. Indicates process noise. Represents the observation vector. Represents the observation matrix. Indicates a direct transmission matrix. Indicates observation noise. and Assume it follows a Gaussian distribution with zero mean; The continuous-time form of the dynamic causal model is as follows:
[0082] in, This represents the derivative of the neural activity state vector of each brain region with respect to time. Represents the intrinsic connectivity matrix, whose elements Indicates the region without external modulation To the area The baseline connectivity strength, where m represents the number of experimental conditions. Indicates the first j Switching variables for each experimental condition. Indicates the first j The modulation matrix of the connection is affected by various experimental conditions (e.g., speech imagination vs. speech perception). This represents the neural activity state vectors of various brain regions (frontal lobe, temporal lobe, auditory cortex of the ear). This represents the input matrix, specifying which regions are driven by external stimuli. Indicates external input.
[0083] The expression for the causal connection parameters between regions is as follows:
[0084] in, Indicates time t Time zone j To the area i Causal connection parameters (effective connection strength). Indicates the area j To the area i The inherent (baseline) connectivity strength, k Index indicating experimental conditions (i.e., index number) k (The experimental conditions are listed below), where m represents the total number of experimental conditions. Indicates time t Time k Switching variables (external modulation input) for each experimental condition. Indicates the first k Experimental conditions for the region j To the area i Modulation parameters for the connection strength between them.
[0085] By comparing and estimating the parameters of the Bayesian model, the parameters of the above model are solved, thereby quantifying the direction and intensity of the information flow from the frontal lobe to the temporal lobe to the auditory cortex of the ear during speech imagination.
[0086] S7. Based on the fused latent space representation and estimated dynamic causal connection parameters, train the speech imagery decoder to output speech category labels, semantic information and / or reconstructed speech features for speech imagery-related EEG signals, and complete the analysis of the data.
[0087] In this embodiment, speech imagery decoding and output specifically includes: training a speech imagery decoder based on the fused latent space representation and estimated dynamic causal connection parameters, and outputting speech category labels, semantic information and / or reconstructed speech features to the EEG signals related to speech imagery, thereby achieving high-precision decoding of speech imagery EEG.
[0088] Example 2 To support the implementation of the data analysis method in Example 1, this invention further designs a matching ear-region brain-computer interface system and a speech imagination experimental paradigm. At the system level, this invention employs a dual-microcontroller structure to achieve synchronous acquisition of intra-ear, periauricular, and bone conduction / accelerometer signals and audio stimulus playback, providing high-precision temporal synchronization and multimodal data interfaces, thus providing a reliable hardware foundation for multi-level modeling. At the experimental paradigm level, this invention constructs multi-condition comparison tasks around typical speech processing stages such as "listening to speech - speech imagination - silent reading - spoken reading," systematically manipulating semantic complexity and cognitive load to adapt the acquired data to the analysis method proposed in this invention, achieving collaborative modeling of cognitive conditions, spatial distribution, and multimodal features. Furthermore, based on existing scalp-intra-ear mutual information modeling work, this invention explicitly incorporates the periauricular region into the modeling structure, constructing a three-level joint model of scalp-periauricular-intra-ear, and extending from static correlation to nonlinear, non-stationary dynamic system modeling. By using state-space modeling and dynamic causal modeling, this invention can track the temporal ordered information flow between the frontal lobe, temporal lobe and auditory cortex of the ear, and reconstruct the signal transmission path in the process of speech imagination, thereby achieving structured and interpretable analysis of speech imagination EEG data.
[0089] like Figure 4 As shown, this invention provides a weak electrical signal analysis system based on multi-level multimodal fusion, used to execute the weak electrical signal analysis method based on multi-level multimodal fusion described in Example 1, including an ear region sensing module, a data acquisition and control module, and a data processing and decoding module: The ear region sensing module includes: Ear area device, used to be fixed in the periauricular region of the subject's ear and to carry various sensing and actuation modules; An intra-ear electrode assembly, placed inside the ear canal, is used to collect electroencephalogram (EEG) signals from the intra-ear region and / or to play audio. Periauricular electrode assembly, distributed along the circumference of the auricle, is used to collect electroencephalogram (EEG) signals from the periauricular region; The bone conduction and acceleration module, fixed at the mastoid process, is used to output bone conduction speech stimulation and collect bone conduction vibration and acceleration signals; The clock synchronization module is used to synchronize the time of electroencephalogram (EEG) signals, bone conduction vibration signals, and acceleration signals in the scalp region, periauricular region, and intraauricular region. The data acquisition and control module adopts a dual microcontroller module, which serves as the master control end and the slave control end respectively. The master control end is used for audio playback control, data management and wireless transmission, while the slave control end is used for multi-channel physiological signal acquisition and preprocessing. The two are connected through a high-speed data cable to achieve timing synchronization. The main control unit includes a power module, a mode switching module, a wireless communication module, and a data synchronization module, used for audio playback control, data management, and wireless transmission; the slave control unit includes a power module, a data acquisition module, a data preprocessing module, and a data synchronization module, used for multi-channel physiological signal acquisition and preprocessing; the left and right dual microcontroller modules control the switching between data acquisition mode and audio playback mode through the mode switching module, and automatically control the opening and closing of the in-ear electrode acquisition channel and record the timestamp of the audio stimulation during the switching to support the time stamping of multi-condition tasks; The communication module is used to transmit synchronized multimodal data to an external computing device; The data processing and decoding module is configured on an external computing device to execute the aforementioned weak electrical signal analysis method based on multi-level multimodal fusion and output the speech imagery decoding results to complete the data analysis.
[0090] In this embodiment, the main body of the ear area device is a support structure that bypasses the ear. The upper end of the support is provided with a connecting rod for connecting the inner ear components, and the lower end is provided with an insulated ear clip for holding the earlobe. The bone conduction and acceleration module is fixedly installed at the lower 1 / 3 position of the support, and the left and right dual microcontroller modules are detachably installed at the upper 1 / 3 position of the support. This ensures the accuracy of mastoid positioning while balancing the overall weight and improving wearing comfort.
[0091] In this embodiment, the in-ear electrode assembly is a hollow cylindrical structure with electrodes arranged in four rings from the inside to the outside along the outer surface of the earplug. The number of electrode points in the four rings are 4, 8, 12 and 16 respectively. The assembly is made of Ag / AgCl or MXene-based conductive materials. The in-ear assembly can be configured into large, medium and small sizes according to the size of the subject's ear canal.
[0092] In this embodiment, the left and right dual microcontroller modules control the switching between data acquisition mode and audio playback mode through a mode switching module. During switching, the opening and closing of the in-ear electrode acquisition channel is automatically controlled and the timestamp of the audio stimulus is recorded to support the time stamping of multi-condition tasks.
[0093] In this embodiment, in order to provide training and validation data for the above data analysis method, the present invention also proposes a speech imagination experimental paradigm, a speech imagination experimental paradigm based on a multi-level multimodal fusion weak electrical signal analysis method, including: A1. Semantic Paradigm Design: Select several independently presented keywords and construct sentences with complete semantics through fixed arrangements to balance phonetic features, semantic complexity, and cognitive load; and divide them into randomly arranged keyword sequences and sentence sequences that constitute complete sentence meanings according to experimental needs; A2. Standardized speech perception acquisition: Subjects wore ear area devices set to audio playback mode and listened to speech stimuli consisting of randomly arranged keywords and complete sentences. Rest periods of no less than a preset duration were set between each group of stimuli. The speech stimuli and corresponding multi-region and multi-modal physiological signals were recorded and labeled through a clock synchronization module. A3. Random sequence speech acquisition: Subjects wore an in-ear component set to acquisition mode, and performed at least two different cognitive depth processing on randomly arranged keywords, including speech imagination, silent reading and / or spoken reading, while simultaneously acquiring EEG signals, bone conduction vibration signals and acceleration signals from the scalp area, periauricular area and in-ear area. A4. Complete Sentence Speech Acquisition: Using the same equipment configuration as the random sequence speech generation stage, the stimuli are upgraded to sentences with complete semantics. The speech imagination, silent reading and / or spoken reading tasks are repeatedly performed to obtain multi-region multimodal signals containing contextual information. A5. Data Quality and Labeling: Pre-experiment training is conducted before the formal experiment. In the formal experiment, real-time data quality monitoring is used to identify and remove severe artifacts. Data from each time period is bound to the corresponding task conditions, semantic labels, and behavioral responses to form a multi-region multimodal EEG dataset adapted to the weak electrical signal analysis method based on multi-level multimodal fusion.
[0094] In this embodiment, as Figure 5 As shown, this speech imagination experimental paradigm includes a pre-training phase and the complete experimental steps.
[0095] The pre-training (familiarization) phase is used to familiarize participants with the experimental procedure, and involves completing the following four phases in sequence: Listening to random sounds: passively listening to random speech stimuli played by the system; Silently reciting random words: silently reciting the words displayed on the screen; Silently reading random words: making mouth shapes to pronounce the words but reading them aloud without making a sound; Reading random words aloud: reading them aloud normally. The complete experimental procedure includes three experiments: Experiment 1: Speech Perception Task. This task was used to collect multi-regional EEG responses under passive auditory conditions. The specific procedure was as follows: randomly play words → short rest → play complete sentences → long rest, repeating multiple times for data collection.
[0096] Experiment 2: Random Sequence Speech Generation Task. Subjects were given three tasks in sequence for randomly presented words: silent reading → short rest → silent reading → short rest → audio reading → long rest. Multiple sets were repeated for data collection.
[0097] Experiment 3: Complete Sentence Speech Generation Task. The stimulus material was upgraded to sentences with complete semantics. Participants performed the following sequence: silent reading → short rest → silent reading → short rest → audio reading. Multiple sets were repeated to collect multi-regional and multimodal physiological signals containing contextual information.
[0098] In this embodiment, the task conditions include performing at least two of the following states sequentially on the same semantic material: a) auditory reception state: passively listening to randomly arranged or complete sentences of speech stimuli; b) speech imagination state: silently reciting the speech content without vocalization; c) silent reading state: the articulatory organs are involved but no sound is produced; d) vocal reading state: accompanied by actual vocal output.
[0099] Pre-experiment training was conducted before the formal experiment to familiarize the subjects with the task process. In the formal experiment, real-time data quality monitoring was used to identify and remove severe artifacts. Data from each time period was bound to the corresponding task conditions, semantic labels and behavioral responses to form a dataset adapted to the weak electrical signal analysis method based on multi-level multimodal fusion described in Example 1.
[0100] The present invention proposes a method and system for analyzing weak electrical signals based on multi-level multi-modal fusion. The main advantages of the present invention are: (1) Multi-level spatial modeling is more complete and the results are more interpretable. For the first time, this invention regards the scalp-periauria-inner ear as three-level nodes of the same speech processing network, constructs a multi-level joint modeling framework, explicitly describes the transmission and transformation of speech-related neural activities between layers, avoids simply treating ear signals as a single alternative channel, and makes the EEG decoding results of speech imagination more structured and interpretable.
[0101] (2) Multimodal nonlinear feature fusion significantly improves decoding accuracy and robustness. Under a unified framework, multimodal signals such as scalp, periauricular, intraauricular EEG, bone conduction, and acceleration are fused. Time-frequency analysis, principal component analysis, and deep neural networks are introduced to achieve nonlinear extraction and adaptive fusion of speech imagery features. This can maintain a high recognition rate and good generalization performance under noise interference and individual differences.
[0102] (3) Dynamic causal modeling helps to reveal the neural mechanism of speech imagination. By using state-space modeling and dynamic causal modeling to quantitatively characterize the temporal ordered information flow between the frontal lobe, temporal lobe and auditory cortex of the ear, a new analytical approach is provided for understanding the dynamic network mechanism of speech imagination and its evolution.
[0103] (4) The system is deeply coupled with the experimental paradigm and method to improve data quality and application feasibility. The dedicated ear area system supports multimodal synchronous acquisition and rapid switching of playback. With the progressive task design of "listening-thinking-silent reading-audio reading", the data is highly matched with the analysis method in terms of time stamp, spatial distribution and cognitive conditions, which is conducive to the subsequent application and transformation in rehabilitation training and auxiliary communication.
[0104] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0105] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0106] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
[0107] The above detailed description further illustrates the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above description is merely a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A weak electrical signal analysis method based on multi-level multi-modal fusion, characterized in that, The weak electrical signal is specifically a multi-regional electroencephalogram (EEG) signal, and the method includes the following steps: S1. Under task conditions, weak EEG signals from the scalp region, periauricular region, and inner ear region are collected simultaneously, along with bone conduction vibration signals and acceleration signals, to obtain multi-regional multimodal signals. S2. Preprocess the acquired multi-region multimodal signals, obtain the time-frequency spectrum representation of each region by performing time-frequency joint analysis, and perform frequency band division and normalization processing; S3. Construct a multi-level joint model of scalp-periauria-inner ear; S4. Based on the normalized time-frequency representation, using the scalp-periauria-intraauria multi-level joint model, nonlinear features are extracted to obtain scalp feature vectors, periauria feature vectors, and intraauria feature vectors. S5. Input the scalp feature vector, periauricular feature vector, intraauricular feature vector, bone conduction vibration signal and acceleration signal into the deep neural network to obtain the fused latent space representation; S6. Construct a state-space model and / or a dynamic causal model to model the system state evolution and information flow between the frontal lobe, temporal lobe and auditory cortex of the ear during speech stimulation and speech imagination, and estimate the causal connection parameters between regions; wherein, the hierarchical transfer topology of the scalp-periauria-intraauria multi-level joint model provides a priori structural constraints for the state-space model and / or dynamic causal model. S7. Based on the fused latent space representation and estimated dynamic causal connection parameters, train the speech imagery decoder to output speech category labels, semantic information and / or reconstructed speech features for speech imagery-related EEG signals, and complete the analysis of the data.
2. The multi-stage multi-modal fusion based weak electrical signal analysis method according to claim 1, wherein, Perform time-frequency joint analysis using at least one of the following methods: Short-time Fourier transform of EEG signals using overlapping time windows yields a time-spectrum representation; in, Indicates the time of the original EEG signal t and frequency f The time-spectral representation of the short-time Fourier transform at a given point. Represents raw brain electrical signals. Indicated by t The time window function centered on j Represents the imaginary unit. f Represents frequency components, t Indicates the center position of the time window. T Indicates the total duration. d Represents an integral infinitesimal operator; Continuous or discrete wavelet transforms are used to construct time-spectrum representations, enhancing the ability to represent transient speech-related neural oscillations. wherein denotes a time-frequency representation, denotes a scale parameter, denotes a translation parameter, denotes a complex conjugate, denotes an original electroencephalogram signal.
3. The multi-stage multi-modal fusion based weak electrical signal analysis method of claim 1, wherein, S3 includes the following steps: A hierarchical transfer function is constructed from the scalp region to the periauricular region and from the periauricular region to the inner ear region. The periauricular region is explicitly introduced as an intermediate transfer link. The signals from the scalp, periauricular region and inner ear are modeled as multi-level outputs driven by the same speech-related potential neural activity. In the multi-level joint model of scalp-periauricular region-inner ear, the transfer noise at each level is modeled and suppressed separately to obtain a multi-level joint model describing the complete transfer path of scalp-periauricular region-inner ear. The construction of the multi-level joint model of scalp-periauricular region-inner ear is completed.
4. The multi-stage multi-modal fusion based weak electrical signal analysis method according to claim 3, characterized in that, The expression for the multi-level joint model of scalp-periauricular-intraauricular region is as follows: wherein, represents the ear region at time the electroencephalogram signal, represents the combined error term, represents the periauricular region electroencephalogram signal, represents the scalp region electroencephalogram signal, represents the composite function.
5. The multi-stage multi-modal fusion based weak electrical signal analysis method according to claim 1, wherein, S4 includes the following steps: The time-spectral representations of each region after normalization are mapped to construct a high-dimensional feature space; Based on the high-dimensional feature space, the principal component features of each region are extracted by solving the eigenvalues and eigenvectors of the kernel matrix; The extracted principal component features of each region are denoised to obtain scalp feature vectors, periauricular feature vectors, and intraauricular feature vectors.
6. The multi-stage multi-modal fusion based weak electrical signal analysis method according to claim 1, wherein, The deep neural network includes: The first feature extraction subnetwork includes several cascaded convolutional and pooling layers, which are used to extract local patterns based on scalp feature vectors, periauricular feature vectors, intraauricular feature vectors, bone conduction vibration signals and acceleration signals, and output the extracted local pattern feature sequence to the second feature modeling subnetwork. The second feature modeling subnetwork is used to process time series dependencies based on local pattern feature sequences and weight the contributions of different time steps to obtain time-weighted features. The time-weighted features are then used as EEG features and input into the feature fusion subnetwork. The feature fusion subnetwork, including a multi-head attention module and a fully connected layer, maps scalp feature vectors, periauricular feature vectors, intraauricular feature vectors, bone conduction vibration signals, and acceleration signals to a unified space. Based on the mapping results, temporally weighted features, and the multi-head attention mechanism, it performs nonlinear fusion of multi-region and multi-modal features to obtain the fused latent space representation. wherein, represents the fused hidden space representation, represents the multi-head attention computation, represents the electroencephalography features, i.e. the time- weighted features, represents the multi-modal concatenation features, including bone conduction features and acceleration features , represents the addition of is a residual connection, used to preserve the original electroencephalography features. The decoding output layer is used to output speech category labels, semantic information, and / or reconstructed speech features based on the fused latent space representation.
7. The multi-stage multi-modal fusion based weak electrical signal analysis method according to claim 1, wherein, The state equations and observation equations of the state-space model are as follows: wherein, represents the hidden state vector at time instant represents the hidden state vector at time instant represents the state transition matrix, represents the input matrix, represents the external input, represents the process noise, represents the observation vector, represents the observation matrix, represents the direct transmission matrix, represents the observation noise; The continuous-time form of the dynamic causal model is as follows: in, This represents the derivative of the neural activity state vector of each brain region with respect to time. Let m represent the intrinsic connectivity matrix, and m represent the number of experimental conditions. Indicates the first j Switching variables for each experimental condition. Indicates the first j Each experimental condition corresponds to the modulation matrix of the connection. This represents a vector representing the neural activity state of each brain region. Represents the input matrix, Indicates external input; The expressions for the causal connection parameters between the regions are as follows: wherein, denotes the time point t denotes the time zone j denotes the zone i denotes the causal connection parameter between the zones denotes the zone j denotes the intrinsic connection strength between the zones i denotes the index of the experimental condition k denotes the time point denotes the switch variable of the nth t experimental condition k denotes the modulation parameter of the connection strength between the zones and the zones k by the nth j experimental condition i .
8. A weak electrical signal analysis system based on multi-level multi-modal fusion, configured to perform the weak electrical signal analysis method based on multi-level multi-modal fusion according to any one of claims 1-7. include: Ear area device, used to be fixed in the periauricular region of the subject's ear and to carry various sensing and actuation modules; An intra-ear electrode assembly, placed inside the ear canal, is used to collect electroencephalogram (EEG) signals from the intra-ear region and / or to play audio. Periauricular electrode assembly, distributed along the circumference of the auricle, is used to collect electroencephalogram (EEG) signals from the periauricular region; The bone conduction and acceleration module, fixed at the mastoid process, is used to output bone conduction speech stimulation and collect bone conduction vibration and acceleration signals; The system consists of two microcontroller modules, one on the left and one on the right, serving as the master and the other on the right. The master module is used for audio playback control, data management, and wireless transmission, while the slave module is used for multi-channel physiological signal acquisition and preprocessing. The master and slave modules are connected via a high-speed data cable to achieve timing synchronization. The clock synchronization module is used to synchronize the time of electroencephalogram (EEG) signals, bone conduction vibration signals, and acceleration signals in the scalp region, periauricular region, and intraauricular region. The communication module is used to transmit synchronized multi-region multimodal signals to external computing devices; An external computing device is configured to execute the aforementioned weak electrical signal analysis method based on multi-level multimodal fusion and output the speech imagery decoding results to complete the analysis of the data.
9. The weak electrical signal analysis system based on multi-stage multi-modal fusion according to claim 8, characterized in that, The main body of the ear area device is a support structure that bypasses the ear. The upper end of the device is provided with a connecting rod for connecting the inner ear components, and the lower end is provided with an insulated ear clip for holding the earlobe. A bone conduction and acceleration module is fixedly installed at a preset position below the support, and left and right dual microcontroller modules can be detachably installed at a preset position above the support. The in-ear electrode assembly is a hollow cylindrical structure with electrodes arranged in four rings from the inside to the outside along the outer surface of the earplug, and is made of conductive material; The left and right dual microcontroller modules control the switching between data acquisition mode and audio playback mode through a mode switching module. During switching, the opening and closing of the in-ear electrode acquisition channel is automatically controlled and the timestamp of the audio stimulus is recorded to support the time stamping of multi-condition tasks.
10. The multi-stage multi-modal fusion based weak electrical signal analysis method of claim 1, wherein, The acquisition of multi-region multimodal signals in step S1 is specifically obtained through the following speech imagination experiment paradigm, including: A1. Semantic Paradigm Design: Select several independently presented keywords and construct sentences with complete semantics through fixed arrangements to balance phonetic features, semantic complexity and cognitive load, and divide them into randomly arranged keyword sequences and sentence sequences that constitute complete sentence meanings according to experimental needs; A2. Standardized speech perception acquisition: Subjects wore ear area devices set to audio playback mode and listened to speech stimuli consisting of randomly arranged keywords and complete sentences. Rest periods of no less than a preset duration were set between each group of stimuli. The speech stimuli and corresponding multi-region multimodal physiological signals were recorded and labeled through a clock synchronization module. A3. Random sequence speech acquisition: Subjects wore an in-ear component set to acquisition mode, and performed at least two different cognitive depth processing on randomly arranged keywords, including speech imagination, silent reading and / or spoken reading, while simultaneously acquiring EEG signals, bone conduction vibration signals and acceleration signals from the scalp area, periauricular area and in-ear area. A4. Complete Sentence Speech Acquisition: Using the same equipment configuration as the random sequence speech generation stage, the stimuli are upgraded to sentences with complete semantics. The speech imagination, silent reading and / or spoken reading tasks are repeatedly performed to obtain multi-region multimodal signals containing contextual information. A5. Data Quality and Labeling: Pre-experiment training is conducted before the formal experiment. In the formal experiment, real-time data quality monitoring is used to identify and remove severe artifacts. Data from each time period is bound to the corresponding task conditions, semantic labels, and behavioral responses to form a multi-region multimodal EEG dataset adapted to the weak electrical signal analysis method based on multi-level multimodal fusion.