Interoceptive neural coding-based brain-computer interface system and method for cardiac-evoked potentials

CN122805296APending Publication Date: 2026-09-25INST OF ACUPUNCTURE & MOXIBUSTION CHINA ACADEMY OF CHINESE MEDICAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610692615.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

HEP作为反映内感受加工的关键神经电生理指标,与情绪调节、焦虑抑郁等情志状态高度相关,但现有研究多依赖实验室高密度脑电,难以在便携设备上实现稳定提取与闭环应用

Benefits of technology

本申请首次实现基于内感受神经编码的便携式脑机接口,拓展了BCI的控制维度与应用场景;为抑郁症、焦虑症等情志病提供可量化、可干预的神经生理靶点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122805296A_ABST
    Figure CN122805296A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of brain-computer interfaces and nerve regulation, in particular to a heart beat evoked potential brain-computer interface system and method based on interoceptive neural coding. The heart beat evoked potential brain-computer interface system comprises an HEP signal extraction and feature coding system, which is used for signal extraction and feature coding on received space-time synchronous aligned electroencephalogram signals and cardiovascular signals, and outputs a stable heart beat evoked potential HEP waveform sequence, an interoceptive original feature vector and a signal quality index; an HEP high-dimensional time sequence feature extraction system, which is used for outputting an HEP mixed feature; an emotion state decoding system, which is used for outputting an emotion state probability vector; and a nerve regulation system, which is used for generating a pulse sequence parameter for stimulating nerves according to an emotion state category corresponding to the emotion state probability vector under the condition that a preset stimulation condition is met. The application first realizes a portable brain-computer interface based on interoceptive neural coding, and expands the control dimension and application scene of the BCI.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of brain-computer interface and neuromodulation technology, and in particular to a brain-computer interface system and method based on endoreceptor neural coding of cardiac evoked potentials. Background Technology

[0002] Brain-computer interface (BCI) technology captures brain signals and converts them into electrical signals, directly establishing a communication and control channel between the brain and external devices to achieve information transmission and control. Heartbeat-evoked potentials (HEPs), as key neurophysiological indicators reflecting interoceptive processing, are highly correlated with mood regulation, anxiety, depression, and other emotional states.

[0003] Currently, brain-computer interface technologies mostly use overt behavior or cortical potentials as control signals, lacking the quantification and utilization of "interoception"—the brain's perception and processing of visceral signals (such as heartbeat). HEP, as a key neurophysiological indicator reflecting interoceptive processing, is highly correlated with mood regulation, anxiety, depression, and other emotional states; however, existing research largely relies on high-density EEG in the laboratory, making it difficult to achieve stable extraction and closed-loop application on portable devices.

[0004] This invention aims to solve the aforementioned problems in brain-computer interface technology and HEP applications, and to realize a non-invasive ear canal brain-computer interface system with cardiac evoked potentials as the core state variable, and its closed-loop neuromodulation method based on interoceptive coding. Summary of the Invention

[0005] In order to at least partially address the aforementioned technical problems, this application provides a brain-computer interface system and method based on endoreceptive neural coding of cardiac evoked potentials.

[0006] In a first aspect, this application provides a brain-computer interface system for cardiac evoked potentials based on interoceptive neural coding, the brain-computer interface system for cardiac evoked potentials comprising: The HEP signal extraction and feature encoding system is used to extract and encode features from received spatiotemporally synchronized EEG and cardiovascular signals, and output stable cardiac evoked potential HEP waveform sequences, original interoceptive feature vectors, and signal quality indicators. The HEP high-dimensional temporal feature extraction system is used to output HEP hybrid features based on stable HEP waveform sequences, intrinsic intuition original feature vectors, and signal quality indicators. The emotional state decoding system is used to encode HEP hybrid features using a dual-branch asymmetric state space coding structure, generate latent feature vectors, and then map the latent feature vectors to a pre-constructed emotional state semantic feature matrix to output an emotional state probability vector. The neural modulation system is used to generate pulse sequence parameters for stimulating nerves based on the emotional state category corresponding to the emotional state probability vector, under preset stimulation conditions.

[0007] Secondly, this application provides a method for heartbeat-evoked neural modulation based on interoceptive neural coding, wherein the heartbeat-evoked neural modulation method is as follows: The received spatiotemporally synchronized EEG and cardiovascular signals are processed for signal extraction and feature encoding, and output stable heartbeat evoked potential (HEP) waveform sequences, original interoceptive feature vectors, and signal quality indices. Based on the stable HEP waveform sequence, the original feature vector of the intrinsic perception, and the signal quality index, output the HEP hybrid features; A dual-branch asymmetric state-space coding structure is used to encode the HEP hybrid features to generate latent feature vectors. After mapping the latent feature vectors to a pre-constructed emotional state semantic feature matrix, the emotional state probability vectors are output. Under preset stimulation conditions, pulse sequence parameters for stimulating nerves are generated based on the emotional state category corresponding to the emotional state probability vector.

[0008] The technical solution provided in this application has the following advantages compared with the prior art: This application is the first to realize a portable brain-computer interface based on interoceptive neural coding, which expands the control dimensions and application scenarios of BCI; and provides quantifiable and interventionist neurophysiological targets for mood disorders such as depression and anxiety. Attached Figure Description

[0009] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 These are schematic diagrams of the cardiac evoked potential brain-computer interface systems based on interoceptive neural coding in various embodiments of this application; Figure 2 These are schematic diagrams of the HEP signal extraction and feature encoding systems of various embodiments of this application; Figure 3 These are schematic diagrams of the HEP high-dimensional temporal feature extraction systems of various embodiments of this application; Figure 4These are schematic diagrams of the residual timing modeling modules in various embodiments of this application; Figure 5 These are schematic diagrams of the emotional state decoding systems of various embodiments of this application; Figure 6 This is a schematic diagram of the neural modulation system in various embodiments of this application. Specific Implementation It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0013] In the following description, the use of suffixes such as "module," "component," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "component," or "unit" may be used interchangeably. Terms such as "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are used only for the convenience of describing this application and simplifying the description. They do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as limiting this application. Terms such as "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0014] Example 1 This invention provides a brain-computer interface system based on endoreceptive neural coding of cardiac evoked potentials, such as... Figure 1 As shown, the cardiac evoked potential brain-computer interface system includes: HEP signal extraction and feature encoding system 10 is used to extract and encode features of received spatiotemporally synchronized EEG and cardiovascular signals, and output stable cardiac evoked potential HEP waveform sequence, endoreceptive feature vector and signal quality index. The HEP high-dimensional temporal feature extraction system 20 is used to output HEP hybrid features based on stable HEP waveform sequences, intrinsic intuition original feature vectors, and signal quality indicators. The emotional state decoding system 30 is used to encode the HEP hybrid features using a dual-branch asymmetric state space coding structure, generate a latent feature vector, and then map the latent feature vector to a pre-constructed emotional state semantic feature matrix to output an emotional state probability vector. The neural modulation system 40 is used to generate pulse sequence parameters for stimulating nerves based on the emotional state category corresponding to the emotional state probability vector under preset stimulation conditions.

[0015] In some embodiments, a stable HEP (Heartbeat Evoked Potential) signal extraction and feature encoding system, such as Figure 2 As shown, this system uses multi-channel EEG and synchronized cardiovascular signals from the ear canal as input layers. Through pipelined processing including spatiotemporal synchronization verification, adaptive artifact suppression, heart rate-dependent baseline correction, continuous weighted superposition of information entropy, and online smooth updates, it outputs a stable HEP waveform sequence, original feature vectors, and signal quality indices (VR, SNR). These three types of outputs will serve as multimodal inputs to the "HEP High-Dimensional Temporal Feature Extraction System," used to further construct high-dimensional HEP hybrid features and complete emotional state decoding.

[0016] The input layer synchronously receives two continuous physiological signals: one is a continuous electroencephalogram (EEG) signal acquired from the ear. ,in C The number of ear canal electrode channels, with a sampling rate of Second, synchronous cardiovascular (ECG) signals. The two signals are synchronized and aligned in time and space at the hardware acquisition end. To facilitate subsequent time-locked interception based on the R-wave (cardiac event), the R-wave time sequence is defined as follows: and in the k Introducing the concept of a relative time axis within the time-locked window of the secondary heartbeat. ,in and These represent the duration of the cut-off points before and after the R wave, respectively.

[0017] In the EEG / cardiovascular signal preprocessing layer, to preserve the main frequency bands of the high-frequency epithelial spectroscopy (HEP) and suppress environmental noise and power line interference, notch filtering and bandpass filtering are applied to EEG and cardiovascular signals, respectively. For EEG signals... The filter is applied sequentially using a high-pass filter, a notch filter, and a low-pass filter, with a cutoff frequency of [value missing]. High-pass filtering (removing slow drift), center frequency is And quality factor is The power frequency notch, and the cutoff frequency is Low-pass filtering was applied to obtain the preprocessed EEG signal. For cardiovascular signals (i.e., electrocardiogram signals) The upper and lower cutoff frequencies are respectively and The bandpass filter is obtained This is used for subsequent R-wave detection.

[0018] The dual-path processing layer mainly performs adaptive cancellation of ECG propagation artifacts and suppression of electrooculography artifacts on the preprocessed EEG signals to reconstruct purified EEG signals; and determines the cardiac event time sequence (R-wave time sequence) of the preprocessed cardiovascular signals.

[0019] During the R-wave confidence assessment phase, the preprocessed cardiovascular signals were analyzed using an improved Pan-Tompkins procedure. The process involves differentiation, squaring, sliding window integration, and adaptive threshold decision to obtain the R-wave time sequence. To quantify the reliability of each detection, the following definition is made: R-wave confidence The RR interval sequence is obtained by Sigmoid mapping based on indices such as candidate peak height, local QRS energy, and consistency with historical RR intervals, and is used for subsequent quality weight calculation. The RR interval sequence is also defined as follows: (001) It can be further used for heart rate estimation and heart rate-dependent baseline correction.

[0020] During the adaptive cancellation of ECG propagation artifacts (CFA-LMS) and suppression of electrooculography (EOG) artifacts, ECG propagation artifacts and EOG artifacts in ear canal EEG acquisition significantly contaminate weak HEP waveforms. To improve SNR without additional electrodes, some implementations employ a combined strategy of "ECG reference adaptive cancellation + multi-channel blind source separation." The preprocessed ECG signal is used as an example. As a reference input, an LMS adaptive filter is established for each EEG signal channel to estimate and cancel out ECG propagation artifacts from the EEG signal. For the ... There are channels, and the adaptive filter weights are denoted as . The reference vector for the electrocardiogram signal is Then the error signal (i.e., the output after artifact removal) is: (002) The weights have been updated to: (003) in This is the step size. All channels... Multichannel EEG after removing ECG propagation artifacts Further investigation Blind source separation is performed (such as ICA / CCA methods or their lightweight implementations) to identify and suppress the source components corresponding to EEG / EMG / motor artifacts, and then reconstruct the purified EEG signal. (004) in Represents the decomposition operator, This represents the reconstruction operator after removing artifact sources.

[0021] The quality gating layer includes adaptive time-locked segmentation, heart rate-dependent adaptive baseline correction, and continuous weighted and continuous multi-level quality (QC) assessment based on information entropy. It is used to adaptively lock and segment the purified EEG signal according to the cardiac event time sequence, extracting cardiac evoked potential waveform segments; applying heart rate-dependent adaptive baseline correction to continuous multi-segment cardiac evoked potential waveforms to obtain baseline-corrected cardiac evoked potential waveform segments; and gating the baseline-corrected cardiac evoked potential waveform segments with R-wave confidence based on quality weights generated by pre-constructed QC quality indicators. In the adaptive time-locked segmentation and heart rate-dependent adaptive baseline correction stages, the R-wave time is used as the basis for the assessment. Zero point, for purifying brainwave signals Intercepting during locking to obtain the first HEP segments of the next heartbeat (i.e., segments of the heartbeat evoked potential waveform): (005) In natural scenarios, heart rate fluctuations can cause waveform overlap and baseline drift between adjacent heartbeats. To address this, a heart rate-dependent adaptive baseline correction is introduced: within each segment, a baseline window is used... Fit a second-order polynomial baseline and subtract the entire segment: (006) (007) When heart rate increases ( When the value decreases, the fitting window can be shortened accordingly or the robustness at the boundary can be improved to reduce the bias caused by the overlap effect.

[0022] In the information entropy weighting and continuous quality assessment stages, traditional HEP extraction often uses hard thresholding to remove artifact segments, which can easily lead to insufficient effective samples or introduce mutations in low signal-to-noise ratio scenarios. This invention proposes an information entropy weighting mechanism and continuous quality assessment, which performs R-wave confidence gating on segments of the baseline-corrected cardiac evoked potential waveform based on quality weights generated from pre-constructed QC quality indicators. Specifically, continuous quality weights are generated for each cardiac segment. This is used for subsequent weighted summation. For the ... For this segment, construct the following multi-level QC quality indicators: peak-to-peak value constraint. Baseline stability threshold Missing / saturation ratio integrity threshold With R-wave confidence Two further global quality metrics are defined for downstream System 2. The first is the Variance Ratio (VR), which characterizes the ratio of baseline energy to event window energy. A low VR indicates significant event-related components and good signal quality, while a high VR indicates that the event window does not differ significantly from the baseline. (008) Next is the signal-to-noise ratio (SNR), which characterizes the strength of the event-related components relative to the noise floor: (009) in Take the average power within the event window. Take the average power within the baseline window. It is a numerically stable term.

[0023] To reflect the degree of disorder / artifacts in waveform morphology, an information entropy index is constructed for the segments. The amplitude distribution of the baseline-corrected segments is discretized into... For each box, the probability distribution is obtained. Define entropy: (010) A larger value usually corresponds to stronger artifacts or greater noise. (In summary) , , , and Generate continuous quality weights Its gating method is as follows: (011) in For the Sigmoid function, This represents the normalized index. This is a configurable coefficient. This continuous weight replaces hard thresholding, allowing low-quality segments to be softly suppressed rather than discarded directly.

[0024] The aggregation layer is used to segment and smooth the baseline-corrected cardiac evoked potential (HEP) waveform, which is gated by R-wave confidence, using a weighted sliding window to obtain a stable HEP waveform sequence at the current moment. During the weighted sliding window superposition and online smoothing update phase, a layer with a length of [missing information] is used. Sliding window for recent The heartbeat segments are weighted and superimposed to obtain the estimated stable HEP waveform sequence at the current moment: (012) To further suppress abrupt changes caused by the sliding window boundary, an exponentially weighted moving average (EWMA) is introduced at the output for smoothing: (013) in This is a smoothing factor.

[0025] The output layer outputs the stable HEP waveform sequence at the current moment, the original intrinsic receptivity feature vector determined based on the stable HEP waveform sequence, and the signal quality index. After obtaining the stable waveform... Subsequently, the original feature vectors oriented towards introspective processing are extracted and combined with quality indicators to output the HEP high-dimensional temporal feature extraction system. Specifically, regarding temporal features, within a given analysis window... Internally defined HEP mean amplitude characteristics (reflecting the intensity of interoceptive processing): (014) Individual baseline amplitude Defined as the system initialization phase (the first M heartbeats collected in the resting state) corresponding to Mean: (015) The estimation is completed and stored when the system starts up, and slowly and adaptively updated with a low learning rate during long-term use, for use by the closed-loop triggering judgment of module three and PID target setting.

[0026] HEP latency characteristics (peak latency) (reflecting interoceptive conduction efficiency): (016) And the hemispherical asymmetry index (if the channel can be divided into left and right ear symmetrical channel groups) : (017) Regarding the frequency domain characteristics of HRV, the RR interval sequence Construct HRV sequences and estimate the power spectral density using the Welch method. Calculate the low-frequency and high-frequency energies and their ratio: (018) (019) in , These are the upper and lower limits of the low-frequency band. , These represent the upper and lower limits of the high-frequency band.

[0027] Regarding spatial / frequency domain coupling characteristics, the cross-channel coherence is calculated (in terms of channels). (For example) (020) in For cross-spectral density; and PAC exponent (in low-frequency phase) With high frequency amplitude (For example) (021) The system ultimately outputs three types of quantities: (1) Stabilize HEP waveform ; (2) Original feature vectors oriented towards introspection: (022) in The proportion of effective segments; (3) Signal quality indicators .

[0028] The above outputs will correspond to the stable HEP waveform, the original eigenvector, and the signal quality index in System 2, respectively. , In Module 2, high-dimensional temporal dynamic feature extraction, prior statistical feature extraction, and cross-modal interactive attention and quality gating fusion are completed to form a high-dimensional hybrid feature representation for use by the emotional state decoding module. When Module 2 receives the above quality indicators, they are denoted as... , , representing the real-time quality estimate at the current heartbeat moment.

[0029] In some embodiments, the HEP high-dimensional temporal feature extraction system is implemented as follows: Figure 3 As shown, high-dimensional temporal feature extraction of HEP can be achieved. This system takes multimodal physiological signals as input and, through multi-scale spatiotemporal feature extraction and quality adaptive fusion mechanisms, outputs a high-dimensional hybrid feature representation containing multidimensional semantic information, providing high-quality feature support for downstream emotional state recognition networks.

[0030] The system receives three types of input information: first, a stable HEP waveform sequence (given in the form of a one-dimensional time sequence); second, raw feature vectors containing statistical moments and frequency domain energy ratios; and third, two signal quality indicators reflecting the current signal acquisition quality: signal variance ratio (VR) and signal-to-noise ratio (SNR). Based on the above multimodal inputs, the system is internally divided into: a time-series dynamic feature extraction module, a residual time-series modeling module, a priori statistical feature extraction module, a cross-modal interactive attention calculation module, a quality gating factor generation unit, and a quality gating fusion unit.

[0031] For the temporal dynamic feature extraction module, a multi-scale one-dimensional convolutional network (1DCNN) is used to extract features from stable HEP waveforms in parallel through multiple branches. The extracted multi-branch features are then concatenated to obtain scale-stitched features. Specifically, using a preprocessed stable HEP waveform sequence (T being the temporal length) as input, a multi-scale one-dimensional convolutional network (1DCNN) is first used for parallel multi-granularity local feature extraction. This module has three parallel feature extraction branches: small kernel convolution, medium kernel convolution, and large kernel convolution. These three branches complement each other in their receptive fields. The small kernel branch focuses on capturing high-frequency local transient distortions (such as spikes, notches, and other detailed morphological features) in the HEP waveform; the large kernel branch focuses on the gradual changes in low-frequency global morphology; and the medium kernel branch takes both into account, extracting medium-scale transitional dynamic features. ReLU activation is performed after each convolution operation. The aim is to utilize different receptive fields to simultaneously capture high-frequency local transient distortions and low-frequency global morphological fluctuations in the HEP waveform. The feature maps output by each scale convolution branch are concatenated along the channel dimension in the feature concatenation layer as shown in Equation (101); (101) Scale splicing features The input feature sequence is the residual time series modeling module, which is as follows: Figure 4 As shown, a high-order temporal feature sequence is output through multi-level temporal dependency modeling, external memory network, and multi-scale feature fusion. This high-order temporal feature sequence is then processed through a layered window attention multi-layer stacked architecture to obtain a deep temporal dynamic feature vector. In detail, the scale-concatenated feature sequence... First, the data is fed into a bidirectional long short-term memory (LSTM) network. By introducing forget gates, input gates, and output gates, the gradient vanishing problem can be effectively alleviated during backpropagation, enabling the modeling of long-range temporal dependencies. Based on this, the bidirectional LSTM sets up an independent LSTM link in both the forward and reverse time directions, so that the network can simultaneously utilize the historical and future context information of the sequence, as shown in Equation (102). (102) in, Let be the input feature vector at time t. For the unidirectional hidden layer dimension, [·||·] indicates that the vectors are concatenated. The complete output feature sequence of a bidirectional LSTM. This structure can effectively capture the deep forward causal dependencies and backward correlation constraints between components in a HEP signal.

[0032] The feature sequence processed by bidirectional LSTM is passed sequentially through a dynamic feature association estimation layer and a dynamic feature convolutional layer to further extract the nonlinear mapping relationship between local features. The dynamic feature association estimation layer adaptively estimates the correlation of the current feature sequence in the time dimension and channel dimension through a learnable parameter matrix and outputs the association weight matrix, which is calculated as shown in Equation (103). (103) in, These are the projection matrices for the query and the key, respectively. The projection dimension is used to normalize the dot product to stabilize the gradient. The dynamic feature convolutional layer uses the aforementioned associated weight matrix. As the modulation coefficients of the adaptive convolution kernel, for the feature sequence The dynamic convolution operation is performed as shown in equation (104); (104) This mechanism enables the receptive field shape and weight distribution of the convolution kernel to dynamically adjust according to the local characteristics of the input signal, rather than using fixed weights. This adaptively enhances the feature representation of key temporal segments in complex HEP waveforms, outputting highly representative mid-to-high-level dynamic features. Features output by dynamic convolutional layers It is divided into two parallel paths: one is passed directly upwards as the main feature flow; the other is diverted to the external memory network to introduce global historical prior knowledge and enhance the semantic expressive power of the current features.

[0033] The internal processing flow of the external memory network is as follows: (1) The spectral clustering pattern extraction module extracts clustering patterns from the feature sequences output by the dynamic feature convolutional layer to determine typical temporal dynamic pattern prototypes. First, the input dynamic features... Construct a similarity matrix Its elements The Gaussian kernel similarity is used to represent the eigenvectors at time steps i and j. Based on this, the normalized Laplacian matrix is ​​calculated. (in The feature vector is a degree matrix, and its eigenvalues ​​are decomposed. The K smallest non-zero eigenvectors are used to form an embedding matrix. K-Means clustering is performed in this low-dimensional spectral space, and the features at T time steps are assigned to K clusters. The center vector of each cluster is calculated. This serves as a typical prototype of a time-series dynamic pattern, representing the current input.

[0034] (2) The memory matrix update module updates the preset memory matrix based on the typical time-series dynamic pattern prototype. The system maintains a global memory matrix. There are N memory slots in total, and each slot stores one record. Historical feature memory of dimensions. In each forward propagation, based on the currently extracted K pattern prototypes... The memory matrix is ​​updated by calculating the similarity between each prototype and each memory slot. The memory slots with the highest similarity are written and updated using an exponential moving average method, so that the memory matrix continuously accumulates the typical feature patterns of historical samples, forming a long-term knowledge base for the system.

[0035] (3) An attention-based aggregation module uses the global mean pooling result of the feature sequence output by the dynamic feature convolutional layer as the query, and extracts the memory feature sequence matching the current heartbeat state from the memory matrix through an attention-weighted retrieval mechanism. This is based on the core dynamic features. The global mean pooling result is used as the query target. The attention-weighted retrieval mechanism extracts the global context memory feature sequence (hereinafter referred to as the technical feature sequence) that is most relevant to the current signal state from the memory matrix, as shown in Equation (105); (105) in, For a trainable projection matrix, For the attention weights of each memory slot, This is the weighted aggregated global context memory feature vector. This mechanism can dynamically retrieve prior patterns highly matching the current heartbeat state from a long-term historical knowledge base using content addressing, effectively compensating for the limitations of insufficient information within a limited temporal window in a single forward inference. After completing external memory retrieval, the system enters a multi-scale feature fusion layer to deeply integrate three feature streams from different levels; path one comes from the output features of the dynamic feature convolutional layer. One path carries mid-to-high-level temporal dynamic information enhanced by bidirectional LSTM and dynamic convolution; the second path comes from the output of the external memory network. After being broadcast and expanded in the time dimension, it aligns with path one and carries the long-term global context prior accumulated by the system; path three is the feature sequence of a bidirectional LSTM module. This process preserves the underlying local details that have not undergone deep transformation. The three features are concatenated along the channel dimension within the fusion layer and then compressed to a unified dimension via a learnable linear projection layer, resulting in multi-scale fused features. This fusion mechanism simultaneously integrates three information levels: low-level details, mid-level temporal dynamic evolution, and high-level global historical memory, constructing a rich, multi-granular, global-local collaborative feature representation of the HEP signal state. To further alleviate the gradient vanishing problem during deep network training and preserve the baseline feature information of the original input signal, this system introduces a global residual connection at the end. The input feature sequence at the very front of the system... The output of the multi-scale feature fusion layer is obtained through a cross-layer main link. The element-wise addition is performed as shown in equation (106); (106) in, The feature sequence is output after summing the global residuals of the linear projection matrix for dimension matching. It simultaneously possesses two aspects of information security: firstly, the high-order semantic representations extracted by the deep modules are... The residual terms fully reflect the characteristics of the original input; on the other hand, the low-order background signal characteristics of the original input are effectively preserved through the residual terms.

[0036] like Figure 2 As shown, the feature sequence output by the residual temporal modeling module is fed into the cross-window dynamic feature interaction and refinement module. This module adopts a hierarchical window attention multi-layer stacked architecture. The high-order temporal feature sequence is processed through the hierarchical window attention multi-layer stacked architecture to obtain a deep temporal dynamic feature vector. Refined self-attention is calculated within a limited local window, and long-range feature interaction across windows is achieved through a window offset mechanism, thus achieving a balance between computational efficiency and global receptive field. Let the input feature sequence be... It is evenly divided in the time dimension. There are *w* non-overlapping local windows, each containing feature vectors from *w* consecutive time steps. For each local window, the feature subsequence... The module sequentially executes window multi-head self-attention and multilayer perceptron (MLP). The MLP sublayer consists of two fully connected layers and a GELU activation function. Its internal hidden layer dimension is usually expanded to 4 times the input dimension to perform nonlinear feature transformation and information integration in the channel dimension, thereby enhancing the model's expressive power.

[0037] To overcome the limitation of fragmented feature information between adjacent windows caused by a single fixed window partitioning, this module alternately employs a shifted window partitioning strategy between adjacent stacked layers. Specifically, standard non-overlapping window partitioning is used in even-numbered layers, while the starting position of the window is shifted along the time axis in odd-numbered layers. After each time step, the layers are re-divided, so that temporal features originally located at the boundaries of adjacent windows are included in the same window for attention computation in the next layer. Through this alternating shift mechanism, the module can gradually establish long-range temporal dependencies across multiple stacked layers without increasing the size of a single window, effectively improving the coverage of global feature interactions. After layer-by-layer refinement and cross-window interactions through L stacked sub-layers, the output features of each window are reassembled to restore the complete sequence, and the temporal dimension is compressed by global average pooling, finally outputting a deep temporal dynamic feature vector. ; The prior statistical feature extraction module alternately performs sequence mixing and channel mixing operations using an alternating hybrid architecture composed of a pre-built sequence mixing multilayer perceptron and a channel mixing multilayer perceptron. This integrates information from the original feature vector in the temporal position dimension and the channel semantic dimension, respectively, to obtain explicit prior statistical features. Specifically, it receives pre-computed original statistical feature vectors. As input (D represents the statistical feature dimension, encompassing explicit prior features such as time-domain statistical moments and frequency-domain energy ratios), the system employs an MLP-Mixer-like alternating mixing architecture. It alternately performs sequence mixing and channel mixing operations, integrating information from the original feature vectors in the temporal position dimension and channel semantic dimension, respectively, to obtain the feature vector. This is to uncover the intrinsic correlations between various statistical feature components. Input statistical feature vector First, the initial feature representation is mapped to a high-dimensional embedding space via a linear feature projection layer. ,in For the number of feature blocks, For the embedding dimension. This projection operation enables subsequent blending modules to integrate information in a unified high-dimensional semantic space. Feature representation This is then passed sequentially through multiple alternating mixing sublayers, each sublayer containing two steps: (1) The Sequence Mixing Multilayer Perceptron (Permute-MLP) is shown in Equation (107). This step transposes the feature matrix and feeds it into the MLP, so that the MLP's action direction expands along the sequence (position) dimension of the features, thereby capturing the cross-positional association patterns between different statistical feature components and realizing global information mixing in the temporal position dimension. After the operation is completed, the transpose restores the original dimensional arrangement and applies residual connections.

[0038] (107) (2) Channel-MLP, as shown in Equation (108); (108) This step independently applies MLP transformation to the feature vector at each location along the channel dimension, integrating information along the channel semantic dimension. This achieves non-linear refinement and recombination of semantic information from each channel, enhancing the expressive power of features at the semantic channel level. Residual connections are also configured to ensure training stability. After stepwise processing through two alternating mixing sublayers, global average pooling is performed on the output feature matrix to obtain explicit prior statistical feature vectors.

[0039] In obtaining deep temporal dynamic features Following explicit prior statistical features, this system incorporates a cross-modal interactive attention computation module. By introducing this mechanism, the global statistical perspective of the explicit prior statistical features guides and weights the extraction of the most relevant components to the current signal state from deep temporal features, achieving deep semantic alignment and complementary fusion of two types of heterogeneous features. Explicit prior statistical features are transformed into a query vector Q via a linear projection matrix, and then into key and value vectors K and V via the same linear projection matrix for cross-modal attention computation. This mechanism uses the global representation of the prior statistical features as the query condition, performing content-adaptive weighted retrieval in the temporal feature space. This allows the fused features to simultaneously possess the macroscopic perception capability of statistical priors and the fine-grained local expression of temporal dynamic features, effectively overcoming the limitations of single-modal representation.

[0040] EEG signals are susceptible to interference from electrooculography (EOG), electromyography (EMG), and motion artifacts during actual acquisition, leading to significant differences in signal quality at different time points. To ensure the final fused features are adaptively robust to signal quality, this invention employs a quality gating factor generation unit. This unit generates quality gating factors from the variance ratio (VR) and signal-to-noise ratio (SNR) of the signal quality metrics via a fully connected network and a sigmoid activation function. The unit uses the VR and SNR as input vectors to quantify signal quality assessment metrics. The continuous quality gating factor is generated by a two-layer fully connected network and a Sigmoid activation function as shown in equation (109); (109) in, For trainable parameters, This is the Sigmoid function. The physical meaning of the quality gate factor g is: when the signal quality is good (low VR, high SNR), The fusion features are fully activated; when the signal is severely contaminated by artifacts, The corresponding feature channels are adaptively suppressed, thereby reducing the interference of low-quality samples on the subsequent recognition network.

[0041] Finally, the cross-modal interaction attention computation fusion output is... The high-dimensional HEP hybrid feature representation is obtained by performing element-wise weighted multiplication with the quality gate factor in the gated fusion unit. It integrates three layers of information: underlying temporal dynamics, global statistical priors, and signal quality adaptive modulation. As the final output of the HEP high-dimensional temporal feature extraction module, it is sent to the downstream emotional state decoding system for further feature mapping and state recognition.

[0042] In some implementations, the emotional state decoding system, such as Figure 5As shown, a dual-branch asymmetric state-space coding structure is used to encode HEP mixture features, generating latent feature vectors. After mapping the latent feature vectors to the emotional state space, the emotional state probabilities are output. In detail, this involves obtaining a high-dimensional HEP mixture feature representation. Subsequently, the emotional state decoding system maps it to a human-understandable space of intrinsic sensory states (i.e., emotional states), achieving end-to-end output of emotional state probability vectors. This system sequentially includes a hierarchical feature encoding module, a semantic similarity decoding module, and so on.

[0043] The hierarchical feature encoding module has a dual-branch asymmetric state-space encoding structure. For one branch, the HEP hybrid features are subjected to sequence flipping, local transformation, and discretization preprocessing before being fed into the state-space model of that branch for sequence state evolution modeling output. For the other branch, the HEP hybrid features are subjected to local transformation and discretization preprocessing before being fed into the state-space model of that branch for sequence state evolution modeling output. The outputs of the two branches are respectively subjected to cross-multiplication operation, then superimposed and merged by channel, and finally reduced in dimension by the top linear projection layer to output the latent feature vector. First, the module proceeds to the hierarchical feature encoding module to generate continuous latent feature vectors. This module is designed as a dual-branch asymmetric state-space encoding structure. Both parallel branches use 1DConv and operators to perform local transformations and discretization preprocessing on the mixed input features before feeding them into their respective State Space Models (SSMs) for sequence state evolution modeling. The difference between the two branches is that the left branch performs an additional sequence flipping operation before feeding the SSM, causing its SSM to scan the feature sequence in reverse along the time axis; the right branch maintains the original temporal direction while scanning forward. This asymmetric design allows the two branches to capture the dynamic evolution of potential states from the causal and anti-causal directions of the temporal sequence, respectively. The two branches complement each other, effectively overcoming the inherent limitations of unidirectional SSMs in capturing bidirectional temporal dependencies. The emotional state decoding system will... Mapped to an emotional state probability vector This probability vector is also the final output, which is transmitted to the neural regulation system.

[0044] The SSM outputs of the two branches are then gatedly fused with their corresponding skip connection features via cross-multiplication to preserve the underlying feature details of each branch before deep transformation. After merging the features of all branches along the channel dimension at each channel stacking unit, the resulting vectors are reduced in dimensionality by a top linear projection layer, outputting a continuous latent feature vector. Continuous latent feature vectors Expanded into matrix form by the feature projection module This forms a latent feature matrix, where N is the predefined number of emotional state categories. Simultaneously, the system pre-constructs an emotional state codebook containing N discrete states, using natural language text labels (such as arousal, anxiety, calmness, depression, apathy, focus, etc.) to semantically describe each emotional state category. These text labels are input into a pre-trained BERT text encoder, which extracts the contextual semantic representation of each label and aligns it to the desired semantic representation using linear projection. The same embedding dimension constitutes the semantic feature matrix of emotional state. Compared to traditional one-hot labels or randomly initialized category embeddings, the pre-trained semantic representations provided by BERT semantic encoders can naturally encode the semantic proximity relationships between emotional states (e.g., "anxiety" and "hysteria" are closer in semantic space than "anxiety" and "calm"), thus providing structured semantic supervision signals for cross-modal alignment and improving the model's discriminative ability in scenarios with blurred emotional category boundaries.

[0045] Next, the semantic similarity decoding module calculates the similarity between the latent feature matrix and the pre-constructed emotional state semantic feature matrix, obtaining a similarity score vector; based on the similarity score vector, it outputs the emotional state probability vector. The cosine similarity of the latent feature matrix obtained through semantic similarity decoding is then used to analyze the... With emotional semantic feature matrix Cross-modal similarity is calculated row by row to obtain the similarity score vector, and the output emotional state probability vector distribution is output after Softmax normalization.

[0046] In some implementations, the closed-loop neuromodulation system (hereinafter referred to as the neuromodulation system) based on interoceptive state feedback is implemented as follows: Figure 6 As shown, this is used to generate pulse sequence parameters for stimulating nerves based on the emotional state category corresponding to the probability of the emotional state, under preset stimulation conditions.

[0047] This system uses transcutaneous vagus nerve stimulation (taVNS) as the actuation method, constructing a two-layer control structure of "outer loop macroscopic strategy setting + inner loop microscopic physiological error closed loop," and introducing anti-habituation perturbation and safety limiting mechanisms. It outputs pulse sequence parameters that can directly drive the stimulation hardware, aiming to solve the technical bottlenecks of existing taVNS technology, such as the lack of individualized physiological indicator feedback and the susceptibility to neural adaptation. The pulse sequence parameters include current intensity. Pulse width Reference frequency Pulse interval .

[0048] The system receives two upstream inputs: one is the emotional state probability vector from the emotional state decoding system of System 2. ( For the number of emotional categories, (or its clustered dominant state cluster) The second part consists of microscopic features from the HEP signal extraction and feature encoding system of System 1, including stable HEP average amplitude features. or original feature vector and signal quality indicators (Used for regulating credibility assessment). The stimulus parameter vector output to the hardware includes frequencies. Pulse width Current emphasis With pulse interval .

[0049] In the dual-feature trigger determination stage, to avoid overstimulation of the normal state, this system adopts a dual-condition triggering mechanism of "macro-negative index + micro-internal perception deviation". The negative emotional index is determined based on the emotional state category corresponding to the emotional state probability vector; a normalized deviation is constructed based on the average amplitude characteristics of the HEP within the current analysis window in the original feature vector and the pre-determined individual baseline amplitude characteristics of the HEP; when the negative emotional index, deviation, and signal quality index are all greater than the corresponding preset threshold parameters, the preset stimulation condition is determined to be met. Specifically, let the set of negative emotional categories be... (Such as anxiety, depression, apathy, etc.), define the Negative Affective Index (NEI): (110) HEP individual baseline amplitude Compared with the average amplitude of HEP in the current analysis window Constructing the normalized deviation: (111) If and only if the following conditions are met: (112) And the signal quality indicators meet (or When this occurs, the taVNS closed-loop intervention is triggered. Among other things... The continuous quality weights output by module one (see formula 011). , , , All parameters are configurable threshold parameters. This mechanism uses both macro-emotional risk and micro-internal perceptual abnormalities as intervention criteria to enhance the specificity of regulation.

[0050] In the outer loop macro-strategy setting phase, a baseline stimulus parameter set corresponding to the emotional state category is mapped based on a pre-set knowledge base. This baseline stimulus parameter set includes baseline frequency, baseline pulse width, and baseline current intensity. Specifically, the outer loop outputs the baseline stimulus parameter set based on the dominant emotional category. Dominant category: (113) Parameter mapping is accomplished using an internally pre-built clinical expert knowledge base lookup table (LUT): (114) in , , They are respectively related to the dominant emotional category The corresponding reference frequency, reference pulse width, and reference current intensity. The outer loop ensures that the stimulus paradigm matches the macroscopic psychopathological characteristics and provides an interpretable initial operating point for the inner loop.

[0051] In the inner loop incremental adaptive PID fine-tuning stage based on HEP ​​error, the adaptive PID adjustment parameters are determined according to the original eigenvector; specifically, the target HEP amplitude is set to... (The setpoint can be taken as the individual baseline or the desired adjustment direction), then the HEP error is: (115) During system initialization, the baseline stimulus parameter set is adjusted based on the adaptive PID adjustment parameters. (The last part, "can be taken," appears to be incomplete and requires further context.) = This means that the initial adjustment target is the individual's baseline; in long-term intervention, adjustments can be made based on clinical goals. Make slow up / down adjustments. Use discrete incremental PID to calculate the current fine-tuning amount: (116) in , , These are the proportional, integral, and derivative coefficients, respectively. The current output is then updated. (117) in For hard limiting operators, and This refers to the lower and upper limits of safety. Within a macro strategy cycle, the inner loop output is considered to revolve around the outer loop benchmark. The perturbation yields the final strength: (118) Similarly, it can also be used for , A small-amplitude adaptive fine-tuning is introduced, but it must meet the charge density and hardware constraints.

[0052] In the anti-habituation perturbation and safe output processing stage, long-term fixed parameter stimulation can easily induce neural adaptation / desensitization. This system introduces perturbations into the pulse timing while ensuring safety. At the reference frequency, a zero-mean random perturbation is introduced to determine the pulse interval. Specifically, at the reference frequency... Next, introduce a zero-mean random perturbation: (119) in This represents the standard deviation of the disturbance, used to control the amplitude of random jitter. And for... Range constraints should be applied to avoid overly dense or overly sparse stimuli: (120) in and These are the lower and upper bounds for the safe pulse interval. To ensure stimulation safety, in addition to current limiting, the amount of single-phase charge also needs to be constrained. : (121) in This is the upper limit of safety for single-phase charge. When When the limit is exceeded, prioritize reducing. or To meet the constraints. Furthermore, as part of the safety hard limiting and protection strategy, when the quality indicators given by Module 1 indicate that the current heart rate estimate is unreliable (e.g. Low or If an anomaly occurs, the system will freeze or slowly decay the inner loop update to avoid mis-tuning of the artifact-driven mechanism. (122) in The mass decay factor can be derived from... , The mapping is obtained.

[0053] Finally, output parameters , , , The signals are converted into hardware-executable PWM / digital-to-analog control instructions to generate a sequence of biphase or symmetrical charge stimulation pulses. The system synchronously records the probability of the input emotional state at each stimulation moment. Microscopic physiological feedback , , and output stimulus parameters , , , Used for subsequent individual baselines With the goal Slow adaptive updates and long-term efficacy assessment.

[0054] In some implementations, the results of the macro-emotional state decoding system are used for outer-loop strategy setting and trigger determination; the neuromodulation system provides micro-internal sensory physiological characteristics for inner-loop error feedback and safety and reliability assessment; together, they constitute a closed-loop control link of decision-making-execution-feedback in closed-loop neuromodulation, realizing precise, safe, and anti-habituation online neurointervention.

[0055] Briefly describe the principles of the embodiments of the present invention.

[0056] In this embodiment of the invention, an ear canal flexible dry electrode array can be used to simultaneously acquire EEG and ECG signals; based on ECG R-wave detection, cardiac events are identified in real time and trigger markers are generated; and based on cardiac events, EEG signals are time-locked and superimposed to extract cardiac evoked potential waveforms and their original characteristic parameters; based on the original characteristics such as HEP amplitude, latency, and topological distribution, an interoceptive functional state vector is constructed; thereby, abnormal emotional patterns can be identified based on the interoceptive state vector, triggering or regulating the output of percutaneous vagus nerve stimulation.

[0057] In practice, the position of the R wave in the ECG signal is detected in real time to generate a cardiac event time series. Using the R wave as a trigger point, EEG signal segments of 200-500 ms before and after the R wave are extracted. Baseline correction and superposition averaging are performed on multiple consecutive EEG segments to extract a stable HEP waveform. Key HEP features are then calculated, including the average amplitude, peak latency, and left-right hemisphere asymmetry index within a specific time window in the frontal / central region (e.g., 200-400 ms after the R wave). The average amplitude of HEP reflects the intensity of interoceptive processing; the HEP latency reflects the efficiency of interoceptive conduction; the coupling index between HEP ​​and HRV reflects the coordination of heart-brain interaction; and the difference between resting-state HEP and stimulus-induced HEP reflects neural plasticity. Based on these features, clustering or classification models are used to identify "normal / abnormal" interoceptive processing states and quantify the degree of abnormality.

[0058] For closed-loop modulation based on interoceptive feedback, an interoceptive state threshold is set: when the HEP amplitude decreases by more than 30% of the baseline or the latency prolongs by more than 50ms, it is determined to be an "interoceptive desensitization" state; and a trigger condition is set: when the dual characteristics of "interoceptive desensitization + HRV reduction" are detected, taVNS stimulation is automatically initiated; for feedback regulation: the HEP amplitude change is monitored in real time during stimulation, and when the HEP amplitude returns to the normal range or reaches the plateau, the stimulation intensity is dynamically reduced or stimulation is terminated; and long-term learning effect recording is possible: the HEP baseline change before and after multiple stimulations is recorded to assess the process of interoceptive function remodeling.

[0059] This invention, for the first time, uses cardiac evoked potentials as the core control signal and state feedback variable of brain-computer interface, realizing the quantification and regulation of the core dimension of "the brain's ability to perceive heartbeats", and providing a new neurophysiological target for the intervention of mood disorders.

[0060] Breaking through the limitation of HEP relying on high-density EEG in the laboratory, stable HEP extraction was achieved for the first time on an ear canal portable device. Through ECG-triggered time-locked superposition and motion artifact suppression algorithms, continuous monitoring of endoesophageal function was achieved in natural scenes.

[0061] Establish a closed-loop control process of "HEP feature recognition → interoceptive state determination → taVNS parameter adjustment → HEP feedback verification" to achieve closed-loop verification from "stimulus output" to "interoceptive function improvement", so that neural regulation can truly act on the neural basis of subjective emotional experience.

[0062] A portable brain-computer interface based on endosensory neural coding for cardiac evoked potentials has been realized, expanding the control dimensions and application scenarios of BCI; To provide quantifiable and interventionist neurophysiological targets for mood disorders such as depression and anxiety; Achieving a closed-loop regulation of interoceptive responses through "stimulus-perception-feedback" improves the accuracy and interpretability of interventions; Provide a wearable tool platform for research on the mechanism of "mind-brain interaction".

[0063] Example 2 This invention provides a method for evoked neuromodulation of heartbeats based on interoceptive neural coding. The method for evoked neuromodulation of heartbeats includes: The received spatiotemporally synchronized EEG and cardiovascular signals are processed for signal extraction and feature encoding, and output stable heartbeat evoked potential (HEP) waveform sequences, original interoceptive feature vectors, and signal quality indices. Based on the stable HEP waveform sequence, the original feature vector of the intrinsic perception, and the signal quality index, output the HEP hybrid features; A dual-branch asymmetric state-space coding structure is used to encode the HEP hybrid features to generate latent feature vectors. After mapping the latent feature vectors to a pre-constructed emotional state semantic feature matrix, the emotional state probability vectors are output. Under preset stimulation conditions, pulse sequence parameters for stimulating nerves are generated based on the emotional state category corresponding to the emotional state probability vector.

[0064] The process involves extracting and encoding features from the received spatiotemporally synchronized EEG and cardiovascular signals, outputting a stable heartbeat evoked potential (HEP) waveform sequence, original interoceptive feature vector, and signal quality indices, including: It receives EEG and cardiovascular signals that are spatiotemporally synchronized and aligned at the hardware acquisition end; The EEG signals were preprocessed using high-pass filtering, notch filtering, and low-pass filtering; the cardiovascular signals were preprocessed using band-pass filtering with upper and lower cutoff frequencies. The preprocessed EEG signals were subjected to adaptive cancellation of ECG propagation artifacts and suppression of electrooculography artifacts to reconstruct purified EEG signals; and the time sequence of cardiac events in the preprocessed cardiovascular signals was determined. Based on the time sequence of cardiac events, the purified EEG signal is adaptively locked and segmented to extract cardiac evoked potential waveform segments; heart rate-dependent adaptive baseline correction is applied to multiple consecutive cardiac evoked potential waveform segments to obtain baseline-corrected cardiac evoked potential waveform segments; R-wave confidence gating is applied to the baseline-corrected cardiac evoked potential waveform segments based on the quality weights generated by the pre-constructed QC quality index. The baseline-corrected cardiac evoked potential waveforms gated by R-wave confidence are segmented, superimposed and smoothed in a weighted sliding window to obtain the stable HEP waveform sequence at the current moment. Output the stable HEP waveform sequence at the current moment, the original intrinsic sensing feature vector determined based on the stable HEP waveform sequence, and the signal quality index.

[0065] In some implementations, pre-constructing QC quality indicators includes: Peak constraint, baseline stability threshold, missing / saturation ratio integrity threshold, and R-wave confidence.

[0066] In some implementations, the HEP hybrid features are output based on the stable HEP waveform sequence, the intrinsic intuition original feature vector, and the signal quality index, including: Temporal dynamic feature extraction: A multi-scale one-dimensional convolutional network is used to extract features from a stable HEP waveform in parallel multi-branch mode, and the extracted multi-branch features are concatenated to obtain scale-concatenated features; Residual temporal modeling: The scale-stitched features are fused with multi-level temporal dependency modeling, external memory network and multi-scale features to output a high-order temporal feature sequence; the high-order temporal feature sequence is processed by a multi-layer stacked architecture of hierarchical window attention to obtain a deep temporal dynamic feature vector; Prior statistical feature extraction: The alternating hybrid architecture composed of a pre-set sequence-mixing multilayer perceptron and a channel-mixing multilayer perceptron is used to perform two types of operations, sequence mixing and channel mixing, respectively, to integrate the information of the original feature vector in the temporal position dimension and the channel semantic dimension, so as to obtain explicit prior statistical features. Cross-modal interaction attention computation: Cross-modal attention computation is performed between the key vector and value vector obtained by linear projection of deep temporal dynamic feature vectors and the query vector obtained by linear projection of explicit prior statistical features; Quality gating factor generation: The variance ratio and signal-to-noise ratio of the signal quality indicators are used to generate quality gating factors through a fully connected network and a Sigmoid activation function; Quality-gated fusion unit: The feature sequence obtained by cross-modal attention calculation is weighted element-wise with the quality gating factor to generate HEP hybrid features.

[0067] In some implementations, the step of fusing scale-stitched features with multi-level temporal dependency modeling, external memory networks, and multi-scale features to output a high-order temporal feature sequence includes: The scale-concatenated feature sequence is fed into a bidirectional long short-term memory network LSTM. The feature sequence output by the bidirectional LSTM is fed into a multi-scale feature fusion layer and then sequentially through a dynamic feature correlation estimation layer and a dynamic feature convolutional layer. The dynamic feature correlation estimation layer adaptively estimates the correlation of the current feature sequence in the time and channel dimensions using a learnable parameter matrix, outputting a correlation weight matrix. The dynamic feature convolutional layer uses the correlation weight matrix as the modulation coefficients of the adaptive convolutional kernel to perform dynamic convolution operations on the feature sequence output by the bidirectional LSTM. The feature sequence output by the dynamic feature convolutional layer is then input into the multi-scale feature fusion layer and an external memory network. The feature sequences output by the bidirectional LSTM, the feature sequences output by the dynamic feature convolutional layer, and the feature sequences output by the external memory network are processed by a multi-scale feature fusion layer to output a high-order temporal feature sequence.

[0068] In some implementations, the external memory network includes: The spectral clustering pattern extraction module is used to extract clustering patterns from the feature sequences output by the dynamic feature convolutional layer and determine typical temporal dynamic pattern prototypes. The memory matrix update module is used to update the preset memory matrix according to the prototype of a typical time-series dynamic pattern. The attention-based aggregation module uses the global mean pooling result of the feature sequence output by the dynamic feature convolutional layer as a query, and extracts the memory feature sequence that matches the current heartbeat state from the memory matrix through an attention-weighted retrieval mechanism.

[0069] In some implementations, the step of encoding the HEP hybrid features using a dual-branch asymmetric state-space coding structure to generate a latent feature vector, and then mapping the latent feature vector to a pre-constructed emotional state semantic feature matrix to output an emotional state probability vector includes: Hierarchical Feature Encoding: A dual-branch asymmetric state-space encoding structure is set up. For one branch, the HEP hybrid features are subjected to sequence flipping, local transformation, and discretization preprocessing, and then fed into the state-space model of that branch for sequence state evolution modeling output. For the other branch, the HEP hybrid features are subjected to local transformation and discretization preprocessing, and then fed into the state-space model of that branch for sequence state evolution modeling output. The outputs of the two branches are respectively subjected to cross-multiplication operation, then superimposed and merged by channel, and finally reduced in dimension by the top linear projection layer to output the latent feature vector. Feature projection: expanding the latent feature vectors into a latent feature matrix; Semantic similarity decoding: The similarity between the latent feature matrix and the pre-constructed semantic feature matrix of emotional state is calculated to obtain a similarity score vector; the emotional state probability vector is output based on the similarity score vector.

[0070] Optionally, the following method can be used to determine whether the preset stimulus conditions are met: The negative emotional index is determined based on the emotional state category corresponding to the emotional state probability vector; A normalized deviation is constructed based on the average amplitude characteristics of HEP within the current analysis window in the original feature vector and the pre-determined baseline amplitude characteristics of individual HEPs. When the negative emotional index, deviation amount, and signal quality index are all greater than the corresponding preset threshold parameters, the preset stimulus conditions are determined to be met.

[0071] Optionally, under preset stimulation conditions, based on the emotional state category corresponding to the emotional state probability vector, the pulse sequence parameters for stimulating nerves include: Based on a pre-set knowledge base, the reference stimulus parameter set corresponding to the emotional state category is mapped; the reference stimulus parameter set includes reference frequency, reference pulse width and reference current intensity. Determine the adaptive PID adjustment parameters based on the original feature vector; The baseline stimulus parameter set is adjusted based on the adaptive PID adjustment parameters; At the reference frequency, a zero-mean random perturbation is introduced to determine the pulse interval in order to resist habituation perturbation; Range constraints are applied to pulse intervals and single-phase charges to ensure safe output limiting. The pulse sequence parameters for stimulating nerves are generated based on adjustments, anti-habituation perturbations, and safe output limiting.

[0072] The specific implementation of Embodiment 2 can be found in Embodiment 1, and it has the corresponding technical effects.

[0073] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0074] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0076] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A brain-computer interface system based on endoreceptive neural coding of cardiac evoked potentials, characterized in that, The cardiac evoked potential brain-computer interface system includes: The HEP signal extraction and feature encoding system is used to extract and encode features from received spatiotemporally synchronized EEG and cardiovascular signals, and output stable cardiac evoked potential HEP waveform sequences, original interoceptive feature vectors, and signal quality indicators. The HEP high-dimensional temporal feature extraction system is used to output HEP hybrid features based on stable HEP waveform sequences, intrinsic intuition original feature vectors, and signal quality indicators. The emotional state decoding system is used to encode the hybrid features of HEP using a dual-branch asymmetric state space coding structure, generate latent feature vectors, and then map the latent feature vectors to a pre-constructed emotional state semantic feature matrix to output an emotional state probability vector. The neural modulation system is used to generate pulse sequence parameters for stimulating nerves based on the emotional state category corresponding to the emotional state probability vector, under preset stimulation conditions.

2. The brain-computer interface system for cardiac evoked potentials according to claim 1, characterized in that, The HEP signal extraction and feature encoding system includes: The input layer is used to receive EEG and cardiovascular signals that are spatiotemporally synchronized and aligned at the hardware acquisition end; The preprocessing layer is used to preprocess EEG signals by high-pass filtering, notch filtering, and low-pass filtering; and to preprocess cardiovascular signals by band-pass filtering with upper and lower cutoff frequencies. A dual-path processing layer is used to adaptively cancel ECG propagation artifacts and suppress electrooculography artifacts on the preprocessed EEG signals, reconstructing clean EEG signals; and to determine the time sequence of cardiac events in the preprocessed cardiovascular signals. The quality gating layer is used to adaptively lock and segment the purified EEG signal according to the time sequence of cardiac events, and extract cardiac evoked potential waveform segments; it applies heart rate-dependent adaptive baseline correction to multiple consecutive cardiac evoked potential waveform segments to obtain baseline-corrected cardiac evoked potential waveform segments; and it performs R-wave confidence gating on the baseline-corrected cardiac evoked potential waveform segments according to the quality weights generated by the pre-constructed QC quality index. The aggregation layer is used to segment and smooth the baseline-corrected cardiac evoked potential waveforms gated by R-wave confidence within a weighted sliding window to obtain a stable HEP waveform sequence at the current moment. The output layer is used to output the stable HEP waveform sequence at the current moment, the original intrinsic sensing feature vector determined based on the stable HEP waveform sequence, and the signal quality index.

3. The brain-computer interface system for cardiac evoked potentials according to claim 2, characterized in that, Pre-constructed QC quality indicators include: Peak constraint, baseline stability threshold, missing / saturation ratio integrity threshold, and R-wave confidence.

4. The cardiac evoked potential brain-computer interface system according to claim 1, characterized in that, The HEP high-dimensional temporal feature extraction system includes: The temporal dynamic feature extraction module is used to perform parallel multi-branch feature extraction on stable HEP waveforms using a multi-scale one-dimensional convolutional network, and then concatenates the extracted multi-branch features to obtain scale-concatenated features. The residual temporal modeling module is used to fuse scale-stitched features with multi-level temporal dependency modeling, external memory network and multi-scale features to output a high-order temporal feature sequence; the high-order temporal feature sequence is processed by a multi-layer stacked architecture of hierarchical window attention to obtain a deep temporal dynamic feature vector. The prior statistical feature extraction module is used to alternately perform two types of operations, sequence mixing and channel mixing, through an alternating mixing architecture composed of a pre-set sequence mixing multilayer perceptron and a channel mixing multilayer perceptron. This is used to integrate the information of the original feature vector in the temporal position dimension and the channel semantic dimension, respectively, to obtain explicit prior statistical features. The cross-modal interaction attention calculation module is used to perform cross-modal attention calculation on the key vector and value vector obtained by linear projection of deep temporal dynamic feature vectors and the query vector obtained by linear projection of explicit prior statistical features. The quality gating factor generation unit is used to generate quality gating factors from the variance ratio and signal-to-noise ratio of signal quality indicators through a fully connected network and a sigmoid activation function; The quality-gated fusion unit is used to weight the feature sequence obtained by cross-modal attention calculation with the quality gating factor element-wise to generate HEP hybrid features.

5. The cardiac evoked potential brain-computer interface system according to claim 4, characterized in that, The process of fusing scale-stitched features through multi-level temporal dependency modeling, external memory networks, and multi-scale features to output a high-order temporal feature sequence includes: The scale-concatenated feature sequence is fed into a bidirectional long short-term memory network LSTM. The feature sequence output by the bidirectional LSTM is fed into a multi-scale feature fusion layer and then sequentially through a dynamic feature correlation estimation layer and a dynamic feature convolutional layer. The dynamic feature correlation estimation layer adaptively estimates the correlation of the current feature sequence in the time and channel dimensions using a learnable parameter matrix, outputting a correlation weight matrix. The dynamic feature convolutional layer uses the correlation weight matrix as the modulation coefficients of the adaptive convolutional kernel to perform dynamic convolution operations on the feature sequence output by the bidirectional LSTM. The feature sequence output by the dynamic feature convolutional layer is then input into the multi-scale feature fusion layer and an external memory network. The feature sequences output by the bidirectional LSTM, the feature sequences output by the dynamic feature convolutional layer, and the feature sequences output by the external memory network are processed by a multi-scale feature fusion layer to output a high-order temporal feature sequence.

6. The brain-computer interface system for cardiac evoked potentials according to claim 5, characterized in that, The external memory network includes: The spectral clustering pattern extraction module is used to extract clustering patterns from the feature sequences output by the dynamic feature convolutional layer and determine typical temporal dynamic pattern prototypes. The memory matrix update module is used to update the preset memory matrix according to the prototype of a typical time-series dynamic pattern. The attention-based aggregation module uses the global mean pooling result of the feature sequence output by the dynamic feature convolutional layer as a query, and extracts the memory feature sequence that matches the current heartbeat state from the memory matrix through an attention-weighted retrieval mechanism.

7. The brain-computer interface system for cardiac evoked potentials according to claim 1, characterized in that, The emotional state decoding system includes: The hierarchical feature encoding module has a dual-branch asymmetric state-space encoding structure. For one branch, the HEP hybrid features are subjected to sequence flipping, local transformation, and discretization preprocessing before being fed into the state-space model of that branch for sequence state evolution modeling output. For the other branch, the HEP hybrid features are subjected to local transformation and discretization preprocessing before being fed into the state-space model of that branch for sequence state evolution modeling output. The outputs of the two branches are respectively subjected to cross-multiplication operations, then superimposed and merged by channel, and finally reduced in dimension by the top linear projection layer to output the latent feature vector. The feature projection module is used to expand the latent feature vectors into a latent feature matrix; The semantic similarity decoding module is used to calculate the similarity between the latent feature matrix and the pre-constructed emotional state semantic feature matrix to obtain a similarity score vector; and outputs an emotional state probability vector based on the similarity score vector.

8. The brain-computer interface system for cardiac evoked potentials according to claim 1, characterized in that, The following method is used to determine whether the preset stimulus conditions are met: The negative emotional index is determined based on the emotional state category corresponding to the emotional state probability vector; A normalized deviation is constructed based on the average amplitude characteristics of HEP within the current analysis window in the original feature vector and the pre-determined baseline amplitude characteristics of individual HEPs. When the negative emotional index, deviation amount, and signal quality index are all greater than the corresponding preset threshold parameters, the preset stimulus conditions are determined to be met.

9. The cardiac evoked potential brain-computer interface system according to any one of claims 1-8, characterized in that, The step of generating pulse sequence parameters for stimulating nerves based on the emotional state category corresponding to the emotional state probability vector includes: Based on a pre-set knowledge base, the reference stimulus parameter set corresponding to the emotional state category is mapped; the reference stimulus parameter set includes reference frequency, reference pulse width and reference current intensity. Determine the adaptive PID adjustment parameters based on the original feature vector; The baseline stimulus parameter set is adjusted based on the adaptive PID adjustment parameters; At the reference frequency, a zero-mean random perturbation is introduced to determine the pulse interval in order to resist habituation perturbation; Range constraints are applied to pulse intervals and single-phase charges to ensure safe output limiting. The pulse sequence parameters for stimulating nerves are generated based on adjustments, anti-habituation perturbations, and safe output limiting.

10. A method for cardiac evoked neural modulation based on interoceptive neural coding, characterized in that, The method for inducing neural modulation of cardiac contraction: The received spatiotemporally synchronized EEG and cardiovascular signals are processed for signal extraction and feature encoding, and output stable heartbeat evoked potential (HEP) waveform sequences, original interoceptive feature vectors, and signal quality indices. Based on the stable HEP waveform sequence, the original feature vector of the intrinsic perception, and the signal quality index, output the HEP hybrid features; A dual-branch asymmetric state-space coding structure is used to encode the HEP hybrid features to generate latent feature vectors. After mapping the latent feature vectors to a pre-constructed emotional state semantic feature matrix, the emotional state probability vectors are output. Under preset stimulation conditions, pulse sequence parameters for stimulating nerves are generated based on the emotional state category corresponding to the emotional state probability vector.