Brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence
By synchronously sampling electrode arrays, near-infrared optical probes, and inertial measurement elements under the same master clock, and combining multi-level artifact suppression and temporal similarity discrimination, cross-modal neural representation sequences are reconstructed. This solves the problems of insufficient signal quality and time delay in existing technologies, and achieves neural intent decoding with high signal-to-noise ratio and temporal continuity, supporting natural and fluent digital human interaction.
Patent Information
- Application Number
- CN202511358144.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-23
AI Technical Summary
Existing brain-computer interface systems have shortcomings in signal quality, spatial resolution, and temporal latency. Multimodal fusion systems lack collaborative optimization strategies for temporal synchronization, energy consistency, and reconstruction of missing regions, resulting in inaccurate neural intent decoding and unstable action mapping.
By using a frequency-synchronized sampling electrode array, near-infrared optical probe, and miniature inertial measurement element under the same master clock, combined with multi-level artifact suppression and temporal similarity discrimination, a cross-modal neural representation sequence is reconstructed. A multi-layer temporal coding network is then used for real-time intent decoding and action mapping to drive the digital human skeleton and facial expressions.
It improves the temporal synchronization and energy consistency of cross-modal signals, eliminates motion artifacts and interpolation faults, enhances the accuracy and robustness of neural intent decoding, supports personalized calibration and real-time feedback, and realizes natural and smooth digital human interaction.
Smart Images

Figure CN120848740B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of brain-computer interface, and particularly relates to a brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence. BACKGROUND
[0002] With the rapid development of brain-computer interface technology and artificial intelligence algorithm, more and more researches attempt to realize the control of virtual characters or digital humans through real-time decoding of electroencephalogram signals or near-infrared optical signals. Traditional brain-computer interface systems are mostly based on single modal signals, for example, electroencephalogram signals collected by a scalp electrode array are mapped to external control instructions through specific frequency band power or phase characteristics; some schemes also use near-infrared optical probes to detect local oxygen concentration changes, and indirectly reflect the intensity of neural activity through hemodynamic response. However, single modal has its own shortcomings in signal quality, spatial resolution and time delay characteristics: electroencephalogram signals are easily disturbed by electromyogram and power frequency interference, and near-infrared optical signals have sampling lag and are sensitive to head movement. In addition, schemes that only use inertial measurement units to compensate for motion artifacts can only achieve a certain degree of smoothing in the low frequency band, and lack targeted suppression of transient artifacts such as frequency electromyogram and electrode sliding. Existing multi-modal fusion systems mainly focus on simple data splicing or weighted fusion, and lack collaborative optimization strategies for key links such as time domain synchronization, energy consistency and missing area reconstruction of the three types of modal signals. SUMMARY
[0003] In view of this, the main purpose of the present application is to provide a brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence, which synchronously samples the electrode array, the near-infrared optical probe and the micro inertial measurement element under the same master clock to form a cross-modal data stream of neural potential fluctuations, local oxygen concentration changes and head displacement information, and sequentially applies multi-stage artifact suppression, gestalt compensation based on time domain similarity discrimination, energy consistency re-calibration and gradual weighted smoothing to reconstruct a high signal-to-noise ratio and continuous cross-modal neural representation sequence; on this basis, a multi-layer time sequence coding network and an adaptive feedback stabilization strategy are used to realize real-time intention decoding and action mapping of the neural representation sequence, to drive the digital human skeleton and expression to realize natural and smooth interaction with low latency; the present application significantly improves the time domain synchronization and energy consistency of cross-modal signals, eliminates motion artifacts and interpolation gaps, enhances the accuracy and robustness of neural intention decoding, supports personalized calibration and real-time feedback, and can stably output coherent and natural digital human actions in complex interaction scenarios.
[0004] The technical scheme adopted by the present application is as follows:
[0005] A brain-controlled digital human interaction system based on a brain-computer interface and artificial intelligence, the system comprising: a multi-modal neural signal acquisition unit, a cross-modal noise suppression and complete form compensation unit, and a digital human interaction control unit; the multi-modal neural signal acquisition unit comprises an electrode array, a near-infrared optical probe, and a miniature inertial measurement element arranged on the scalp surface and behind the ear of a user at a predetermined interval, and uses the same master clock trigger frequency sampling to obtain three types of modal data, including: neural potential fluctuation, local blood oxygen concentration change, and head displacement information, and the three types of modal data are injected into an on-chip ring-shaped double buffer in real time, and a unified timestamp is attached when writing in the on-chip ring-shaped double buffer, cross-modal time alignment is achieved, and a cross-modal data stream is obtained; the cross-modal noise suppression and complete form compensation unit is used for multi-stage artifact suppression on the cross-modal data stream, and then complete form compensation based on time domain similarity discrimination is performed to obtain a cross-modal neural representation sequence; and the digital human interaction control unit is used for realizing digital human control according to the cross-modal neural representation sequence.
[0006] Further, the process of multi-stage artifact suppression on the cross-modal data stream by the cross-modal noise suppression and complete form compensation unit comprises: calculating the median value and the range in the first sliding time window for the three types of modal data respectively; if the sum of the ranges of the three types of modal data in the first sliding time window is less than a set steady-state threshold, the maximum value of the ranges of the current three types of modal data is recorded as a zero-level steady-state template; the cross-modal data stream is maintained in three layers of increasing length sliding time windows respectively according to the channels, and the median absolute deviation of each layer is updated in real time; the current layer median absolute deviation is multiplied by a sub-window coefficient, and then the range proportion of the zero-level steady-state template is superimposed to obtain a multi-layer fusion threshold; if the absolute amplitude of the value of any type of modal data exceeds the fusion threshold of the corresponding layer and the duration is shorter than the minimum duration of a neural event, the timestamp range of the three types of modal data is uniformly marked as a first-level artifact section; fixed-length normal waveforms are extracted before and after the first-level artifact section, respectively, and cubic spline interpolation is performed on the neural potential fluctuation and the local blood oxygen concentration change three times, and bilinear interpolation is performed on the head displacement information, and then the first-level cleaning matrix is spliced.
[0007] Further, the range proportion of the zero-level steady-state template is defined as the ratio of the zero-level steady-state template to the sum of the ranges of the three types of modal data in the first sliding time window.
[0008] Furthermore, the cross-modal noise suppression and gestalt compensation unit performs gestalt compensation based on temporal similarity discrimination on the cross-modal data stream. This process includes: detecting and indexing abnormal sampling gaps or residual regions in the three types of modal data to form a gap list; retrieving candidate segments that meet the bidirectional correlation condition in historical time periods based on the gap context waveform and constructing a candidate set; performing multi-scale dynamic temporal alignment on the candidate segments and filtering alignment candidates whose matching errors meet the threshold requirements; performing similarity-weighted fusion on the alignment candidates based on the alignment error and eliminating local abnormal spikes through filtering to obtain interpolated segments; performing amplitude recalibration based on energy vectors on the interpolated segments to achieve cross-modal energy consistency; and performing gradual weighted smoothing on both ends of the interpolated segments and concatenating them to the first-level cleaning matrix to form a second-level cleaning matrix.
[0009] Furthermore, the method for detecting and indexing abnormal sampling segments or residuals in the three types of modal data to form a gap list includes: traversing the unified time axis of the primary cleaning matrix to detect sampling segments or segments in any type of modal data where the absolute value of the interpolation residual exceeds twice the local median absolute deviation; and recording the starting index for the detected segments. Terminate index Root mean square value of residuals of nerve potential fluctuations Root mean square value of residuals of local blood oxygen concentration changes The root mean square value of the residuals with head displacement information ; Generate gap tuples The gaps are then written into the gap list; if the interval between two consecutive gaps is less than the minimum period of a neural event, they are merged into a single gap entry to prevent fragmented reconstruction from causing boundary tearing.
[0010] Furthermore, the process of retrieving candidate segments that meet the bidirectional correlation condition in historical time periods based on the gap context waveform and constructing a candidate set includes: for each gap, considering the segments before and after the gap... The context waveform in milliseconds is used as the query sample; the closest one before the gap timestamp. Within the historical period of seconds, according to A sliding window with a step size is used to extract candidate fragments, and positive Pearson correlation coefficients are calculated for each query sample. and inverse Pearson correlation coefficient Only when both the positive and negative Pearson correlation coefficients are above the threshold. Furthermore, the joint correlation error of the three types of modal data Less than the threshold Only when the time is right will the candidate segment be added to the candidate queue and sorted from high to low according to the following formula: If the number of candidate segments is insufficient If you enter a search term, the search window will automatically expand. and lower the threshold To the lower limit Continue until the quantity requirement is met.
[0011] Furthermore, the process of performing multi-scale dynamic temporal alignment on candidate fragments and selecting alignment candidates whose matching errors meet the threshold requirements includes: ranking the top... Each candidate fragment undergoes multi-scale dynamic time warping within the scale set. The time axis is aligned layer by layer; within each scale layer, the main peak sequence of the head displacement information channel is used as an anchor point constraint, requiring the global step size change rate to not exceed [a certain value]. ; Calculate the total alignment error between the aligned fragment and the query sample across the three modalities. and will The fragments are written into the alignment candidate set; This is the alignment error threshold; if the alignment candidate set is empty, the search window will be automatically expanded. and lower the threshold to the lower limit Then, re-search and construct a candidate set.
[0012] Furthermore, the alignment candidate set is normalized according to the following formula to obtain the first... The weighted coefficient vector of each segment : ;in, For the first Alignment error of each segment; perform element-wise weighted summation of three modal data components at each time sample point to synthesize the initial interpolation segment. ;exist In the detection of local spikes, if the spike amplitude exceeds four times the average of the five adjacent samples, it is corrected using the neighborhood Savitzky-Golay filtering method to prevent isolated spikes caused by weight imbalance; the interpolation segments are calculated. Energy vector of millisecond waveform on three types of modal data ; Calculate the front and back ends of the gap context respectively. Energy vector of millisecond waveform on three types of modal data and Constructing an energy transfer matrix ;right Perform channel-by-channel amplitude recalibration to obtain interpolation segments with consistent energy. If the recalibration coefficient of any channel is out of range If the corresponding candidate segment is removed from the alignment candidate set and recalculated, the process continues until all channel coefficients are valid.
[0013] Furthermore, in Cut off from both ends millisecond ramp window, the ramp window adopts a cosine half-wave function ; wherein, is time, is the length of the ramp window; the interpolation segment in the ramp window and the gap context are weighted and superimposed according to to output the interpolation segment after boundary smoothing ; the gap position of the first-level cleaning matrix is written back to update the second-level cleaning matrix.
[0014] With the above technical solutions, the present application has the following beneficial effects: the present application realizes comprehensive cleaning and accurate reconstruction of cross-modal data streams of neural potential fluctuations, local blood oxygen concentration changes and head displacement information by frequency synchronous sampling of the electrode array, near-infrared optical probe and inertial measurement element under the same master clock, combined with multi-stage artifact suppression and complete form compensation based on time domain similarity discrimination, so that the system input has high signal-to-noise ratio and time continuity. On this basis, the multi-layer time sequence coding network is used to decode the cross-modal neural representation sequence after cleaning in real time, and the decoding confidence and the consistency degree of digital human action are coupled in a closed loop through adaptive feedback stabilization, which not only ensures the delicate restoration of digital human skeleton and expression, but also controls the whole end-to-end delay within an acceptable range. Thanks to the energy consistency re-calibration and gradual weighting smoothing processing, the present application can effectively eliminate the spectral leakage caused by gap splicing, improve the continuity of the interpolation segment and the context in amplitude and phase, and thus suppress the action jitter and frame loss phenomenon in complex interactive scenarios. In addition, the system supports personalized calibration in the first round of resting state, which can automatically adjust the threshold and decoding model parameters according to the head shape and physiological characteristics of different users, realizing rapid adaptation to diversified use environment. Overall, the present application has significant improvement in accuracy, stability and real-time performance compared with the prior art, and provides reliable and efficient technical support for immersive digital human interaction. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 A system structure schematic diagram of a brain-controlled digital human interaction system based on a brain-computer interface and artificial intelligence is provided for the embodiment of the present application.
[0016] Figure 2 A working principle schematic diagram of a multi-modal neural signal acquisition unit is provided for the embodiment of the present application.
[0017] Figure 3 A processing flowchart of multi-stage artifact suppression in a cross-modal noise suppression and complete form compensation unit is provided for the embodiment of the present application.
[0018] Figure 4 A system architecture schematic diagram of a digital human interaction control unit is provided for the embodiment of the present application. DETAILED DESCRIPTION
[0019] All of the features disclosed in this specification, and / or all of the steps of any method or process specified in this specification, can be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.
[0020] Any of the features disclosed in this specification, unless explicitly stated otherwise, can be replaced by alternative features serving the same, equivalent or similar purpose, to the features disclosed. Thus, unless expressly stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features.
[0021] REFERENCE Figure 1 A brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence, the system comprises: a multi-modal neural signal acquisition unit, a cross-modal noise suppression and gestalt compensation unit, a digital human interaction control unit; the multi-modal neural signal acquisition unit comprises an electrode array, a near-infrared optical probe and a miniature inertial measurement element arranged on the scalp surface and behind the ear of a user at a predetermined interval, and frequency sampling is triggered by the same master clock to obtain three types of modal data, including: neural potential fluctuation, local blood oxygen concentration change and head displacement information, and the three types of modal data are injected into an on-chip ring-shaped double buffer in real time, and a unified time stamp is attached when writing in the on-chip ring-shaped double buffer, cross-modal time alignment is realized, and a cross-modal data stream is obtained; the cross-modal noise suppression and gestalt compensation unit is used for multi-stage artifact suppression on the cross-modal data stream, and then gestalt compensation based on time domain similarity discrimination is performed to obtain a cross-modal neural representation sequence; the digital human interaction control unit is used for realizing digital human control according to the cross-modal neural representation sequence.
[0022] The electrode array captures the microvolt-level transient changes of cortical electrical activity, the near-infrared optical probe detects the light absorption difference of hemoglobin in the cortical microcirculation, and the miniature inertial measurement element senses the rigid body motion of the skull; the three respectively reflect different physical mechanisms of neural excitation, metabolic coupling and external action, and the frequency sampling driven by the same master clock ensures event-level synchronization, so that the subsequent algorithm can understand the brain-blood flow-motion coupling state at the same time from the cross-modal perspective.
[0023] The principle of the cross-modal noise suppression and gestalt compensation unit can be divided into two layers: the first layer is aimed at artifact suppression, and a robust statistics combining stationary template and sliding time window is adopted to estimate the median level and range of the signal baseline stable zone under the assumption of no model, and then the transient impact is detected according to the change of the instantaneous median absolute deviation. The essence of this is to separate the "fast and sharp" non-physiological noise from the "slow and wide" real neural waveform using a statistical quantity insensitive to abnormalities, and through the cross-modal joint threshold, any abnormality of a single sensor can trigger the three-mode synchronous rejection, thereby avoiding time domain misjudgment. The second layer is aimed at gestalt compensation, and the system first scans the primary cleaning matrix to label the sampling empty segment and residual abnormal zone, and then searches for the candidate segment with high bidirectional correlation in the historical window according to the context waveform. Bidirectional correlation ensures the time structure symmetry of the candidate segment in forward prediction and backward playback, and also makes the neural potential fluctuation and local blood oxygen concentration change present consistent phase relationship on both sides of the gap. Then the purpose of multi-scale dynamic time alignment is to find the paired path that minimizes the overall distortion of the three modalities under different time stretching scales, so that under the constraint of head displacement information as an anchor point, local compression or stretching of the time axis not only conforms to the kinematic coherence, but also does not distort the neural and blood flow rhythm excessively.
[0024] The candidate segment with low matching error is weighted and fused, and the weight is proportional to the reciprocal of the alignment error, which embodies the Bayesian idea of "the more similar, the more credible"; after the weight fusion, local filtering is performed to prevent isolated spikes caused by excessive weight of a certain segment. The principle of energy vector amplitude rescaling is that any signal interpolation must preserve the cross-modal energy conservation, in other words, the energy distribution of the interpolation segment in the neural potential fluctuation, local blood oxygen concentration change and head displacement information channels should be consistent with the context, here the energy vector is understood as the time domain power integral, through the energy transfer matrix, the interpolation segment is scaled to be consistent with the average energy of the context, so that the interpolation neither weakens the amplitude of neural events nor amplifies artifacts. If the rescaling coefficient is too large or too small, it indicates that the candidate segment and the context of the gap are not matched in amplitude, so the system reselects the candidate set to ensure physiological reasonableness. The boundary gradual smoothing adopts a cosine half-wave window, because this window function transitions naturally at zero derivative, and can continuously eliminate the frequency spectrum discontinuity at the junction of the interpolation segment and the original segment, thereby avoiding the introduction of frequency leakage in the frequency domain. After the above two layers of processing, the cross-modal neural representation sequence not only retains the coupling information of neuron discharge and blood flow dynamics, but also eliminates motion artifacts and empty segment loss, providing high signal-to-noise ratio, high time resolution, and cross-modal energy consistent input for digital human interaction control unit.
[0025] The digital human interaction control unit (DHU) in the overall system is responsible for the real-time mapping of cross-modal neural representation sequences to digital human action commands. Its implementation process can be summarized into four consecutive stages: initialization and adaptive calibration, real-time intent decoding, action synthesis and distribution, and interaction quality feedback and stabilization. In the initialization and adaptive calibration stage, the data access manager first registers and subscribes to cross-modal neural representation sequences on the on-chip pass-through bus. The system recommends a two-level circular buffer for zero-copy transfer: the first-level buffer's write end corresponds to the output thread of the cross-modal noise suppression and gestalt compensation unit; the second-level buffer's read end is bound to the main processing thread of the DHU. The read and write ends maintain queue synchronization through register-level timestamp comparison, ensuring that the maximum buffer latency does not exceed fifty milliseconds. Subsequently, the calibration task scheduler is started, guiding the user to execute a predefined set of actions (e.g., gaze switching, simple tilting and shaking, clenching and releasing a fist), while simultaneously recording the cross-modal neural representation sequence and the digital human's standard action label pair for subsequent intra-domain fine-tuning of the decoding network. To facilitate rapid convergence, a strategy of freezing the backbone parameters and updating only the biases of higher-order attention and output layers can be adopted. A typical learning rate of one to two times ten to the power of negative four is sufficient to meet the convergence requirement within ten seconds. After calibration, a snapshot of the calibration weights is generated for loading during the real-time inference stage.
[0026] In the real-time intent decoding stage, the cross-modal neural representation sequence first enters the sequence normalization pipeline. The pipeline applies sliding zero-mean correction and percentile pruning to neural potential fluctuations, local blood oxygen concentration changes, and head displacement information, respectively, and then concatenates them along the feature dimension to form a multi-channel tensor. This tensor is then fed into a multi-layer temporal coding network. The recommended structure is a single 3D depthwise separable convolution to achieve millisecond-level local feature extraction, followed by bidirectional gated recurrent units to capture contextual dependencies, and then two layers of lightweight Transformer encoders with relative position encoding are stacked. Finally, attention pooling converges the vector into a driving intent vector. To ensure inference latency is controlled within 20 milliseconds, half-precision computation can be enabled on the deployment end and bound to a specific GPU stream. Once generated, the driving intent vector is immediately fed into the action mapper. The action mapper has a built-in skeletal joint mapping table and an expression shape mapping table, both stored in YAML format. Fields include the target bone node name, rotation quaternion index, and expression blendshape weight index, etc. The motion mapper, based on the semantic components of the driving intent vector, calls the linear projection layer to output joint rotation increments and facial expression weight increments, and adds a temporal smoother to generate frame-to-frame transitions. The smoother employs a first-order Butterworth filter and automatically reduces the cutoff frequency at key motion boundaries to balance response speed and visual continuity.
[0027] The action synthesis and distribution stage is managed by an action bus dispatcher. This dispatcher compresses each frame of joint rotation increments into sixteen-bit fixed-point format and expression weight increments into eight-bit unsigned integers after vector quantization, and then encapsulates them into a unified action frame structure. The action frame structure contains a timestamp, a frame number, an array of joint rotations, an array of expression weights, and a CRC check field. The dispatcher pushes the action frame to the digital human rendering engine via zero-copy shared memory or NamedPipe; if the system needs to face multiple display terminals, a WebSocket bridge can be added to broadcast the action frame to each rendering terminal after being packaged into JSON or Protobuf. In order to ensure synchronization under network fluctuations, the bridge has embedded adaptive bitrate logic: when the end-to-end latency exceeds two hundred milliseconds or the packet loss rate exceeds five percent, the action frame downsamples from sixty frames per second to thirty frames per second, and notifies the decoding end to expand the shared memory buffer by one hundred percent.
[0028] The interaction quality feedback and stabilization stage is evaluated by two complementary indicators: digital human motion consistency and decoding confidence. Motion consistency is evaluated by comparing the mean square error between the actual pose and expression weight values read back in the rendering thread and the target values in the action frame; decoding confidence is derived from the entropy of the attention weight distribution output by the intent decoding network. The feedback controller performs exponential smoothing on both indicators to obtain short-term and long-term trends. When the short-term trend worsens and the long-term trend is good, the system infers that it is a transient artifact, triggering the cross-modal noise suppression and gestalt compensation unit to increase the cleaning level. If the long-term trend also continues to worsen, the adaptive recalibration sub-process is entered: freeze the rendering refresh thread, only keep the data access and decoding threads, and the user performs a simple instruction demonstration for a few seconds, the system updates the decoding network high-level parameters online, this process does not reset the joint mapping table, so it can recover the interaction within thirty seconds. If the decoding confidence drops below thirty percent and does not recover for five seconds, the system activates the safety exit logic: the action bus dispatcher immediately issues zero-increment frames, and the digital human slowly resets to the resting pose, and automatically restarts the entire process when the confidence recovers to fifty percent or more.
[0029] The thread organization of the whole digital human interaction control unit is recommended to adopt an event-driven model: the main processing thread is responsible for pulling the cross-modal neural representation sequence from the ring buffer and starting the intention decoding; the action bus scheduling thread runs independently and is responsible for frame compression and distribution; the feedback control thread periodically reads the action execution reward and attention entropy and updates the global control variables. Lightweight messages are passed between threads through lock-free queues, and GPU stream synchronization primitives are used to ensure timing consistency during core computation. In software implementation, it is recommended to deploy the decoding network based on PyTorch or TensorRT, SkeletonAPI can be bound to Unity or UnrealEngine, and the communication layer can use ZeroMQ or ROS topics; in hardware, a single NVIDIA RTX 4060 level GPU can meet the real-time requirements of sixty frames per second and an end-to-end latency of less than 120 milliseconds.
[0030] Further, the cross-modal noise suppression and gestalt compensation unit, the process of multi-level artifact suppression on the cross-modal data stream includes: calculating the median and range of the first sliding time window for the three types of modal data respectively; if the sum of the ranges of the three types of modal data in the first sliding time window is less than the set steady state threshold, the maximum value of the range of the current three types of modal data is recorded as the zero-level steady state template; three layers of sliding time windows with increasing length are maintained for the cross-modal data stream according to the channel, and the median absolute deviation of each layer is updated in real time; the median absolute deviation of the current layer is multiplied by the sub-window coefficient, and the range proportion of the zero-level steady state template is added to obtain the fusion threshold of multiple layers; if the absolute amplitude of the value of any type of modal data exceeds the fusion threshold of the corresponding layer and the duration is shorter than the minimum duration of the neural event, the timestamp range of the three types of modal data is uniformly marked as a first-level artifact section; fixed-length normal waveforms are extracted before and after the first-level artifact section, respectively, and three times of spline interpolation are performed on the neural potential fluctuation and local oxygen concentration change, respectively, and then the bilinear interpolation is performed on the head displacement information, and then the first-level cleaning matrix is spliced.
[0031] After the system is started, the electrode array, near-infrared optical probe and miniature inertial measurement element are written into the on-chip buffer synchronously, and the cross-modal data stream is then sent to the front-end buffer of the unit. The algorithm calculates the median and range of the three types of modal data immediately when the buffer is full of the first sliding time window, and quickly judges whether the current sequence is in a physiological resting state by comparing the sum of the ranges with the preset steady state threshold; if it is determined to be resting, the maximum value of the range of the three types of modal data is fixed as a zero-level steady state template, which is considered as a ruler for cross-modal background noise in the future and participates in threshold fusion in the form of range proportion.
[0032] In the subsequent real-time running phase, the algorithm maintains three layers of sliding time windows with increasing length for each channel. This increasing strategy ensures that both short-term anomalies and long-term trends can be captured simultaneously. The median absolute deviation of the corresponding channel is updated in real time when each layer of sliding time window expires, making the threshold calculation insensitive to transient impact and adaptive to baseline drift. In order to unify the local fluctuations within the channel with the cross-modal resting amplitude, the algorithm multiplies the current layer median absolute deviation with the sub-window coefficient and adds the range proportion of the zero-level steady-state template to form a multi-layer fusion threshold array. This fusion threshold array has global alignment characteristics for the three types of modal data, which can effectively avoid false alarms caused by sensor sensitivity differences.
[0033] When the absolute amplitude of any type of modal data breaks through the corresponding layer of fusion threshold and the duration is shorter than the minimum duration of neural events, the algorithm will consider that the impact lacks neurological significance and is more likely to be caused by electrode slippage, light obstruction or transient head shaking. Therefore, the same time stamp of the three types of modal data is uniformly marked as a first-level artifact segment. This cross-modal synchronous marking prevents artifact residues caused by channel delay differences. After the first-level artifact segment annotation is completed, the algorithm extracts fixed-length normal waveforms before and after the segment as context for subsequent interpolation reconstruction.
[0034] In the interpolation phase, the system follows the principle of signal prototype continuity: cubic spline interpolation is used for neural potential fluctuations and local blood oxygen concentration changes. This high-order polynomial interpolation has second-order derivative continuity in the time domain, which can maintain waveform smoothness without excessive distortion of peaks. Bilinear interpolation is used for head displacement information because the physical meaning of this information is rigid body translation and rotation. Linear interpolation can maintain consistent displacement increments and avoid introducing false frequencies in low sampling intervals. After interpolation, the three types of modal data are spliced into a once-level cleaning matrix according to the original time axis. The matrix maintains time dimension alignment and channel dimension separation in structure, which facilitates gap detection in the subsequent gestalt compensation phase and ensures that the signals received by the digital human interaction control unit meet the minimum time delay and the highest integrity. The entire multi-level artifact suppression process uses a steady-state template initialization, an increasing length sliding window adaptive update, a median absolute deviation driven fusion threshold generation, and a cross-modal synchronous marking strategy, so that the system can immediately eliminate most motion artifacts and sensor burst noise without relying on complex models and priori. The subsequent interpolation mechanism restores the broken waveform to a continuous trajectory that conforms to the physiological law through physically consistent signal reconstruction, laying the foundation for the subsequent gestalt compensation unit to further retrieve historical segments and fill in the gaps.
[0035] Further, the range proportion of the zero-level steady-state template is defined as the ratio of the sum of the range of the zero-level steady-state template and the three types of modal data within the first sliding time window.
[0036] The extreme difference proportion of the zero-level steady-state template is defined as the ratio of the maximum extreme difference of the three types of modal data in the first sliding time window to the sum of the extreme differences of the zero-level steady-state template and the three types of modal data. The fundamental purpose is to provide a unified and adaptive amplitude baseline for the subsequent multi-layer fusion threshold. When the cross-modal data stream enters the cross-modal noise suppression and gestalt compensation unit, it has not been cleaned up yet. The absolute amplitude, signal form and noise distribution of the three types of modal data often differ by orders of magnitude. If the amplitude index of a single channel is directly used in threshold design, it is easy to lead to the fusion threshold being dominated by the accidental peak of a certain modality, thereby reducing the robustness of artifact detection. On the contrary, if the average is completely processed, the protection of the most sensitive modality will be weakened, increasing the risk of missed detection. Therefore, the algorithm first calculates the extreme difference of the three types of modal data independently in the first sliding time window, and selects the maximum one as the zero-level steady-state template, which represents the highest background noise amplitude that the system may encounter in the resting state. At the same time, in order to avoid the maximum extreme difference being amplified too much in the subsequent threshold fusion, the system compares the maximum extreme difference with the sum of the three types of extreme differences to obtain the extreme difference proportion of the zero-level steady-state template. The proportion naturally falls between zero and one, which is equivalent to normalizing the maximum extreme difference, making it a relative quantity independent of dimension and comparable across modalities, rather than an absolute amplitude. When calculating the multi-layer fusion threshold, the median absolute deviation multiplied by the sub-window coefficient can only reflect the local fluctuations within the channel, and still lacks the alignment of the global noise background. By superimposing the extreme difference proportion of the zero-level steady-state template, all channels share the same resting amplitude reference, ensuring that the threshold size in cross-modal joint determination is neither excessively biased towards high-noise channels nor ignores low-noise channels.
[0037] Further, the cross-modal noise suppression and gestalt compensation unit includes the following steps in the process of gestalt compensation based on time domain similarity discrimination: detecting and indexing the sampling empty segment or residual abnormal area existing in the three types of modal data to form a gap list; based on the gap context waveform, searching for candidate segments that meet the bidirectional correlation condition in the historical period and constructing a candidate set; performing multi-scale dynamic time alignment on the candidate segments, and screening alignment candidates whose matching errors meet the threshold requirements; performing similarity weighted fusion on the alignment candidates according to the alignment error, and eliminating local abnormal peaks through filtering processing to obtain an interpolation segment; performing amplitude re-labeling based on energy vector on the interpolation segment to realize cross-modal energy consistency; performing gradual weighting smoothing processing on both ends of the interpolation segment, and splicing to the secondary cleaning matrix to form a secondary cleaning matrix.
[0038] Firstly, the sampling empty section or residual abnormal area existing in the three types of modal data is detected and indexed, forming a gap list. The internal mechanism of this detection is to use the unified time axis obtained in the primary cleaning stage to calculate the local residual and local stationary statistics for each time sampling point. When the absolute value of the residual continuously exceeds the threshold and the duration exceeds the minimum empty section threshold, it is identified as a sampling empty section. When the residual presents isolated peak state and the local energy is significantly higher than the average energy of the context, it is marked as a residual abnormal area. Since the three types of modal data have synchronicity in physical coupling, the gap index is unified in the gap list through timestamp alignment, ensuring that the subsequent reconstruction remains consistent across modalities. Then the algorithm uses the gap context waveform as a search template to search for candidate segments that meet the bidirectional correlation condition in the historical period and construct a candidate set. Bidirectional correlation requires that the candidate segment simultaneously obtains high correlation coefficients in forward matching and reverse matching to ensure that its time sequence structure can both predict the front edge of the gap and fit the rear edge of the gap, thereby maximizing the preservation of the true phase relationship of neural discharge-blood flow coupling.
[0039] To fully adapt to different rhythms of deformation, each candidate segment in the candidate set then enters the multi-scale dynamic time alignment module. The algorithm performs alignment at multiple time stretching scales and uses the main peak sequence in the head displacement information as an anchor point constraint to limit global step changes and prevent excessive time stretching. After alignment, the matching error is calculated and the alignment candidate set with an error less than a set threshold is selected. This step uses error as an indicator to eliminate large mismatched segments, ensuring that the reconstruction is based on high confidence samples. The system then performs similarity weighted fusion of the alignment candidates based on the alignment error. The similarity weighting follows the principle that the smaller the error, the higher the weight, allowing multiple candidate segments to collaboratively complete the gap in the time domain. At the same time, local abnormal peaks that may occur during the fusion process are eliminated through neighborhood filtering, and the interpolated section is output.
[0040] To make the energy distribution of the interpolation segment consistent with the gap context on the three types of modal data, the system performs amplitude rescaling based on the energy vector for the interpolation segment. The energy vector is regarded as a cross-modal power characterization here, and the rescaling process calculates scaling coefficients independently for each channel. The coefficients are then checked as a whole to see if they fall within a reasonable physiological range. If any of them are out of range, the candidate segment is reselected and the process is repeated to ensure consistency in cross-modal energy. After amplitude correction of the interpolation segment, it still needs to be seamlessly spliced with the original waveform. To this end, the algorithm applies a gradual weighting smoothing process to the two ends of the interpolation segment. The gradual window is defined according to the cosine half-wave function, so that the first derivative at the endpoints is continuous, thereby avoiding spectral leakage caused by amplitude jumps. The processed interpolation segment is written back to the corresponding gap position in the primary cleaning matrix to form a secondary cleaning matrix. The secondary cleaning matrix is not only continuous in time domain without gaps, but also consistent with the surrounding waveforms in amplitude, energy and phase structure, providing high-integrity and high-SNR cross-modal neural representation sequences for the subsequent intention decoding of the digital human interactive control unit.
[0041] Further, the method for detecting and indexing the sampling empty segment or residual abnormal area existing in the three types of modal data to form a gap list comprises: detecting the sampling empty segment or the segment with an absolute value of the splicing residual exceeding twice the local median absolute deviation of any type of modal data by traversing the uniform time axis of the primary cleaning matrix; recording the start index , the end index , the residual root mean square value of the neural potential fluctuation , the residual root mean square value of the local blood oxygen concentration change , and the residual root mean square value of the head displacement information ; generating a gap tuple , and writing it into the gap list; if the interval between two consecutive gaps is less than the minimum period of neural events, they are merged into a single gap entry to prevent boundary tearing caused by fragmented reconstruction.
[0042] The algorithm traverses the primary cleaning matrix frame by frame, calculates the absolute value of the three types of modal splicing residual at each sampling point, and compares it with the local median absolute deviation of the corresponding channel updated in real time with a sliding window. When any channel satisfies and continuously spans at least two sampling periods, it is determined that the interval is a potential gap candidate; The values of the sampling empty segment or the absolute value of the splicing residual are , , and ; if the same interval is accompanied by a sampling empty segment flag, it is directly classified as a sampling empty segment. After detecting the candidate segment, the system records its start index and end index Record and calculate the root mean square error of the three types of modal splicing residuals in the interval , , wherein is the number of sampling points in the section, and the three residual power measures collectively describe the influence of the gap on the neural potential fluctuation, local blood oxygen concentration change, and head displacement information. Subsequently, the algorithm generates a gap tuple and writes it to the gap list. The gap tuple provides both time domain positioning information and cross-modal error quantification, providing necessary input for subsequent candidate segment retrieval and energy vector amplitude rescaling.
[0043] To avoid signal boundary tearing caused by multiple annotations of short-time repeated artifacts, the system continues to sequentially check the intervals between adjacent entries in the gap list along the unified time axis. If the difference between the start index of the two entries and the end index of the previous entry is less than the minimum period of neural events, it is considered that the two abnormalities belong to the same physiological or artifact cause, and gap merging is immediately performed: the new takes the original value of the former, the new takes the original value of the latter, and the root mean square error is recalculated using the sum of the squares of the two segments to maintain energy conservation. This merging strategy can effectively prevent fragmentation reconstruction during candidate retrieval and dynamic time warping, thereby reducing the number of interpolation segment boundaries, reducing the complexity of boundary smoothing, and ensuring the calculation of the rescaling matrix based on complete segment statistics rather than broken segment averages, improving the timing continuity and amplitude stability of brain-controlled digital human motion decoding.
[0044] Further, based on the gap context waveform, the process of retrieving candidate segments that meet the bidirectional correlation condition in the historical period and constructing the candidate set includes: for each gap, the context waveform of the gap before and after milliseconds is used as the query sample; in the historical section seconds before the gap timestamp, candidate segments are extracted in a sliding window with a step size of milliseconds, and the forward Pearson correlation coefficient and the reverse Pearson correlation coefficient are calculated for the query sample; only when the forward Pearson correlation coefficient and the reverse Pearson correlation coefficient are both higher than the threshold , and the joint correlation error of the three types of modal data is less than the threshold , the candidate segment is pushed into the candidate queue, and is sorted from high to low according to the following formula: If the number of candidate segments is less than , the search window is automatically expanded and the threshold is reduced to the lower limit Continue until the quantity requirement is met.
[0045] To ensure the physiological consistency of the reconstructed segment with the original waveform in terms of phase and energy structure, the cross-modal noise suppression and gestalt compensation unit first extracts the phases before and after each notch for each notch. The context waveform in milliseconds is used as the query sample. This context simultaneously preserves the joint temporal features of three modalities: neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. Then, the most recent time before the gap timestamp... Within the historical segment of seconds A sliding window with a fixed step size is used to extract candidate fragments. This strategy balances retrieval speed and coverage density, ensuring sufficient sampling of the historical pattern library with limited computing resources. For each candidate fragment, the system first calculates its positive Pearson correlation coefficient with the query sample on the original time axis. Then, the candidate fragments are time-reversed and aligned with the query samples to calculate the inverse Pearson correlation coefficient. The design principle of bidirectional correlation is that the waveforms of neural discharge and metabolism coupling often exhibit symmetrical characteristics. High positive correlation ensures rhythm matching, while high negative correlation can capture time-reversed symmetrical patterns, thereby improving the measurement accuracy of the overall similarity of waveforms on both sides of the gap.
[0046] The system further performs joint correlation error analysis on the three types of modal data. , It combines the mismatches of positive and negative directions and is a normalized error index for measuring the overall similarity of candidate segments; only when , and Only then does the system determine that the joint temporal pattern of the fragment across the three modalities is highly consistent with the query sample, and thus pushes it into the candidate queue. The candidate queue adopts a... Sorting the keys in descending order of priority, this average relevance metric balances bidirectional matching while also being numerically aligned directly to... This interval facilitates the rapid selection of the top-ranked items during the subsequent dynamic time alignment phase. Perform multi-scale alignment on high-confidence segments. If the number of candidate segments in the queue is insufficient... The system immediately triggers an adaptive window expansion mechanism, expanding the historical search window. Gradually expand to And simultaneously set the relevant thresholds Decrease to the lower bound with a linear step size This progressive search framework ensures that sufficient candidate samples can be obtained even when samples are sparse or the user's specific brainwave patterns are observed, while also... A hard constraint is set for the minimum similarity to prevent quality degradation. The entire candidate set construction process relies on four mechanisms: context constraint, bidirectional correlation, joint correlation error, and adaptive window expansion. The selected segments simultaneously satisfy high structural homogeneity in the three channels of neural potential fluctuation, local oxygen concentration change, and head displacement information, providing a high confidence starting point for subsequent multi-scale dynamic time warping and energy vector amplitude rescaling, thereby ensuring the motion coherence and signal stability of the brain-controlled digital human interaction link at the system level.
[0047] Further, multi-scale dynamic time warping is performed on the candidate segments, and the process of screening alignment candidates with matching errors meeting threshold requirements includes: Performing multi-scale dynamic time warping on the top candidate segments, aligning the time axis layer by layer on the scale set ; in each scale layer, the main peak sequence of the head displacement information channel is used as an anchor constraint, and the global step change rate is required to be no more than ; calculate the total alignment error of the three types of modal data of the aligned segments and the query sample, and write the segment into the alignment candidate set; The alignment error threshold is set; if the alignment candidate set is empty, the search window is automatically expanded and the threshold is lowered to the lower limit , and the candidate set is constructed by re-searching.
[0048] Multi-scale dynamic time warping is designed to maximize the fit of cross-modal rhythms in the context of gaps while ensuring real-time performance. Therefore, the algorithm first takes out the top candidate segments from the candidate queue, and performs multi-scale dynamic time warping on each candidate segment. Multi-scale refers to the system expanding the time axis layer by layer on the scale set , with representing time compression to half the original, with representing time stretching to twice the original, and representing equal ratio alignment baseline. In the warping process, each scale layer uses the main peak sequence of the head displacement information channel as an anchor constraint. These main peaks usually correspond to the phase extreme values of cranial rigid body motion, and have obvious synchronization relationship with neural potential fluctuation and local oxygen concentration change in physical coupling, so they are selected as anchor points to prevent phase drift during cross-modal alignment. To limit the time axis from being excessively stretched or compressed, the algorithm specifies that the global step change rate should not exceed , that is, the local scaling ratio difference from the original ratio is limited to the interval
[0049] .After alignment, the system calculates the alignment residuals for each channel of the three types of modal data and sums them up to obtain the total alignment error. Its calculation method maintains energy consistency with the aforementioned joint correlation error index, giving the error a cross-modal normalization property; when If the phase and amplitude of the candidate fragment match the query sample at the current scale, the fragment is immediately added to the alignment candidate set. If a fragment within the same scale layer has not yet passed the threshold test, it automatically proceeds to the next scale to continue alignment until all scales have been traversed. After completing the three-scale traversal, if the alignment candidate set is still empty, the system determines the history window. Length or similarity threshold The criteria were too strict, so a windowing and re-search mechanism was invoked. Expand by arithmetic increments and Linearly decreasing to the lower limit Then, the entire process of candidate retrieval, bidirectional relevance screening, and multi-scale dynamic time warping is executed again; this progressive threshold-lowering strategy ensures that acceptable candidate segments can still be found even when the user's brainwave morphology changes slowly or sampling noise increases, while also... A minimum relevance cap was set to avoid introducing morphological bias due to excessive relaxation of standards. The fragments finally written into the alignment candidate set are not only homogeneous with the query samples on the time axis, but also meet the threshold requirements in terms of cross-modal coupling strength and energy distribution, laying a high-confidence foundation for subsequent similarity-weighted fusion and energy vector magnitude recalibration. In this way, the reconstructed waveform in the secondary cleaning matrix can maintain physiological rationality while achieving high compatibility with the real-time requirements of the digital human interaction control unit, thus allowing the brain-controlled digital human to maintain a smooth, stable, and latency-controllable interactive experience in complex action sequences.
[0050] Furthermore, the alignment candidate set is normalized according to the following formula to obtain the first... The weighted coefficient vector of each segment : ;in, For the first Alignment error of each segment; perform element-wise weighted summation of three modal data components at each time sample point to synthesize the initial interpolation segment. ;exist In the detection of local spikes, if the spike amplitude exceeds four times the average of the five adjacent samples, it is corrected using the neighborhood Savitzky-Golay filtering method to prevent isolated spikes caused by weight imbalance; the interpolation segments are calculated. Energy vector of millisecond waveform on three types of modal data ; Calculate the front and back ends of the gap context respectively. Energy vector of millisecond waveform on three types of modal data and Constructing an energy transfer matrix ;right Perform channel-by-channel amplitude recalibration to obtain interpolation segments with consistent energy. If the recalibration coefficient of any channel is out of range If the corresponding candidate segment is removed from the alignment candidate set and recalculated, the process continues until all channel coefficients are valid.
[0051] After multi-scale dynamic time warping is completed and an alignment candidate set is obtained, the system needs to fuse several candidate segments with different alignment errors into a continuous and energy-consistent reconstructed waveform. The algorithm first relies on... The normalization formula is used to calculate the first The weighted coefficient vector of each segment, where The total alignment error is taken from the multi-scale dynamic temporal alignment output of the previous step. Normalization is achieved by summing all coefficients by segment dimension and then dividing each by the sum, ensuring that the total weight of all candidate segments is always 1. Simultaneously, segments with smaller errors are implicitly assigned higher weights, directly mapping alignment accuracy to fusion contribution. Subsequently, the system performs element-wise weighted accumulation of the three modal data components at each time sample point on a unified time axis to obtain the initial interpolation segment. Because the small local phase differences between different candidate segments may still cause weight superposition at very few time points, thus introducing isolated amplitude spikes, the system immediately... The internal process performs local spike detection: iterates through each sampling point, and if the current sampling value exceeds four times the average of the five sampling points to its left and right, it is identified as a spike and the neighborhood Savitzky-Golay filter is called to re-estimate the smoothed values of the point and three to five neighboring sampling points. The polynomial least squares property of Savitzky-Golay ensures that the peak value is softened while preserving the higher-order morphological information, avoiding isolated spikes caused by weight imbalance that could disrupt the coupling relationship between nerve potential fluctuations, local blood oxygen concentration changes and head displacement information.
[0052] After the spike correction is completed, the system uses Calculate energy vectors on three types of modal data using milliseconds as the time window. Energy is defined using the sum of squares of samples within a window, normalized by the number of sampling points to ensure comparability of energy across channels. Simultaneously, the same length is taken for both the front and back ends of the notch context. Millisecond waveform, calculate energy vector and Three sets of energy vectors were used to construct the energy transfer matrix. Each diagonal element in the matrix reflects the scaling ratio required for the interpolation segment in the corresponding channel, so that the average energy of the interpolation segment matches the geometric center of the energy on both sides of the gap context; this is achieved channel-by-channel scaling. Multiply the corresponding scaling factor of the matrix, the system generates consistent energy interpolation segment At this time, the algorithm checks whether all channel recalibration coefficients fall between If any channel coefficient exceeds the boundary, it means that the energy distribution of the candidate segment of that channel is too different from the context, which may be caused by abnormal noise or mispairing, so the corresponding candidate segment is removed from the alignment candidate set, and the remaining segment weight is normalized and returned to the step of recalculating the weighted cumulative sum energy vector, and the whole process is iterated until all channel coefficients are legal.
[0053] This "remove-normalize-recompute" closed loop ensures that the interpolation segment will not be imbalanced as a whole due to a single high or low energy segment, thereby maintaining the cross-modal energy consistency of neural potential fluctuations, local oxygen concentration changes and head displacement information. The final Both in the time domain and the phase domain, the details of the minimum error candidate segment are fully inherited, and in the energy domain, the migration matrix and the context are used to complete accurate alignment, which lays a solid foundation for subsequent gradual weighting smoothing and secondary cleaning matrix splicing, and also ensures that the digital human interaction control unit can receive a stable amplitude and coherent phase cross-modal neural representation sequence when decoding, thereby continuously outputting coherent and natural digital human actions.
[0054] Further, at the beginning and end of the , a millisecond fade window is intercepted, and the fade window uses a cosine half-wave function ; where is the time, is the length of the fade window; the interpolation segment and the gap context in the fade window are weighted and superimposed according to , and the interpolation segment after boundary smoothing is output; and is written back to the gap position of the primary cleaning matrix, which is updated to the secondary cleaning matrix.
[0055] After obtaining the energy-consistent interpolation segment , the splicing boundary of the interpolation segment and the gap context still needs to be smoothed to eliminate the sense of discontinuity caused by the phase difference between the interpolation and the original waveform. The algorithm intercepts a millisecond fade window at the beginning and end of the , and weights the interpolation segment samples according to the cosine half-wave function inside each fade window; where is the time from the beginning of the fade window, with a value range . The function has a weight of zero at , and a weight of one at , and the weight outside the gap context is Because they are complementary, the two waveforms are cosine-weighted and superimposed linearly within the transition window region, resulting in a smooth transition. The cosine half-wave function has the property that its zeroth and first derivatives are both zero at both ends, ensuring that not only the amplitude but also the slope is continuous at the splicing point, thus avoiding sharp transition bands and frequency leakage in the frequency domain. The output after weighted superposition is the smoothed interpolation segment. The system immediately... Write back to the gap interval of the primary cleaning matrix, replace the original gap labels, and update it with the secondary cleaning matrix. At this point, the secondary cleaning matrix no longer contains any breaks on the time axis and is seamlessly aligned with the surrounding waveforms in terms of amplitude, phase, and energy. This provides the digital human interaction control unit with a highly complete, high signal-to-noise ratio cross-modal neural representation sequence, ensuring that the brain-controlled digital human maintains continuous, stable, and low-latency action output in subsequent interaction scenarios.
[0056] The formula for weighted fusion of candidate fragments is written as follows:
[0057] ;
[0058] in This indicates that on a unified timeline, for channels At the moment The values of the initial interpolation segment obtained; For the first The resampled waveform of a candidate segment aligned with the gap after multi-scale dynamic time warping; The total alignment error of the three modalities of the corresponding segment; the smaller the error, the higher the similarity. (Fractional part) Define the weighted coefficient vector Ensure that the sum of the weights of all candidate fragments is and Inversely proportional to alignment accuracy, high-precision segments contribute more to the fusion process; This represents the number of segments that will enter the alignment candidate set.
[0059] The formula for energy vector magnitude recalibration and interpolation segment scaling is written as:
[0060] ;
[0061] in Indicates the initial interpolation segment in the channel above Energy vector components calculated using milliseconds as the window; and These are the energy components of the long waveform in the front and back ends of the gap context, respectively. The range of values for the channel-by-channel amplitude recalibration coefficient is subject to... Restraints are necessary to ensure physiological rationality; For the final interpolation segment with consistent energy, energy conservation between the interpolation segment and the context is achieved across the cross-modal scale. The first formula adaptively fuses the phase-amplitude information of multiple segments according to the error, and the second formula precisely aligns the fusion result with the context in the energy domain. The system can output a cross-modal neural representation sequence that is both phase-continuous and energy-balanced, thereby meeting the dual requirements of temporal integrity and amplitude consistency of the brain-controlled digital human interaction link.
[0062] Suppose the electrode array worn by the user contains The near-infrared optical probe contains one electrode. Transmitter-receiver pair, miniature inertial measurement element output Axial acceleration and axial angular velocity, total One channel; master clock sampling rate set to [value]. The length of the first sliding time window after system startup is taken as... (Right now (Sampling points). Within this window, the ranges of the three modalities are respectively , (Relative changes in near-infrared light absorption) The sum of the ranges is Set steady-state threshold ,because The system confirms that it is in a steady state and will This is denoted as the zero-order steady-state template; its range ratio is... After entering real-time operation, settings are configured for each channel. Layered sliding window: According to EEG number Taking the channel as an example, during a certain period of time Time corresponding The absolute deviation of the median within the window is Sub-window coefficients are taken The fusion threshold is At this point, the absolute amplitude of the channel was detected to be present. Crossing but continuing Minimum duration of neural events Meanwhile, neither the fNIRS nor the IMU channels exceeded the threshold, so the three modes were uniformly labeled. This is a first-level artifact section. Samples were taken from both the front and back. Normal waveforms are processed by cubic spline interpolation of EEG and fNIRS, and bilinear interpolation of IMU, generating a primary cleaning matrix. During continued scanning, ... arrive The absolute value of EEG plugging residuals was found. And accompanied by fNIRS sampling gaps, recording , , Generate gap tuples Then, with respect to both sides of the gap. The waveform is the query sample, in the most recent Within the window of history Step size sliding, to obtain candidate segments Item. Calculate the positive correlation coefficient. With inverse correlation coefficient After that, only Conditions satisfy and ,according to Sort by first The data enters a multi-scale dynamic time warping process.
[0063] Taking the first segment as an example, its scale Top alignment error ,scale With scale The errors are respectively and Select the record with the minimum error. The final candidate set is then aligned and retained. The segment, its error vector .according to Obtain normalized weights .right Interpolation interval and three modes Sample-by-sample weighted summation Detect spikes in S_raw, only in the EEG channel. Peak was found at the location Exceeding four times the neighborhood mean, applying a Savitzky-Golay third-order window length. Sample filtering correction Then take Calculate the energy vector (The units are respectively) , variance, The average of the gap context energy vectors is as follows: Diagonal elements of the energy transfer matrix All fell into Therefore, it can be recalibrated directly: .right Take from both the front and back Gradient window, press Overlay with the gap context to obtain the smoothed version .Will Write back to the primary cleaning matrix After the segment, the secondary cleaning matrix is completed. The digital human interaction control unit checks the phase consistency between this segment and the context EEG-fNIRS-IMU in real time, improving it to... Overall end-to-end latency is controlled within This meets the needs of smooth interaction for brain-controlled digital humans.
[0064] After completing the secondary cleaning matrix, the system... Perform a window aggregation on the cross-modal neural representation sequences, and then... Lu EEG, fNIRS and The IMU channel values are normalized by channel and then concatenated to obtain the length. The instantaneous vector is then appended with a second-order difference to obtain the total dimension. Then connect them together The frame length is The timing block input to the digital human interaction control unit is a first-layer bidirectional gated loop network; the hidden dimension of this network is set. Because the time step is Therefore, the output dimensions of the forward and reverse hidden state tensors are... The system obtains the hidden state vector through time-dimensional max pooling. Then, through linear transformation Will Projected to Dimensional driving intent space, in which , The driving intent vector is then... generate The joint angle increment, , , For hyperbolic tangent. The joint mapping table specifies the first... Dimension mapping to the head quaternion, Mapped to the cervical segment, Corresponding to the degrees of freedom of the two arms, torso, and legs; for example: in Get it at all times The digital human's head immediately rotates slightly around three axes, with the motion frames in... Rate packaged as Push to the rendering end. To suppress micro-judder, the motion bus scheduler performs a first-order Butterworth filter on consecutive frame increments, with a cutoff frequency set to... The system is in Frame consistency error detected less than the threshold Decoding confidence Therefore, the current decoding weights are maintained. If persistent [issues] occur at the joints of both arms... low amplitude fluctuations and The feedback control thread immediately raises the cross-modal noise suppression level to the sliding window. and put Temporary increase To tighten the candidate relevance. One complete interaction cycle from cross-modal neural representation sequence generation to digital human rendering end output motion frame lasts , where the cleaning reconstruction takes time , the intention decoding and action mapping take time , the frame compression and transmission take time , and the rest is the GPU synchronization and rendering overhead; After running continuously , the average of the overall motion consistency is measured , the standard deviation is , the end-to-end delay drift is , which shows that this specific numerical scheme can stably support the smooth interaction of brain-controlled digital humans under real hardware configuration.
[0065] Figure 2 The working principle and data acquisition process of the multi-modal neural signal acquisition unit of the application are shown in detail. The acquisition unit adopts a distributed sensor array configuration, and electrodes, near-infrared optical probes and miniature inertial measurement elements are arranged on the scalp and behind the ears of the user at a predetermined interval to realize the synchronous acquisition of three types of different modal neural signals. Specifically, the electrode array is responsible for acquiring the neural potential fluctuation signal, which reflects the electrical activity changes of the brain cortical neuron group, has high time resolution characteristics, and the sampling frequency usually reaches more than 500Hz. The near-infrared optical probe is used to monitor the local blood oxygen concentration changes, and through the detection of the spectral absorption difference of hemoglobin and deoxyhemoglobin, it indirectly reflects the local metabolic activity of the brain, and the signal has good spatial resolution but relatively slow time response. The miniature inertial measurement element mainly acquires head displacement information, including three-axis acceleration and angular velocity data, which is used to monitor and compensate the motion artifact interference caused by head motion to the other two types of signals. The system adopts a unified master clock triggering mechanism to ensure the frequency synchronous sampling of the three types of modal data. All the acquired raw data is injected into the on-chip ring-shaped double buffer in real time, and the buffer adopts a ping-pong cache architecture, automatically adds a unified time stamp label during data writing, and realizes the accurate time alignment of cross-modal data. This design effectively solves the time synchronization problem caused by the inherent delay difference of different sensors, and lays a reliable time reference for subsequent cross-modal data fusion and analysis. The capacity of the ring-shaped double buffer is designed considering the real-time processing requirements and hardware resource limitations, which can minimize the system delay while ensuring data integrity.
[0066] Figure 3The complete processing flow of multi-level artifact suppression in cross-modal noise suppression and gestalt compensation unit is systematically demonstrated. The processing procedure is divided into three core stages: noise detection, threshold calculation and signal repair, aiming to effectively eliminate the influence of various physiological and non-physiological artifacts on the quality of neural signals. In the noise detection stage, the system first calculates the median and range statistics of the first sliding time window for three types of modal data. When the sum of the range of the three types of modal data in the first sliding time window is less than the preset steady-state threshold limit, the system records the maximum value of the range of the three types of modal data as a zero-level steady-state template, which serves as a reference for subsequent noise discrimination. Subsequently, the system maintains three layers of sliding time windows with increasing length for cross-modal data streams per channel. This multi-level window design can adapt to artifacts of different time scales, with the first layer of windows capturing short burst noise, and the second and third layers of windows detecting long-duration artifact patterns. In the threshold calculation stage, the system updates the median absolute deviation of each layer of sliding windows in real time and generates dynamic thresholds through fusion algorithms. The specific calculation method is as follows: multiply the median absolute deviation of the current layer by a predefined sub-window coefficient, and then add the range proportion of the zero-level steady-state template to obtain the multi-layer fusion threshold. This adaptive threshold mechanism can dynamically adjust the detection sensitivity according to the statistical characteristics of the signal, effectively balancing the noise detection accuracy and false detection rate. When the absolute amplitude of any type of modal data exceeds the corresponding layer fusion threshold and the duration is shorter than the minimum duration of neural events, the system uniformly marks the corresponding time stamp range of the three types of modal data as a first-level artifact segment. Fixed-length normal waveforms are extracted before and after the artifact segment as repair references, and corresponding interpolation algorithms are used for different modal data characteristics: cubic spline interpolation for neural potential fluctuations and local oxygen concentration changes, and bilinear interpolation for head displacement information, finally generating a first-level cleaning matrix.
[0067] Figure 4The system architecture and control implementation mechanism of the digital human interaction control unit are comprehensively described. The unit receives the cross-modal neural representation sequence processed by cross-modal noise suppression and gestalt compensation as input, and finally realizes precise real-time control of the digital human through multi-level analysis and conversion. The digital human interaction control unit adopts a modular design architecture, including four core functional modules: motion analysis, expression control, speech synthesis, and behavior planning. The motion analysis module is responsible for extracting the user's motion intention information from the neural representation sequence, and mapping specific neural activity patterns to corresponding digital human limb motion instructions through pattern recognition algorithms. This module uses machine learning methods to establish a non-linear mapping relationship between neural signal features and action categories, and can identify motion intentions including hands, head, torso, and other parts. The expression control module specifically processes facial expression-related neural signals, and generates corresponding digital human facial expression control parameters by analyzing neural representation features related to emotions and facial muscle activity. This module considers the standardization requirements of the facial action coding system and can accurately control the expression changes of key areas such as eyebrows, eyes, and mouth, achieving natural and smooth facial expression expression. The speech synthesis module converts speech intention-related neural signals into speech output control instructions. This module first extracts speech-related feature parameters from neural representations, including pitch, speech rate, emotional color, etc., and then generates corresponding speech signals through neural network speech synthesis technology to realize the intelligent speech interaction function of the digital human. The behavior planning module is responsible for integrating the outputs of each control module, and performing global behavior coordination and optimization. This module considers the time sequence and spatial constraints of actions to ensure that the behaviors of the digital human are consistent in time and space, avoiding unreasonable action conflicts. The control effect evaluation system monitors various performance indicators of digital human control in real time, including response delay, control accuracy, and stability. Through the establishment of a real-time feedback mechanism, the system can continuously optimize control parameters to ensure the smoothness and accuracy of digital human interaction. The evaluation results show that the system response delay is controlled within 100 milliseconds, the control accuracy is above 95%, and the stability index is rated as excellent, meeting the technical requirements of real-time interaction applications.
[0068] Although the specific embodiments of the present application are described above, those skilled in the art should understand that these specific embodiments are only illustrative, and those skilled in the art can make various omissions, substitutions and changes to the details of the above method and system without departing from the principles and essence of the present application. For example, combining the above method steps, performing substantially the same function in substantially the same manner to achieve substantially the same result according to the same method is within the scope of the present application. Therefore, the scope of the present application is only limited by the appended claims.
Claims
1. A brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence, characterized in that, The system comprises: a multimodal neural signal acquisition unit, a cross-modal noise suppression and gestalt compensation unit, and a digital human interaction control unit. The multimodal neural signal acquisition unit includes an electrode array, a near-infrared optical probe, and a miniature inertial measurement unit deployed at predetermined intervals on the user's scalp and behind the ears. It uses the same master clock to trigger sampling at a frequency to acquire three types of modal data: neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. These three types of modal data are injected into an on-chip circular double buffer in real time, and a unified timestamp is appended during writing to the on-chip circular double buffer to achieve cross-modal time alignment and obtain a cross-modal data stream. The cross-modal noise suppression and gestalt compensation unit performs multi-level artifact suppression on the cross-modal data stream and then performs gestalt compensation based on temporal similarity discrimination to obtain a cross-modal neural representation sequence. The digital human interaction control unit is used to control the digital human based on the cross-modal neural representation sequence. The cross-modal noise suppression and gestalt compensation unit performs multi-level artifact suppression on the cross-modal data stream, including: calculating the median and range within the first sliding time window for each of the three modalities; if the sum of the ranges of the three modalities within the first sliding time window is less than a set steady-state threshold, then the maximum value of the ranges of the current three modalities is recorded as the zero-level steady-state template; maintaining three layers of sliding time windows with increasing magnitude for each channel of the cross-modal data stream, and updating the median absolute deviation of each layer in real time; multiplying the current layer's median absolute deviation by... The sub-window coefficients are then superimposed with the range ratio of the zero-level steady-state template to obtain the fusion threshold of the multi-layer. If the absolute amplitude of the value of any modal data exceeds the fusion threshold of the corresponding layer and the duration is shorter than the minimum duration of the neural event, the timestamp range of the three modal data is uniformly marked as the first-level artifact segment. Fixed-length normal waveforms are extracted before and after the first-level artifact segment, and cubic spline interpolation is performed on the neural potential fluctuations and local blood oxygen concentration changes, and bilinear interpolation is performed on the head displacement information. Then, they are combined into a first-level cleaning matrix. The cross-modal noise suppression and gestalt compensation unit performs gestalt compensation based on temporal similarity discrimination on cross-modal data streams. This process includes: detecting and indexing abnormal sampling gaps or residual regions in the three types of modal data to form a gap list; retrieving candidate segments that meet bidirectional correlation conditions in historical time periods based on the gap context waveform and constructing a candidate set; performing multi-scale dynamic temporal alignment on the candidate segments and filtering alignment candidates whose matching errors meet threshold requirements; performing similarity-weighted fusion on the alignment candidates based on alignment errors and eliminating local abnormal spikes through filtering to obtain interpolated segments; performing amplitude recalibration based on energy vectors on the interpolated segments to achieve cross-modal energy consistency; and performing gradual weighted smoothing on both ends of the interpolated segments and concatenating them to a first-level cleaning matrix to form a second-level cleaning matrix.
2. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 1, characterized in that, The range ratio of the zero-order steady-state template is defined as the ratio of the zero-order steady-state template to the sum of the ranges of the three types of modal data within the first sliding time window.
3. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 2, characterized in that, The method for detecting and indexing abnormal sampling segments or residuals in the three types of modal data to form a gap list includes: traversing the unified time axis of the primary cleaning matrix to detect sampling segments or segments in any type of modal data where the absolute value of the interpolation residual exceeds twice the local median absolute deviation; and recording the starting index for the detected segments. Terminate index Root mean square value of residuals of nerve potential fluctuations Root mean square value of residuals of local blood oxygen concentration changes The root mean square value of the residuals with head displacement information ; Generate gap tuples The gaps are then written into the gap list; if the interval between two consecutive gaps is less than the minimum period of a neural event, they are merged into a single gap entry to prevent fragmented reconstruction from causing boundary tearing.
4. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 3, characterized in that, Based on the gap context waveform, the process of retrieving candidate segments that meet the bidirectional correlation condition in historical time periods and constructing a candidate set includes: for each gap, considering the segments before and after the gap... The context waveform in milliseconds is used as the query sample; the closest one before the gap timestamp. Within the historical period of seconds, according to A sliding window with a step size is used to extract candidate fragments, and positive Pearson correlation coefficients are calculated for each query sample. and inverse Pearson correlation coefficient Only when both the positive and negative Pearson correlation coefficients are above the threshold. Furthermore, the joint correlation error of the three types of modal data Less than the threshold Only when the time is right will the candidate segment be added to the candidate queue and sorted from high to low according to the following formula: If the number of candidate segments is insufficient If you enter a search term, the search window will automatically expand. and lower the threshold to the lower limit Continue until the quantity requirement is met.
5. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 4, characterized in that, The process of performing multi-scale dynamic temporal alignment on candidate segments and selecting alignment candidates whose matching errors meet the threshold requirements includes: ranking the top... Each candidate fragment undergoes multi-scale dynamic time warping within the scale set. The time axis is aligned layer by layer; within each scale layer, the main peak sequence of the head displacement information channel is used as an anchor point constraint, requiring the global step size change rate to not exceed [a certain value]. ; Calculate the total alignment error between the aligned fragment and the query sample across the three modalities. and will The fragments are written into the alignment candidate set; This is the alignment error threshold; if the alignment candidate set is empty, the search window will be automatically expanded. and lower the threshold to the lower limit Then, re-search and construct a candidate set.
6. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 5, characterized in that, The alignment candidate set is normalized according to the following formula to obtain the first... The weighted coefficient vector of each segment : ;in, For the first Alignment error of each segment; perform element-wise weighted summation of three modal data components at each time sample point to synthesize the initial interpolation segment. ;exist In the detection of local spikes, if the spike amplitude exceeds four times the average of the five adjacent samples, it is corrected using the neighborhood Savitzky-Golay filtering method to prevent isolated spikes caused by weight imbalance; the interpolation segments are calculated. Energy vector of millisecond waveform on three types of modal data ; Calculate the front and back ends of the gap context respectively. Energy vector of millisecond waveform on three types of modal data and Constructing an energy transfer matrix ;right Perform channel-by-channel amplitude recalibration to obtain interpolation segments with consistent energy. If the recalibration coefficient of any channel is out of range If the corresponding candidate segment is removed from the alignment candidate set and recalculated, the process continues until all channel coefficients are valid.
7. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 6, characterized in that, exist Cut off from both ends Millisecond gradient window, the gradient window uses a cosine half-wave function ;in, For time, The length of the gradient window; the interpolation segments within the gradient window are then matched with the gap context according to... Perform weighted superposition and output the interpolated segment with smoothed boundaries. ;Will Write back the gap position of the primary cleaning matrix and update it to the secondary cleaning matrix.
Citation Information
Patent Citations
Multi-mode AI glasses vision-electroencephalogram cooperative control method, device and equipment
CN120085760A
Motor imagery recognition method combining electroencephalogram and functional near infrared spectrum
CN120392020A