Brain-control digital human interaction system based on brain-computer interface and artificial intelligence

By combining synchronous sampling and multi-level artifact suppression techniques in the brain-computer interface system with temporal similarity discrimination and multi-layer temporal coding, the signal quality and latency problems of existing systems are solved, and high signal-to-noise ratio cross-modal neural representation sequence reconstruction and stable digital human interaction are achieved.

CN120848740AActive Publication Date: 2025-10-28SICHUAN WUTONG TECH CO LTD

Patent Information

Application Number
CN202511358144.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-10-28
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing brain-computer interface systems have shortcomings in signal quality, spatial resolution, and latency characteristics. Single-modal signals are susceptible to interference, and multimodal fusion lacks collaborative optimization strategies, resulting in inaccurate neural intent decoding and unstable action mapping.

Method used

By synchronously sampling under the same master clock using an electrode array, near-infrared optical probe, and miniature inertial measurement element, and combining multi-level artifact suppression and temporal similarity discrimination, cross-modal neural representation sequences are reconstructed. A multi-layer temporal coding network is then used for real-time intent decoding and action mapping to drive the digital human skeleton and facial expressions.

Benefits of technology

It improves the temporal synchronization and energy consistency of cross-modal signals, eliminates motion artifacts and interpolation faults, enhances the accuracy and robustness of neural intent decoding, supports personalized calibration and real-time feedback, and realizes natural and smooth digital human interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848740A_ABST
    Figure CN120848740A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of brain-computer interfaces, and particularly relates to a brain-control digital human interaction system based on a brain-computer interface and artificial intelligence, and the system comprises a multi-modal neural signal collection unit, a cross-modal noise suppression and completion compensation unit, and a digital human interaction control unit. The multi-modal neural signal acquisition unit adopts the same main clock to trigger frequency sampling so as to respectively obtain three types of modal data, and injects the three types of modal data into the on-chip annular double buffer areas in real time; the cross-modal noise suppression and completion compensation unit is used for carrying out multilevel artifact suppression on a cross-modal data stream and then carrying out completion compensation based on time domain similarity discrimination to obtain a cross-modal neural representation sequence; and the digital human interaction control unit is used for realizing digital human control according to the cross-modal nerve representation sequence. According to the invention, coherent and natural digital human actions can be stably output in a complex interaction scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of brain-computer interface technology, specifically relating to a brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence. Background Technology

[0002] With the rapid development of brain-computer interface (BCI) technology and artificial intelligence algorithms, an increasing number of studies are attempting to control virtual characters or digital humans through real-time decoding of EEG or near-infrared optical signals. Traditional BCI systems are mostly based on single-modal signals, such as EEG signals acquired using scalp electrode arrays, which are mapped to external control commands through specific frequency band power or phase characteristics. Some schemes also use near-infrared optical probes to detect changes in local blood oxygen concentration, indirectly reflecting the intensity of neural activity through hemodynamic responses. However, single-modal approaches have limitations in signal quality, spatial resolution, and temporal delay: EEG signals are susceptible to electromyography (EMG) and power frequency interference, while near-infrared optical signals suffer from sampling lag and are sensitive to head movements. Furthermore, schemes using only inertial measurement units (IMUs) to compensate for motion artifacts often only achieve a certain degree of smoothing in the low-frequency band, lacking targeted suppression of transient artifacts such as frequency EMG and electrode slippage. Existing multimodal fusion systems also mainly focus on simple data splicing or weighted fusion, lacking collaborative optimization strategies for key aspects such as temporal synchronization, energy consistency, and reconstruction of missing regions among the three modalities. Summary of the Invention

[0003] Therefore, the main objective of this invention is to provide a brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence. By synchronously sampling electrode arrays, near-infrared optical probes, and miniature inertial measurement elements at the same master clock frequency, a cross-modal data stream of neural potential fluctuations, local blood oxygen concentration changes, and head displacement information is formed. Multi-level artifact suppression, gestalt compensation based on temporal similarity discrimination, energy consistency recalibration, and gradual weighted smoothing are then sequentially applied to reconstruct a high signal-to-noise ratio and continuous, complete cross-modal neural representation sequence. Based on this, a multi-layer temporal coding network and an adaptive feedback stabilization strategy are used to perform real-time intention decoding and action mapping on the neural representation sequence, driving the digital human skeleton and facial expressions to achieve natural, smooth, and low-latency interaction. This invention significantly improves the temporal synchronization and energy consistency of cross-modal signals, eliminates motion artifacts and interpolation faults, enhances the accuracy and robustness of neural intention decoding, supports personalized calibration and real-time feedback, and can stably output coherent and natural digital human actions in complex interactive scenarios.

[0004] The technical solution adopted in this invention is as follows: A brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence is disclosed. The system includes: a multimodal neural signal acquisition unit, a cross-modal noise suppression and gestalt compensation unit, and a digital human interaction control unit. The multimodal neural signal acquisition unit includes an electrode array, a near-infrared optical probe, and a miniature inertial measurement element arranged at predetermined intervals on the user's scalp and behind the ears. It uses the same master clock to trigger frequency sampling to acquire three types of modal data, including: neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. These three types of modal data are injected into an on-chip circular double buffer in real time, and a unified timestamp is added during writing to the on-chip circular double buffer to achieve cross-modal time alignment and obtain a cross-modal data stream. The cross-modal noise suppression and gestalt compensation unit performs multi-level artifact suppression on the cross-modal data stream and then performs gestalt compensation based on temporal similarity discrimination to obtain a cross-modal neural representation sequence. The digital human interaction control unit is used to control the digital human based on the cross-modal neural representation sequence.

[0005] Furthermore, the cross-modal noise suppression and gestalt compensation unit performs multi-level artifact suppression on the cross-modal data stream, including: calculating the median and range within the first sliding time window for each of the three modal data types; if the sum of the ranges of the three modal data types within the first sliding time window is less than a set steady-state threshold, then the maximum value among the ranges of the current three modal data types is recorded as the zero-level steady-state template; maintaining three layers of sliding time windows with increasing magnitude for each channel of the cross-modal data stream, and updating the median absolute deviation of each layer in real time; using the current layer's median absolute deviation... The difference is multiplied by the sub-window coefficient, and then the range ratio of the zero-level steady-state template is added to obtain the fusion threshold of the multi-layer. If the absolute amplitude of the value of any modal data exceeds the fusion threshold of the corresponding layer and the duration is shorter than the minimum duration of the neural event, the timestamp range of the three modal data is uniformly marked as the first-level artifact segment. Normal waveforms of fixed length are extracted before and after the first-level artifact segment, and cubic spline interpolation is performed on the neural potential fluctuations and local blood oxygen concentration changes, and bilinear interpolation is performed on the head displacement information. Then, they are combined into a first-level cleaning matrix.

[0006] Furthermore, the range ratio of the zero-order steady-state template is defined as the ratio of the zero-order steady-state template to the sum of the ranges of the three types of modal data within the first sliding time window.

[0007] Furthermore, the cross-modal noise suppression and gestalt compensation unit performs gestalt compensation based on temporal similarity discrimination on the cross-modal data stream. This process includes: detecting and indexing abnormal sampling gaps or residual regions in the three types of modal data to form a gap list; retrieving candidate segments that meet the bidirectional correlation condition in historical time periods based on the gap context waveform and constructing a candidate set; performing multi-scale dynamic temporal alignment on the candidate segments and filtering alignment candidates whose matching errors meet the threshold requirements; performing similarity-weighted fusion on the alignment candidates based on the alignment error and eliminating local abnormal spikes through filtering to obtain interpolated segments; performing amplitude recalibration based on energy vectors on the interpolated segments to achieve cross-modal energy consistency; and performing gradual weighted smoothing on both ends of the interpolated segments and concatenating them to the first-level cleaning matrix to form a second-level cleaning matrix.

[0008] Furthermore, the method for detecting and indexing abnormal sampling segments or residuals in the three types of modal data to form a gap list includes: traversing the unified time axis of the primary cleaning matrix to detect sampling segments or segments in any type of modal data where the absolute value of the interpolation residual exceeds twice the local median absolute deviation; and recording the starting index for the detected segments. Terminate index Root mean square value of residuals of nerve potential fluctuations Root mean square value of residuals of local blood oxygen concentration changes The root mean square value of the residuals with head displacement information ; Generate gap tuples The gaps are then written into the gap list; if the interval between two consecutive gaps is less than the minimum period of a neural event, they are merged into a single gap entry to prevent fragmented reconstruction from causing boundary tearing.

[0009] Furthermore, the process of retrieving candidate segments that meet the bidirectional correlation condition in historical time periods based on the gap context waveform and constructing a candidate set includes: for each gap, considering the segments before and after the gap... The context waveform in milliseconds is used as the query sample; the closest one before the gap timestamp. Within the historical period of seconds, according to A sliding window with a step size is used to extract candidate fragments, and positive Pearson correlation coefficients are calculated for each query sample. and inverse Pearson correlation coefficient Only when both the positive and negative Pearson correlation coefficients are above the threshold. Furthermore, the joint correlation error of the three types of modal data Less than the threshold Only when the time is right will the candidate segment be added to the candidate queue and sorted from high to low according to the following formula: If the number of candidate segments is insufficient If you enter a search term, the search window will automatically expand. and lower the threshold To the lower limit Continue until the quantity requirement is met.

[0010] Furthermore, the process of performing multi-scale dynamic temporal alignment on candidate fragments and selecting alignment candidates whose matching errors meet the threshold requirements includes: ranking the top... Each candidate fragment undergoes multi-scale dynamic time warping within the scale set. The time axis is aligned layer by layer; within each scale layer, the main peak sequence of the head displacement information channel is used as an anchor point constraint, requiring the global step size change rate to not exceed [a certain value]. ; Calculate the total alignment error between the aligned fragment and the query sample across the three modalities. , and The fragments are written into the alignment candidate set; This is the alignment error threshold; if the alignment candidate set is empty, the search window will be automatically expanded. and lower the threshold To the lower limit Then, re-search and construct a candidate set.

[0011] Furthermore, the alignment candidate set is normalized according to the following formula to obtain the first... The weighted coefficient vector of each segment : ;in, For the first Alignment error of each segment; perform element-wise weighted summation of three modal data components at each time sample point to synthesize the initial interpolation segment. ;exist In the detection of local spikes, if the spike amplitude exceeds four times the average of the five adjacent samples, it is corrected using the neighborhood Savitzky-Golay filtering method to prevent isolated spikes caused by weight imbalance; the interpolation segments are calculated. Energy vector of millisecond waveform on three types of modal data ; Calculate the front and back ends of the gap context respectively. Energy vector of millisecond waveform on three types of modal data and Constructing an energy transfer matrix ;right Perform channel-by-channel amplitude recalibration to obtain interpolation segments with consistent energy. If the recalibration coefficient of any channel is out of range If the corresponding candidate segment is removed from the alignment candidate set and recalculated, the process continues until all channel coefficients are valid.

[0012] Furthermore, in Cut off from both ends Millisecond gradient window, the gradient window uses a cosine half-wave function ;in, For time, The length of the gradient window; the interpolation segments within the gradient window are then matched with the gap context according to... Perform weighted superposition and output the interpolated segment with smoothed boundaries. ;Will Write back the gap position of the primary cleaning matrix and update it to the secondary cleaning matrix.

[0013] By employing the above technical solutions, this invention achieves the following beneficial effects: By synchronously sampling the electrode array, near-infrared optical probe, and inertial measurement element under the same master clock, and combining multi-level artifact suppression with gestalt compensation based on temporal similarity discrimination, this invention achieves comprehensive cleaning and accurate reconstruction of cross-modal data streams containing neural potential fluctuations, local blood oxygen concentration changes, and head displacement information, resulting in high signal-to-noise ratio and temporal continuity for the system input. Furthermore, a multi-layer temporal coding network is used to perform real-time intent decoding on the cleaned cross-modal neural representation sequence. Adaptive feedback stabilization is used to couple the decoding confidence and digital human motion consistency in a closed loop, ensuring both the detailed restoration of the digital human's skeleton and expressions and keeping the end-to-end latency within an acceptable range. Thanks to energy consistency recalibration and gradual weighted smoothing, this invention effectively eliminates spectral leakage caused by gap splicing, improves the amplitude and phase coherence of the interpolation segment and context, and thus suppresses motion jitter and frame loss in complex interactive scenarios. Furthermore, the system supports personalized calibration during the initial resting state, automatically adjusting thresholds and decoding model parameters based on different users' head shapes and physiological characteristics, enabling rapid adaptation to diverse usage environments. Overall, this invention significantly improves upon existing technologies in terms of accuracy, stability, and real-time performance, providing reliable and efficient technical support for immersive digital human interaction. Attached Figure Description

[0014] Figure 1 A schematic diagram of the system structure of a brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the working principle of the multimodal neural signal acquisition unit provided in an embodiment of the present invention; Figure 3 A flowchart illustrating the multi-level artifact suppression process in the cross-modal noise suppression and gestalt compensation unit provided in this embodiment of the invention; Figure 4 This is a schematic diagram of the system architecture of the digital human interaction control unit provided in an embodiment of the present invention. Detailed Implementation

[0015] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.

[0016] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.

[0017] refer to Figure 1 A brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence is disclosed. The system includes: a multimodal neural signal acquisition unit, a cross-modal noise suppression and gestalt compensation unit, and a digital human interaction control unit. The multimodal neural signal acquisition unit includes an electrode array, a near-infrared optical probe, and a miniature inertial measurement element arranged at predetermined intervals on the user's scalp and behind the ears. It uses the same master clock to trigger frequency sampling to acquire three types of modal data, including: neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. These three types of modal data are injected into an on-chip circular double buffer in real time, and a unified timestamp is added when writing to the on-chip circular double buffer to achieve cross-modal time alignment and obtain a cross-modal data stream. The cross-modal noise suppression and gestalt compensation unit is used to perform multi-level artifact suppression on the cross-modal data stream, and then perform gestalt compensation based on temporal similarity discrimination to obtain a cross-modal neural representation sequence. The digital human interaction control unit is used to realize digital human control based on the cross-modal neural representation sequence.

[0018] Electrode arrays capture microvolt-level transient changes in cortical electrical activity, near-infrared optical probes detect differences in light absorption of hemoglobin in the cortical microcirculation, and miniature inertial measurement elements sense rigid body motion of the skull. Each of these three elements reflects different physical mechanisms of neural excitation, metabolic coupling, and external action, while frequency sampling driven by the same master clock ensures event-level synchronization, enabling subsequent algorithms to understand the brain-blood flow-motor coupling state at the "same moment" from a cross-modal perspective.

[0019] The principle of the cross-modal noise suppression and gestalt compensation unit can be divided into two layers: The first layer is directional artifact suppression, which uses robust statistics combining steady-state templates and sliding time windows. Under model-free assumptions, it first estimates the median level and range of the stable region of the signal baseline, and then detects transient impacts based on the change in instantaneous median absolute deviation. The essence of this is to separate "fast and sharp" non-physiological noise from "slow and wide" real neural waveforms with statistics that are insensitive to anomalies, and to enable the simultaneous elimination of three modes by any single sensor anomaly through cross-modal joint thresholding, thereby avoiding temporal misjudgment. The second layer is directional gestalt compensation, which follows the consistency principle of temporal similarity and metabolic coupling of neural signals: the system first scans the primary cleaning matrix, marks the sampling gaps and residual abnormal regions, and then retrieves candidate segments with high bidirectional correlation in the historical window based on the context waveform. Bidirectional correlation ensures the temporal structure symmetry of candidate segments in forward prediction and backward playback, and also makes the fluctuation of neural potential and the change of local blood oxygen concentration present a consistent phase relationship on both sides of the gap. The purpose of performing multi-scale dynamic time alignment is to find the pairing path that minimizes the overall distortion of the three modes under different time scaling scales, so that under the constraint of head displacement information as the anchor point, the local compression or stretching of the time axis is both kinematically coherent and does not excessively distort the neural and blood flow rhythms.

[0020] Candidate segments with low matching errors are weighted and fused, with the weight proportional to the reciprocal of the alignment error, reflecting the Bayesian principle of "the more similar, the more reliable." After weighted fusion, local filtering is performed to prevent isolated spikes caused by excessive weighting of any segment. The principle of energy vector amplitude recalibration is that any signal interpolation must preserve cross-modal energy conservation. In other words, the energy distribution of the interpolated segment in the three channels of neural potential fluctuations, local blood oxygen concentration changes, and head displacement information should be of the same order of magnitude as the gap context. Here, the energy vector is understood as a time-domain power integral. The energy transfer matrix is ​​used to scale the entire interpolated segment to match the average energy of the context, so that the interpolation neither weakens the amplitude of neural events nor amplifies artifacts. If the recalibration coefficient is too large or too small, it indicates that the candidate segment does not match the gap context in amplitude. The system then backtracks the candidate set for re-selection to ensure physiological rationality. The boundary gradient smoothing uses a cosine half-wave window because this window function transitions naturally at zero derivative, continuously eliminating spectral discontinuities at the junction of the interpolated segment and the original segment, thereby avoiding the introduction of frequency leakage in the frequency domain. The cross-modal neural representation sequence obtained after the above two layers of processing not only retains the coupling information of neuronal firing and hemodynamics, but also eliminates motion artifacts and missing segments, providing a high signal-to-noise ratio, high temporal resolution, and cross-modal energy consistent input for the digital human interaction control unit.

[0021] The digital human interaction control unit (DHU) in the overall system is responsible for the real-time mapping of cross-modal neural representation sequences to digital human action commands. Its implementation process can be summarized into four consecutive stages: initialization and adaptive calibration, real-time intent decoding, action synthesis and distribution, and interaction quality feedback and stabilization. In the initialization and adaptive calibration stage, the data access manager first registers and subscribes to cross-modal neural representation sequences on the on-chip pass-through bus. The system recommends a two-level circular buffer for zero-copy transfer: the first-level buffer's write end corresponds to the output thread of the cross-modal noise suppression and gestalt compensation unit; the second-level buffer's read end is bound to the main processing thread of the DHU. The read and write ends maintain queue synchronization through register-level timestamp comparison, ensuring that the maximum buffer latency does not exceed fifty milliseconds. Subsequently, the calibration task scheduler is started, guiding the user to execute a predefined set of actions (e.g., gaze switching, simple tilting and shaking, clenching and releasing a fist), while simultaneously recording the cross-modal neural representation sequence and the digital human's standard action label pair for subsequent intra-domain fine-tuning of the decoding network. To facilitate rapid convergence, a strategy of freezing the backbone parameters and updating only the biases of higher-order attention and output layers can be adopted. A typical learning rate of one to two times ten to the power of negative four is sufficient to meet the convergence requirement within ten seconds. After calibration, a snapshot of the calibration weights is generated for loading during the real-time inference stage.

[0022] In the real-time intent decoding stage, the cross-modal neural representation sequence first enters the sequence normalization pipeline. The pipeline applies sliding zero-mean correction and percentile pruning to neural potential fluctuations, local blood oxygen concentration changes, and head displacement information, respectively, and then concatenates them along the feature dimension to form a multi-channel tensor. This tensor is then fed into a multi-layer temporal coding network. The recommended structure is a single 3D depthwise separable convolution to achieve millisecond-level local feature extraction, followed by bidirectional gated recurrent units to capture contextual dependencies, and then two layers of lightweight Transformer encoders with relative position encoding are stacked. Finally, attention pooling converges the vector into a driving intent vector. To ensure inference latency is controlled within 20 milliseconds, half-precision computation can be enabled on the deployment end and bound to a specific GPU stream. Once generated, the driving intent vector is immediately fed into the action mapper. The action mapper has a built-in skeletal joint mapping table and an expression shape mapping table, both stored in YAML format. Fields include the target bone node name, rotation quaternion index, and expression blendshape weight index, etc. The motion mapper, based on the semantic components of the driving intent vector, calls the linear projection layer to output joint rotation increments and facial expression weight increments, and adds a temporal smoother to generate frame-to-frame transitions. The smoother employs a first-order Butterworth filter and automatically reduces the cutoff frequency at key motion boundaries to balance response speed and visual continuity.

[0023] The motion synthesis and distribution phase is managed by the motion bus scheduler. After vector quantization, the scheduler compresses the joint rotation increments of each frame into a 16-bit fixed-point format and the expression weight increments into 8-bit unsigned integers, then encapsulates them into a unified motion frame structure. The motion frame structure includes a timestamp, frame number, joint rotation array, expression weight array, and a CRC checksum field. The scheduler pushes the motion frames to the digital human rendering engine via zero-copy shared memory or NamedPipe. If the system needs to be displayed on multiple devices, a WebSocket bridge can be added to convert the motion frames into JSON or Protobuf, package them, and broadcast them to each rendering terminal. To ensure synchronization under network fluctuations, the bridge embeds adaptive bitrate logic: when the end-to-end latency exceeds 200 milliseconds or the packet loss rate exceeds 5%, the motion frame downsampling rate smoothly decreases from 60 frames per second to 30 frames per second, and the decoder is notified to double the size of the shared memory buffer.

[0024] The interaction quality feedback and stabilization phase is evaluated cyclically using two complementary metrics: motion consistency and decoding confidence. Motion consistency is assessed by comparing the joint poses and facial expression weights read back from the rendering thread with the target values ​​in the motion frames using mean squared error. Decoding confidence is derived from the attention weight distribution entropy output by the intent decoding network. The feedback controller performs exponential smoothing on the two metrics to obtain short-term and long-term trends. When the short-term trend deteriorates while the long-term trend remains positive, the system infers transient artifacts and triggers cross-modal noise suppression and gestalt compensation units to improve the cleaning level. If the long-term trend also continues to deteriorate, an adaptive recalibration subprocess is initiated: the rendering refresh thread is frozen, retaining only the data access and decoding threads. The user then performs a few seconds of simple command demonstration, and the system updates the higher-level parameters of the decoding network online. This process does not reset the joint mapping table, thus allowing interaction to resume within 30 seconds. If the decoding confidence drops below 30% and does not recover for five consecutive seconds, the system activates a safe exit logic: the motion bus scheduler immediately sends out zero-increment frames, the digital human slowly resets to a resting posture, and the entire process automatically restarts when the confidence recovers to above 50%.

[0025] The thread organization of the entire digital human interaction control unit is recommended to adopt an event-driven model: the main processing thread is responsible for pulling cross-modal neural representation sequences from the circular buffer and initiating intent decoding; the action bus scheduling thread runs independently, responsible for frame compression and distribution; the feedback control thread periodically reads action execution reports and attention entropy, and updates global control variables. Lightweight messages are passed between threads through lock-free queues, and GPU stream synchronization primitives are used to ensure timing consistency during the core computation phase. In terms of software implementation, it is recommended to deploy the decoding network based on PyTorch or TensorRT. The Skeleton API part can be bound to Unity or Unreal Engine, and the communication layer can use ZeroMQ or ROS topics. In terms of hardware, a single NVIDIA RTX 4060-level GPU can meet the real-time requirements of 60 frames per second and end-to-end latency of less than 120 milliseconds.

[0026] Furthermore, the cross-modal noise suppression and gestalt compensation unit performs multi-level artifact suppression on the cross-modal data stream, including: calculating the median and range within the first sliding time window for each of the three modal data types; if the sum of the ranges of the three modal data types within the first sliding time window is less than a set steady-state threshold, then the maximum value among the ranges of the current three modal data types is recorded as the zero-level steady-state template; maintaining three layers of sliding time windows with increasing magnitude for each channel of the cross-modal data stream, and updating the median absolute deviation of each layer in real time; using the current layer's median absolute deviation... The difference is multiplied by the sub-window coefficient, and then the range ratio of the zero-level steady-state template is added to obtain the fusion threshold of the multi-layer. If the absolute amplitude of the value of any modal data exceeds the fusion threshold of the corresponding layer and the duration is shorter than the minimum duration of the neural event, the timestamp range of the three modal data is uniformly marked as the first-level artifact segment. Normal waveforms of fixed length are extracted before and after the first-level artifact segment, and cubic spline interpolation is performed on the neural potential fluctuations and local blood oxygen concentration changes, and bilinear interpolation is performed on the head displacement information. Then, they are combined into a first-level cleaning matrix.

[0027] After the system starts, the electrode array, near-infrared optical probe, and micro inertial measurement element are synchronously written into the on-chip buffer. The cross-modal data stream is then sent to the pre-buffer of this unit. When the buffer is full and the first sliding time window is full, the algorithm immediately calculates the median and range of the three types of modal data. By comparing the sum of the ranges with the preset steady-state threshold, it quickly determines whether the current sequence is in a physiological resting state. If it is determined to be resting, the maximum value of the range of the three types of modal data is solidified as the zero-level steady-state template. The zero-level steady-state template is subsequently regarded as a scale for cross-modal background noise and participates in threshold fusion in the form of range proportion.

[0028] In the subsequent real-time operation phase, the algorithm maintains three layers of sliding time windows with increasing degrees for each channel. This incremental strategy ensures that short-term anomalies and long-term trends can be captured simultaneously. When each sliding time window expires, the median absolute deviation of the corresponding channel is updated in real time, making the threshold calculation insensitive to instantaneous shocks but adaptive to baseline drift. To unify the local fluctuations within a channel with the cross-modal resting amplitude, the algorithm multiplies the current layer's median absolute deviation by the sub-window coefficient and then superimposes the range ratio of the zero-order steady-state template to form a multi-layered fusion threshold array. This fusion threshold array has global alignment characteristics for the three types of modal data, effectively avoiding false alarms in a single channel due to differences in sensor sensitivity.

[0029] When the absolute amplitude of any modal data exceeds the fusion threshold of the corresponding layer and its duration is shorter than the minimum duration of a neural event, the algorithm considers the impact to lack neurological significance and more likely caused by electrode slippage, optical path obstruction, or momentary head jerking. Therefore, it uniformly marks the synchronous timestamps of the three modal data as a first-level artifact segment. This cross-modal synchronous marking prevents artifact residue caused by channel delay differences. After the first-level artifact segment is marked, the algorithm extracts a fixed length of normal waveform before and after the segment as context for subsequent interpolation and reconstruction.

[0030] During the interpolation stage, the system adheres to the principle of signal prototype continuity: cubic spline interpolation is used for neural potential fluctuations and local blood oxygen concentration changes. This high-order polynomial interpolation possesses second-order derivative continuity in the time domain, maintaining waveform smoothness and avoiding excessive distortion of peaks. Bilinear interpolation is used for head displacement information because the physical meaning of this information is rigid body translation and rotation. Linear interpolation can maintain consistent displacement increments and avoid introducing spurious frequencies in low sampling intervals. After interpolation, the three types of modal data are combined into a first-order cleansing matrix according to the original time axis. The matrix maintains time dimension alignment and channel dimension separation in its structure, which facilitates gap detection in the subsequent gestalt compensation stage and ensures that the signal received by the digital human interaction control unit meets the requirements of minimum latency and maximum integrity. The entire multi-level artifact suppression process, through steady-state template initialization, adaptive updating of incremental sliding windows, fusion threshold generation driven by median absolute deviation, and cross-modal synchronous labeling strategy, enables the system to immediately eliminate most motion artifacts and sensor burst noise without relying on complex models and priors. The subsequent interpolation mechanism restores the broken waveform to a continuous trajectory that conforms to physiological laws through physically consistent signal reconstruction, laying the foundation for the subsequent gestalt compensation unit to further retrieve historical fragments and fill gaps.

[0031] Furthermore, the range ratio of the zero-order steady-state template is defined as the ratio of the zero-order steady-state template to the sum of the ranges of the three types of modal data within the first sliding time window.

[0032] The range ratio of the zero-level steady-state template is defined as the ratio of the zero-level steady-state template to the sum of the ranges of the three modal data within the first sliding time window. Its fundamental purpose is to provide a unified and adaptive amplitude baseline for subsequent multi-layer fusion thresholding. When the cross-modal data stream first enters the cross-modal noise suppression and gestalt compensation unit, it has not undergone any cleaning, and the absolute amplitudes, signal morphology, and noise distributions of the three modal data often differ by orders of magnitude. If the amplitude index of a single channel is directly used in the threshold design, the fusion threshold is easily dominated by the occasional spike of a certain modality, thus reducing the robustness of artifact detection. Conversely, if it is completely averaged, the protection for the most sensitive modality will be weakened, increasing the risk of missed detections. Therefore, the algorithm first independently calculates the range of the three modal data within the first sliding time window and selects the largest one as the zero-level steady-state template to represent the highest background noise amplitude that the system may encounter in the resting state. Meanwhile, to prevent this maximum range from being over-amplified in subsequent threshold fusion, the system compares this maximum range with the sum of the three ranges to obtain the range proportion of the zero-level steady-state template. This proportion naturally falls between zero and one, effectively normalizing the maximum range and making it a dimension-independent, cross-modal comparable relative quantity, rather than an absolute magnitude. Subsequently, when calculating the fusion threshold for multiple layers, the median absolute deviation multiplied by the sub-window coefficient only reflects local fluctuations within a channel and still lacks alignment with the global noise background. By superimposing the range proportion of the zero-level steady-state template, all channels share the same resting amplitude reference, thus ensuring that the threshold size in cross-modal joint judgment is neither overly biased towards high-noise channels nor neglects low-noise channels.

[0033] Furthermore, the cross-modal noise suppression and gestalt compensation unit performs gestalt compensation based on temporal similarity discrimination on the cross-modal data stream. This process includes: detecting and indexing abnormal sampling gaps or residual regions in the three types of modal data to form a gap list; retrieving candidate segments that meet the bidirectional correlation condition in historical time periods based on the gap context waveform and constructing a candidate set; performing multi-scale dynamic temporal alignment on the candidate segments and filtering alignment candidates whose matching errors meet the threshold requirements; performing similarity-weighted fusion on the alignment candidates based on the alignment error and eliminating local abnormal spikes through filtering to obtain interpolated segments; performing amplitude recalibration based on energy vectors on the interpolated segments to achieve cross-modal energy consistency; and performing gradual weighted smoothing on both ends of the interpolated segments and concatenating them to the first-level cleaning matrix to form a second-level cleaning matrix.

[0034] First, sampling gaps or residual anomalies in the three modalities are detected and indexed to form a gap list. The underlying mechanism of this detection is to use a unified time axis obtained in the first-stage cleaning phase to calculate local residuals and local steady-state statistics for each time sampling point. When the absolute value of the residual continuously exceeds a threshold and the duration exceeds the minimum gap threshold, it is identified as a sampling gap; when the residual exhibits an isolated peak and the local energy is significantly higher than the average energy of the context, it is marked as a residual anomaly. Since the three modalities are physically coupled synchronously, the gap index is uniformly written into the gap list through timestamp alignment to ensure that subsequent reconstruction maintains cross-modal consistency. Subsequently, the algorithm uses the gap context waveform as a retrieval template to retrieve candidate segments that meet the bidirectional correlation condition in the historical time period through a sliding window and constructs a candidate set. Bidirectional correlation requires that the candidate segments achieve high correlation coefficients in both forward and backward matching to ensure that their temporal structure can both predict the leading edge of the gap and fit the trailing edge of the gap, thereby preserving the true phase relationship of neural discharge-blood flow coupling to the maximum extent.

[0035] To fully adapt to variations in rhythm, each candidate segment in the candidate set then enters a multi-scale dynamic temporal alignment module. The algorithm performs alignment at multiple temporal scaling scales, using the main peak sequence in the head displacement information as an anchor constraint to limit global step size changes and prevent excessive time stretching. After alignment, the matching error is calculated, and alignment candidate sets with errors less than a set threshold are selected. This step uses error as an indicator to remove segments with large mismatches, ensuring that the reconstruction is based on high-confidence samples. The system then performs similarity-weighted fusion on the alignment candidates based on the alignment error. The similarity weighting follows the principle that the smaller the error, the higher the weight, enabling multiple candidate segments to collaboratively fill gaps in the temporal domain. Simultaneously, neighborhood filtering is used to eliminate potential local abnormal peaks during the fusion process, outputting an interpolated segment.

[0036] To ensure the energy distribution of the interpolated segments across the three modalities remains consistent with the gap context, the system performs amplitude recalibration based on the energy vector. The energy vector is treated as a cross-modal power representation. The recalibration process independently calculates scaling factors for each channel, then checks whether the coefficients fall within a reasonable physiological range. If any exceed the limit, candidate segments are re-selected and the process is repeated to guarantee cross-modal energy consistency. After amplitude correction, the interpolated segments still need seamless splicing with the original waveform. Therefore, the algorithm applies a gradual weighted smoothing process to both ends of the interpolated segments. The gradual window is defined according to the cosine half-wave function, ensuring the continuity of the first derivative at the endpoints, thus avoiding spectral leakage due to amplitude jumps. The processed interpolated segments are written back to the corresponding gap positions in the first-level cleaning matrix, forming a second-level cleaning matrix. Compared to the first-level cleaning matrix, the second-level cleaning matrix is ​​not only continuous and uninterrupted in the time domain, but also consistent with the surrounding waveforms in amplitude, energy, and phase structure, providing a high-completeness and high-signal-to-noise-ratio cross-modal neural representation sequence for subsequent intent decoding in the digital human interaction control unit.

[0037] Furthermore, the method for detecting and indexing abnormal sampling segments or residuals in the three types of modal data to form a gap list includes: traversing the unified time axis of the primary cleaning matrix to detect sampling segments or segments in any type of modal data where the absolute value of the interpolation residual exceeds twice the local median absolute deviation; and recording the starting index for the detected segments. Terminate index Root mean square value of residuals of nerve potential fluctuations Root mean square value of residuals of local blood oxygen concentration changes The root mean square value of the residuals with head displacement information ; Generate gap tuples The gaps are then written into the gap list; if the interval between two consecutive gaps is less than the minimum period of a neural event, they are merged into a single gap entry to prevent fragmented reconstruction from causing boundary tearing.

[0038] The algorithm iterates through the primary cleaning matrix frame by frame, calculates the absolute value of the interpolation residuals of the three modes at each sampling point, and compares it with the local median absolute deviation of the corresponding channel, which is updated in real time using a sliding window method. Comparison, when any channel satisfies And if the interval spans at least two consecutive sampling periods, it is determined to be a potential gap candidate; This represents the absolute value of the sampling gap or interpolation residual. The values ​​are respectively , and If the same interval is also accompanied by a sampling empty segment marker, it is directly classified as a sampling empty segment. After detecting a candidate segment, the system sets its starting index. With Termination Index Record the data and calculate the root mean square value of the residuals for the three types of modal interpolation residuals within the interval. , , ,in The three residual power measures, representing the number of sampling points within the segment, collectively describe the intensity of the gap's influence on neural potential fluctuations, local blood oxygenation changes, and head displacement information. The algorithm then generates gap tuples. The information is written into the gap list. The gap tuple provides both temporal positioning information and cross-modal error quantization, providing necessary input for subsequent candidate fragment retrieval and energy vector magnitude recalibration.

[0039] To avoid signal boundary tearing caused by multiple annotations due to short-term repetitive artifacts, the system continues to check the intervals between adjacent entries in the gap list sequentially along a unified time axis. If the difference between the start index of two entries and the destination end index of the previous entry is found to be less than the minimum period of the neural event, it is considered that these two abnormalities belong to the same physiological or artifact cause, and a gap merging operation is immediately performed: new Take the original value of the former, and the new value The original value of the latter is taken, and the root mean square value of the residual is recalculated by summing the squares of the two residuals and then taking the square root again to maintain energy conservation. This merging strategy can effectively prevent fragmented reconstruction during the candidate retrieval and dynamic time alignment stages, thereby reducing the number of interpolation segment boundaries, reducing the complexity of boundary gradient smoothing, and ensuring that the recalibration matrix calculation for cross-modal energy consistency is based on the statistics of complete segments rather than the average of fragmented segments, thus improving the temporal continuity and amplitude stability during the decoding of brain-controlled digital human actions.

[0040] Furthermore, the process of retrieving candidate segments that meet the bidirectional correlation condition in historical time periods based on the gap context waveform and constructing a candidate set includes: for each gap, considering the segments before and after the gap... The context waveform in milliseconds is used as the query sample; the closest one before the gap timestamp. Within the historical period of seconds, according to A sliding window with a step size is used to extract candidate fragments, and positive Pearson correlation coefficients are calculated for each query sample. and inverse Pearson correlation coefficient Only when both the positive and negative Pearson correlation coefficients are above the threshold. Furthermore, the joint correlation error of the three types of modal data Less than the threshold Only when the time is right will the candidate segment be added to the candidate queue and sorted from high to low according to the following formula: If the number of candidate segments is insufficient If you enter a search term, the search window will automatically expand. and lower the threshold To the lower limit Continue until the quantity requirement is met.

[0041] To ensure the physiological consistency of the reconstructed segment with the original waveform in terms of phase and energy structure, the cross-modal noise suppression and gestalt compensation unit first extracts the phases before and after each notch for each notch. The context waveform in milliseconds is used as the query sample. This context simultaneously preserves the joint temporal features of three modalities: neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. Then, the most recent time before the gap timestamp... Within the historical segment of seconds A sliding window with a fixed step size is used to extract candidate fragments. This strategy balances retrieval speed and coverage density, ensuring sufficient sampling of the historical pattern library with limited computing resources. For each candidate fragment, the system first calculates its positive Pearson correlation coefficient with the query sample on the original time axis. Then, the candidate fragments are time-reversed and aligned with the query samples to calculate the inverse Pearson correlation coefficient. The design principle of bidirectional correlation is that the waveforms of neural discharge and metabolism coupling often exhibit symmetrical characteristics. High positive correlation ensures rhythm matching, while high negative correlation can capture time-reversed symmetrical patterns, thereby improving the measurement accuracy of the overall similarity of waveforms on both sides of the gap.

[0042] The system further performs joint correlation error analysis on the three types of modal data. , It combines the mismatches of positive and negative directions and is a normalized error index for measuring the overall similarity of candidate segments; only when , and Only then does the system determine that the joint temporal pattern of the fragment across the three modalities is highly consistent with the query sample, and thus pushes it into the candidate queue. The candidate queue adopts a... Sorting the keys in descending order of priority, this average relevance metric balances bidirectional matching while also being numerically aligned directly to... This interval facilitates the rapid selection of the top-ranked items during the subsequent dynamic time alignment phase. Perform multi-scale alignment on high-confidence segments. If the number of candidate segments in the queue is insufficient... The system immediately triggers an adaptive window expansion mechanism, expanding the historical search window. Gradually expand to And simultaneously set the relevant thresholds Decrease to the lower bound with a linear step size This progressive search framework ensures that sufficient candidate samples can be obtained even when samples are sparse or the user's specific brainwave patterns are observed, while also... A hard constraint on minimum similarity is set to prevent quality degradation. The entire candidate set construction process relies on four mechanisms—contextual constraints, bidirectional correlation, joint correlation error, and adaptive windowing—to ensure that the selected segments simultaneously satisfy high structural homogeneity across the three channels of neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. This provides a high-confidence starting point for subsequent multi-scale dynamic time alignment and energy vector amplitude recalibration, thereby ensuring the action coherence and signal stability of the brain-controlled digital human interaction link at the system level.

[0043] Furthermore, the process of performing multi-scale dynamic temporal alignment on candidate fragments and selecting alignment candidates whose matching errors meet the threshold requirements includes: ranking the top... Each candidate fragment undergoes multi-scale dynamic time warping within the scale set. The time axis is aligned layer by layer; within each scale layer, the main peak sequence of the head displacement information channel is used as an anchor point constraint, requiring the global step size change rate to not exceed [a certain value]. ; Calculate the total alignment error between the aligned fragment and the query sample across the three modalities. , and The fragments are written into the alignment candidate set; This is the alignment error threshold; if the alignment candidate set is empty, the search window will be automatically expanded. and lower the threshold To the lower limit Then, re-search and construct a candidate set.

[0044] Multi-scale dynamic temporal alignment is designed to maximize the fit to cross-modal rhythms within the gap context while ensuring real-time performance. Therefore, the algorithm first retrieves the top-ranked candidates from the candidate queue. Each candidate fragment is subjected to multi-scale dynamic time warping; multi-scale refers to the system operating across a set of scales. The timeline unfolds layer by layer, with This means the time is compressed to half of its original value. This indicates that the time has been stretched to twice its original value, while This serves as the baseline for proportional alignment. During the normalization process, each scale layer uses the main peak sequence of the head displacement information channels as anchor points. These main peaks typically correspond to the phase extrema of the rigid body motion of the skull and have a clear synchronous relationship with nerve potential fluctuations and local blood oxygen concentration changes in physical coupling. Therefore, selecting them as anchor points can prevent phase drift during cross-modal alignment. To limit waveform distortion caused by excessive stretching or compression of the time axis, the algorithm stipulates that the global step size change rate must not exceed [a certain value]. That is, the difference between the local scaling ratio and the original scaling ratio is limited across all time interpolation segments. Within the range.

[0045] After alignment, the system calculates the alignment residuals for each channel of the three types of modal data and sums them up to obtain the total alignment error. Its calculation method maintains energy consistency with the aforementioned joint correlation error index, giving the error a cross-modal normalization property; when If the phase and amplitude of the candidate fragment match the query sample at the current scale, the fragment is immediately added to the alignment candidate set. If a fragment within the same scale layer has not yet passed the threshold test, it automatically proceeds to the next scale to continue alignment until all scales have been traversed. After completing the three-scale traversal, if the alignment candidate set is still empty, the system determines the history window. Length or similarity threshold The criteria were too strict, so a windowing and re-search mechanism was invoked. Expand by arithmetic increments and Linearly decreasing to the lower limit Then, the entire process of candidate retrieval, bidirectional relevance screening, and multi-scale dynamic time warping is executed again; this progressive threshold-lowering strategy ensures that acceptable candidate segments can still be found even when the user's brainwave morphology changes slowly or sampling noise increases, while also... A minimum relevance cap was set to avoid introducing morphological bias due to excessive relaxation of standards. The fragments finally written into the alignment candidate set are not only homogeneous with the query samples on the time axis, but also meet the threshold requirements in terms of cross-modal coupling strength and energy distribution, laying a high-confidence foundation for subsequent similarity-weighted fusion and energy vector magnitude recalibration. In this way, the reconstructed waveform in the secondary cleaning matrix can maintain physiological rationality while achieving high compatibility with the real-time requirements of the digital human interaction control unit, thus allowing the brain-controlled digital human to maintain a smooth, stable, and latency-controllable interactive experience in complex action sequences.

[0046] Furthermore, the alignment candidate set is normalized according to the following formula to obtain the first... The weighted coefficient vector of each segment : ;in, For the first Alignment error of each segment; perform element-wise weighted summation of three modal data components at each time sample point to synthesize the initial interpolation segment. ;exist In the detection of local spikes, if the spike amplitude exceeds four times the average of the five adjacent samples, it is corrected using the neighborhood Savitzky-Golay filtering method to prevent isolated spikes caused by weight imbalance; the interpolation segments are calculated. Energy vector of millisecond waveform on three types of modal data ; Calculate the front and back ends of the gap context respectively. Energy vector of millisecond waveform on three types of modal data and Constructing an energy transfer matrix ;right Perform channel-by-channel amplitude recalibration to obtain interpolation segments with consistent energy. If the recalibration coefficient of any channel is out of range If the corresponding candidate segment is removed from the alignment candidate set and recalculated, the process continues until all channel coefficients are valid.

[0047] After multi-scale dynamic time warping is completed and an alignment candidate set is obtained, the system needs to fuse several candidate segments with different alignment errors into a continuous and energy-consistent reconstructed waveform. The algorithm first relies on... The normalization formula is used to calculate the first The weighted coefficient vector of each segment, where The total alignment error is taken from the multi-scale dynamic temporal alignment output of the previous step. Normalization is achieved by summing all coefficients by segment dimension and then dividing each by the sum, ensuring that the total weight of all candidate segments is always 1. Simultaneously, segments with smaller errors are implicitly assigned higher weights, directly mapping alignment accuracy to fusion contribution. Subsequently, the system performs element-wise weighted accumulation of the three modal data components at each time sample point on a unified time axis to obtain the initial interpolation segment. Because the small local phase differences between different candidate segments may still cause weight superposition at very few time points, thus introducing isolated amplitude spikes, the system immediately... The internal process performs local spike detection: iterates through each sampling point, and if the current sampling value exceeds four times the average of the five sampling points to its left and right, it is identified as a spike and the neighborhood Savitzky-Golay filter is called to re-estimate the smoothed values ​​of the point and three to five neighboring sampling points. The polynomial least squares property of Savitzky-Golay ensures that the peak value is softened while preserving the higher-order morphological information, avoiding isolated spikes caused by weight imbalance that could disrupt the coupling relationship between nerve potential fluctuations, local blood oxygen concentration changes and head displacement information.

[0048] After the spike correction is completed, the system uses Calculate energy vectors on three types of modal data using milliseconds as the time window. Energy is defined using the sum of squares of samples within a window, normalized by the number of sampling points to ensure comparability of energy across channels. Simultaneously, the same length is taken for both the front and back ends of the notch context. Millisecond waveform, calculate energy vector and Three sets of energy vectors were used to construct the energy transfer matrix. Each diagonal element in the matrix reflects the scaling ratio required for the interpolation segment in the corresponding channel, so that the average energy of the interpolation segment matches the geometric center of the energy on both sides of the gap context; this is achieved channel-by-channel scaling. Multiplying by the scaling factor corresponding to this matrix, the system generates interpolation segments with consistent energy. At this point, the algorithm checks whether the recalibration coefficients of all channels fall within the range of... If any channel coefficient exceeds the boundary, it indicates that the energy distribution of the candidate segment in that channel is too far from the context, which may be caused by abnormal noise or mispairing. Therefore, the corresponding candidate segment is removed from the alignment candidate set, and the weights of the remaining segments are restored and normalized. The process is then repeated until all channel coefficients are valid.

[0049] This "elimination-normalization-recalculation" closed loop ensures that the interpolated segments do not become unbalanced due to a single segment with abnormally high or low energy, thus maintaining the cross-modal energy consistency of neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. The final result... It not only fully inherits the details of the candidate segment with the smallest error in the time and phase domains, but also achieves precise alignment with the context in the energy domain through the transfer matrix. This lays a solid foundation for subsequent gradual weighted smoothing and secondary cleaning matrix splicing, and also ensures that the digital human interaction control unit can receive cross-modal neural representation sequences with stable amplitude and coherent phase when decoding, thereby continuously outputting coherent and natural digital human movements.

[0050] Furthermore, in Cut off from both ends Millisecond gradient window, the gradient window uses a cosine half-wave function ;in, For time, The length of the gradient window; the interpolation segments within the gradient window are then matched with the gap context according to... Perform weighted superposition and output the interpolated segment with smoothed boundaries. ;Will Write back the gap position of the primary cleaning matrix and update it to the secondary cleaning matrix.

[0051] When an interpolation segment with consistent energy is obtained Afterwards, the splicing boundary between the interpolated segment and the gap context still needs to be smoothed to eliminate the discontinuity caused by the slight phase difference between the interpolation and the original waveform. The algorithm in The front and back ends are each cut to a length of 1. A millisecond gradient window is used, and within each gradient window, the interpolation segment samples are processed according to a cosine half-wave function. Weighting is performed; where The time taken from the start of the gradient window, with a range of values. This function is in The time weight is zero, in The time weight is one, and the weight of the outside of the gap context is... Because they are complementary, the two waveforms are cosine-weighted and superimposed linearly within the transition window region, resulting in a smooth transition. The cosine half-wave function has the property that its zeroth and first derivatives are both zero at both ends, ensuring that not only the amplitude but also the slope is continuous at the splicing point, thus avoiding sharp transition bands and frequency leakage in the frequency domain. The output after weighted superposition is the smoothed interpolation segment. The system immediately... Write back to the gap interval of the primary cleaning matrix, replace the original gap labels, and update it with the secondary cleaning matrix. At this point, the secondary cleaning matrix no longer contains any breaks on the time axis and is seamlessly aligned with the surrounding waveforms in terms of amplitude, phase, and energy. This provides the digital human interaction control unit with a highly complete, high signal-to-noise ratio cross-modal neural representation sequence, ensuring that the brain-controlled digital human maintains continuous, stable, and low-latency action output in subsequent interaction scenarios.

[0052] The formula for weighted fusion of candidate fragments is written as follows: ; in This indicates that on a unified timeline, for channels At the moment The values ​​of the initial interpolation segment obtained; For the first The resampled waveform of a candidate segment aligned with the gap after multi-scale dynamic time warping; The total alignment error of the three modalities of the corresponding segment; the smaller the error, the higher the similarity. (Fractional part) Define the weighted coefficient vector Ensure that the sum of the weights of all candidate fragments is 1. and Inversely proportional to alignment accuracy, high-precision segments contribute more to the fusion process; The number of segments that enter the alignment candidate set.

[0053] The formula for energy vector magnitude recalibration and interpolation segment scaling is written as: ; in Indicates the initial interpolation segment in the channel above Energy vector components calculated using milliseconds as the window; and These are the energy components of the long waveform in the front and back ends of the gap context, respectively. The range of values ​​for the channel-by-channel amplitude recalibration coefficient is subject to... Restraints are necessary to ensure physiological rationality; For the final interpolation segment with consistent energy, energy conservation between the interpolation segment and the context is achieved across the cross-modal scale. The first formula adaptively fuses the phase-amplitude information of multiple segments according to the error, and the second formula precisely aligns the fusion result with the context in the energy domain. The system can output a cross-modal neural representation sequence that is both phase-continuous and energy-balanced, thereby meeting the dual requirements of temporal integrity and amplitude consistency of the brain-controlled digital human interaction link.

[0054] Suppose the electrode array worn by the user contains The near-infrared optical probe contains one electrode. Transmitter-receiver pair, miniature inertial measurement element output Axial acceleration and axial angular velocity, total One channel; master clock sampling rate set to [value]. The length of the first sliding time window after system startup is taken as... (Right now (Sampling points). Within this window, the ranges of the three modalities are respectively , (Relative changes in near-infrared light absorption) The sum of the ranges is Set steady-state threshold ,because The system confirms that it is in a steady state and will Let it be denoted as the zero-order steady-state template; its range ratio is After entering real-time operation, settings are configured for each channel. Layered sliding window: According to EEG number Taking the channel as an example, during a certain period of time Time corresponding The absolute deviation of the median within the window is Sub-window coefficients are taken The fusion threshold is At this point, the absolute amplitude of the channel was detected to be present. Crossing but continuing Minimum duration of neural events Meanwhile, neither the fNIRS nor the IMU channels exceeded the threshold, so the three modes were uniformly labeled. This is a first-level artifact section. Samples were taken from both the front and back. Normal waveforms are processed by cubic spline interpolation of EEG and fNIRS, and bilinear interpolation of IMU, generating a primary cleaning matrix. During continued scanning, ... arrive The absolute value of EEG plugging residuals was found. And accompanied by fNIRS sampling gaps, recording , , Generate gap tuples Then, with respect to both sides of the gap. The waveform is the query sample, in the most recent Within the window of history Step size sliding, to obtain candidate segments Item. Calculate the positive correlation coefficient. With inverse correlation coefficient After that, only Conditions satisfy and ,according to Sort by first The data enters a multi-scale dynamic time warping process.

[0055] Taking the first segment as an example, its scale Top alignment error ,scale With scale The errors are respectively and Select the record with the minimum error. The final alignment of the candidate set is then retained. The segment, its error vector .according to Obtain normalized weights .right Interpolation interval and three modes Sample-by-sample weighted summation Detect spikes in S_raw, only in the EEG channel. Peak was found at the location Exceeding four times the neighborhood mean, applying a Savitzky-Golay third-order window length. Sample filtering correction Then take Calculate the energy vector (The units are respectively) , variance, The average of the gap context energy vectors is as follows: Diagonal elements of the energy transfer matrix All fell into Therefore, it can be recalibrated directly: .right Take from both the front and back Gradient window, press Overlay with the gap context to obtain a smoothed result. .Will Write back to the primary cleaning matrix After the segment, the secondary cleaning matrix is ​​completed. The digital human interaction control unit checks the phase consistency between this segment and the context EEG-fNIRS-IMU in real time, improving it to... Overall end-to-end latency is controlled within This meets the needs of smooth interaction for brain-controlled digital humans.

[0056] After completing the secondary cleaning matrix, the system... Perform a window aggregation on the cross-modal neural representation sequences, and then... Lu EEG, fNIRS and The IMU channel values ​​are normalized by channel and then concatenated to obtain the length. The instantaneous vector is then appended with a second-order difference to obtain the total dimension. Then connect them together The frame length is The timing block input to the digital human interaction control unit is a first-layer bidirectional gated loop network; the hidden dimension of this network is set. Because the time step is Therefore, the output dimensions of the forward and reverse hidden state tensors are... The system obtains the hidden state vector through time-dimensional max pooling. Then, through linear transformation Will Projected to Dimensional driving intent space, in which , The driving intent vector is then... generate The joint angle increment, , , For hyperbolic tangent. The joint mapping table specifies the first... Dimension mapping to the head quaternion, Mapped to the cervical segment, Corresponding to the degrees of freedom of the two arms, torso, and legs; for example: in Get it at all times The digital human's head immediately rotates slightly around three axes, with the motion frames in... Rate packaged as Push to the rendering end. To suppress micro-judder, the motion bus scheduler performs a first-order Butterworth filter on consecutive frame increments, with a cutoff frequency set to... The system is in Frame consistency error detected less than the threshold Decoding confidence Therefore, the current decoding weights are maintained. If persistent [issues] occur at the joints of both arms... low amplitude fluctuations and The feedback control thread immediately raises the cross-modal noise suppression level to the sliding window. and put Temporary increase To tighten the candidate relevance. The duration of a complete interaction loop from the generation of cross-modal neural representation sequences to the output of motion frames at the digital human rendering end. The cleaning and reconstruction process is time-consuming. Intent decoding and action mapping time Frame compression and transmission time The remainder is GPU synchronization and rendering overhead; continuous operation The mean of overall motion consistency was measured later. Standard deviation End-to-end delay drift This demonstrates that the specific numerical scheme can stably support smooth interaction of brain-controlled digital humans under real-world hardware configurations.

[0057] Figure 2 This paper details the working principle and data acquisition process of the multimodal neural signal acquisition unit of this invention. The acquisition unit employs a distributed sensor array configuration, deploying electrode arrays, near-infrared optical probes, and miniature inertial measurement units at predetermined intervals on the user's scalp and behind the ears to achieve simultaneous acquisition of three different modalities of neural signals. Specifically, the electrode array is responsible for acquiring neural potential fluctuation signals, which reflect changes in the electrical activity of neuronal populations in the cerebral cortex. These signals have high temporal resolution, with sampling frequencies typically exceeding 500 Hz. The near-infrared optical probe monitors changes in local blood oxygen concentration, indirectly reflecting local brain metabolic activity by detecting differences in the spectral absorption of hemoglobin and deoxyhemoglobin. This signal has good spatial resolution but a relatively slow temporal response. The miniature inertial measurement unit primarily acquires head displacement information, including triaxial acceleration and angular velocity data, used to monitor and compensate for motion artifacts caused by head movements to the other two types of signals. The system employs a unified master clock triggering mechanism to ensure synchronous sampling of the three modalities of data. All acquired raw data is injected in real-time into an on-chip circular dual buffer. This buffer employs a ping-pong caching architecture, automatically attaching a uniform timestamp during data writing to achieve precise time alignment across modal data. This design effectively solves the time synchronization problem caused by inherent latency differences between different sensors, laying a reliable time benchmark for subsequent cross-modal data fusion and analysis. The capacity design of the circular dual buffer fully considers real-time processing requirements and hardware resource constraints, minimizing system latency while ensuring data integrity.

[0058] Figure 3This paper systematically demonstrates the complete processing flow of multi-level artifact suppression in a cross-modal noise suppression and gestalt compensation unit. The process is divided into three core stages: noise detection, threshold calculation, and signal restoration, aiming to effectively eliminate the impact of various physiological and non-physiological artifacts on neural signal quality. In the noise detection stage, the system first calculates the median and range statistics within the first sliding time window for each of the three modalities. When the sum of the ranges of the three modalities within the first sliding time window is less than a preset steady-state threshold, the system records the maximum value of the current ranges of the three modalities as the zero-level steady-state template, which serves as the benchmark reference for subsequent noise discrimination. Subsequently, the system maintains three layers of progressively increasing sliding time windows for each channel of the cross-modal data stream. This multi-layered window design can adapt to artifact characteristics at different time scales. The first layer captures short-term burst noise, while the second and third layers are used to detect artifact patterns with longer durations. In the threshold calculation stage, the system updates the median absolute deviation of each sliding window in real time and generates a dynamic threshold through a fusion algorithm. The specific calculation method is as follows: multiply the absolute deviation of the current layer median by a predefined sub-window coefficient, and then superimpose the range ratio of the zero-level steady-state template to obtain the multi-layer fusion threshold. This adaptive threshold mechanism can dynamically adjust the detection sensitivity according to the statistical characteristics of the signal, effectively balancing the accuracy of noise detection and the false detection rate. When the absolute amplitude of any type of modal data exceeds the corresponding layer fusion threshold and the duration is shorter than the minimum duration of the neural event, the system uniformly marks the corresponding timestamp range of the three types of modal data as a first-level artifact segment. Normal waveforms of fixed length are extracted before and after the artifact segment as repair references, and corresponding interpolation algorithms are used for the characteristics of different modal data: cubic spline interpolation is performed on neural potential fluctuations and local blood oxygen concentration changes, and bilinear interpolation is performed on head displacement information, finally splicing to generate a first-level cleaning matrix.

[0059] Figure 4This paper comprehensively describes the system architecture and control implementation mechanism of the digital human interaction control unit. This unit receives cross-modal neural representation sequences, processed with cross-modal noise suppression and gestalt compensation, as input. Through multi-level parsing and transformation, it ultimately achieves precise real-time control of the digital human. The digital human interaction control unit adopts a modular design architecture, comprising four core functional modules: motion parsing, facial expression control, speech synthesis, and behavior planning. The motion parsing module is responsible for extracting the user's movement intention information from the neural representation sequence, mapping specific neural activity patterns to corresponding digital human limb movement commands through pattern recognition algorithms. This module uses machine learning methods to establish a non-linear mapping relationship between neural signal features and movement categories, enabling the recognition of movement intentions from multiple parts, including the hands, head, and torso. The facial expression control module specifically processes neural signals related to facial expressions, generating corresponding digital human facial expression control parameters by analyzing neural representation features related to emotions and facial muscle activity. This module considers the standardization requirements of facial motion coding systems, enabling precise control of facial expression changes in key areas such as eyebrows, eyes, and mouth, achieving natural and fluent facial expression expression. The speech synthesis module converts neural signals related to speech intentions into speech output control commands. This module first extracts speech-related feature parameters from neural representations, including pitch, speech rate, and emotional tone. Then, it generates corresponding speech signals using neural network speech synthesis technology to achieve intelligent voice interaction for the digital human. The behavior planning module integrates the outputs of various control modules, coordinating and optimizing global behavior. This module considers the temporal sequence and spatial constraints of actions, ensuring that the digital human's various behaviors are coordinated in time and space, avoiding unreasonable action conflicts. The control effect evaluation system monitors various performance indicators of the digital human's control in real time, including response latency, control accuracy, and stability. By establishing a real-time feedback mechanism, the system can continuously optimize control parameters, ensuring the smoothness and accuracy of the digital human's interaction. Evaluation results show that the system response latency is controlled within 100 milliseconds, control accuracy reaches over 95%, and stability is rated as excellent, meeting the technical requirements for real-time interactive applications.

[0060] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Those skilled in the art can omit, substitute, and modify the details of the above methods and systems in various ways without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result according to substantially the same method falls within the scope of the present invention. Therefore, the scope of the present invention is defined only by the appended claims.

Claims

1. A brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence, characterized in that, The system includes: a multimodal neural signal acquisition unit, a cross-modal noise suppression and gestalt compensation unit, and a digital human interaction control unit. The multimodal neural signal acquisition unit comprises an electrode array, a near-infrared optical probe, and a miniature inertial measurement unit arranged at predetermined intervals on the user's scalp and behind the ears. It uses the same master clock to trigger sampling at a specific frequency to acquire three types of modal data: neural potential fluctuations, local blood oxygen concentration changes, and head displacement information. These three types of modal data are injected into an on-chip circular double buffer in real time, and a unified timestamp is appended during writing to the on-chip circular double buffer to achieve cross-modal time alignment and obtain a cross-modal data stream. The cross-modal noise suppression and gestalt compensation unit performs multi-level artifact suppression on the cross-modal data stream and then performs gestalt compensation based on temporal similarity discrimination to obtain a cross-modal neural representation sequence. The digital human interaction control unit is used to control the digital human based on the cross-modal neural representation sequence.

2. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 1, characterized in that, The cross-modal noise suppression and gestalt compensation unit performs multi-level artifact suppression on the cross-modal data stream, including: calculating the median and range within the first sliding time window for each of the three modalities; if the sum of the ranges of the three modalities within the first sliding time window is less than a set steady-state threshold, then the maximum value of the ranges of the current three modalities is recorded as the zero-level steady-state template; maintaining three layers of sliding time windows with increasing magnitude for each channel of the cross-modal data stream, and updating the median absolute deviation of each layer in real time; multiplying the current layer's median absolute deviation by... The sub-window coefficients are then superimposed with the range ratio of the zero-level steady-state template to obtain the fusion threshold of multiple layers. If the absolute amplitude of the value of any modal data exceeds the fusion threshold of the corresponding layer and the duration is shorter than the minimum duration of the neural event, the timestamp range of the three modal data is uniformly marked as the first-level artifact segment. Fixed-length normal waveforms are extracted before and after the first-level artifact segment, and cubic spline interpolation is performed on the neural potential fluctuations and local blood oxygen concentration changes, and bilinear interpolation is performed on the head displacement information. Then, they are combined into a first-level cleaning matrix.

3. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 2, characterized in that, The range ratio of the zero-order steady-state template is defined as the ratio of the zero-order steady-state template to the sum of the ranges of the three types of modal data within the first sliding time window.

4. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 3, characterized in that, The cross-modal noise suppression and gestalt compensation unit performs gestalt compensation based on temporal similarity discrimination on cross-modal data streams. This process includes: detecting and indexing abnormal sampling gaps or residual regions in the three types of modal data to form a gap list; retrieving candidate segments that meet bidirectional correlation conditions in historical time periods based on the gap context waveform and constructing a candidate set; performing multi-scale dynamic temporal alignment on the candidate segments and filtering alignment candidates whose matching errors meet threshold requirements; performing similarity-weighted fusion on the alignment candidates based on alignment errors and eliminating local abnormal spikes through filtering to obtain interpolated segments; performing amplitude recalibration based on energy vectors on the interpolated segments to achieve cross-modal energy consistency; and performing gradual weighted smoothing on both ends of the interpolated segments and concatenating them to a first-level cleaning matrix to form a second-level cleaning matrix.

5. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 4, characterized in that, The method for detecting and indexing abnormal sampling segments or residuals in the three types of modal data to form a gap list includes: traversing the unified time axis of the primary cleaning matrix to detect sampling segments or segments in any type of modal data where the absolute value of the interpolation residual exceeds twice the local median absolute deviation; and recording the starting index for the detected segments. Terminate index Root mean square value of residuals of nerve potential fluctuations Root mean square value of residuals of local blood oxygen concentration changes The root mean square value of the residuals with head displacement information ; Generate gap tuples The gaps are then written into the gap list; if the interval between two consecutive gaps is less than the minimum period of a neural event, they are merged into a single gap entry to prevent fragmented reconstruction from causing boundary tearing.

6. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 5, characterized in that, Based on the gap context waveform, the process of retrieving candidate segments that meet the bidirectional correlation condition in historical time periods and constructing a candidate set includes: for each gap, considering the segments before and after the gap... The context waveform in milliseconds is used as the query sample; the closest one before the gap timestamp. Within the historical period of seconds, according to A sliding window with a step size is used to extract candidate fragments, and positive Pearson correlation coefficients are calculated for each query sample. and inverse Pearson correlation coefficient Only when both the positive and negative Pearson correlation coefficients are above the threshold. Furthermore, the joint correlation error of the three types of modal data Less than the threshold Only when the time is right will the candidate segment be added to the candidate queue and sorted from high to low according to the following formula: If the number of candidate segments is insufficient If you enter a search term, the search window will automatically expand. and lower the threshold To the lower limit Continue until the quantity requirement is met.

7. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 6, characterized in that, The process of performing multi-scale dynamic temporal alignment on candidate segments and selecting alignment candidates whose matching errors meet the threshold requirements includes: ranking the top... Each candidate fragment undergoes multi-scale dynamic time warping within the scale set. The time axis is aligned layer by layer; within each scale layer, the main peak sequence of the head displacement information channel is used as an anchor point constraint, requiring the global step size change rate to not exceed [a certain value]. ; Calculate the total alignment error between the aligned fragment and the query sample across the three modalities. , and The fragments are written into the alignment candidate set; This is the alignment error threshold; if the alignment candidate set is empty, the search window will be automatically expanded. and lower the threshold To the lower limit Then, re-search and construct a candidate set.

8. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 7, characterized in that, The alignment candidate set is normalized according to the following formula to obtain the first... The weighted coefficient vector of each segment : ;in, For the first Alignment error of each segment; perform element-wise weighted summation of three modal data components at each time sample point to synthesize the initial interpolation segment. ;exist In the detection of local spikes, if the spike amplitude exceeds four times the average of the five adjacent samples, it is corrected using the neighborhood Savitzky-Golay filtering method to prevent isolated spikes caused by weight imbalance; the interpolation segments are calculated. Energy vector of millisecond waveform on three types of modal data ; Calculate the front and back ends of the gap context respectively. Energy vector of millisecond waveform on three types of modal data and Constructing an energy transfer matrix ;right Perform channel-by-channel amplitude recalibration to obtain interpolation segments with consistent energy. If the recalibration coefficient of any channel is out of range If the corresponding candidate segment is removed from the alignment candidate set and recalculated, the process continues until all channel coefficients are valid.

9. The brain-controlled digital human interaction system based on brain-computer interface and artificial intelligence as described in claim 8, characterized in that, exist Cut off from both ends Millisecond gradient window, the gradient window uses a cosine half-wave function ;in, For time, The length of the gradient window; the interpolation segments within the gradient window are then matched with the gap context according to... Perform weighted superposition and output the interpolated segment with smoothed boundaries. ;Will Write back the gap position of the primary cleaning matrix and update it to the secondary cleaning matrix.

Citation Information

Patent Citations

  • Virtual digital human generation and interaction optimization system based on multi-modal data fusion

    CN119888027A

  • Multi-mode AI glasses vision-electroencephalogram cooperative control method, device and equipment

    CN120085760A

  • Motor imagery recognition method combining electroencephalogram and functional near infrared spectrum

    CN120392020A

  • Multi-modal electroencephalogram classification method under multi-source interference based on multi-head attention mechanism

    CN120429713A

  • Motion state monitoring and feedback system and method based on electroencephalogram signals

    CN120570598A

Cited By

  • Cable production data acquisition method and system

    CN121029868A

  • Glasses end content issuing and local confirmation method in private network environment

    CN121692079A

  • Glasses end content delivery and local confirmation method in private network environment

    CN121692079B

  • Multi-modal neural signal synchronous acquisition method and system for Chinese speech brain-computer interface, electronic equipment and storage medium

    CN121879584A

  • Multi-modal neural signal synchronous acquisition method and system for Chinese speech brain-computer interface, electronic device and storage medium

    CN121879584B