Brain-computer AI headset-based EEG-speech hybrid interaction method, device and equipment
Through the hybrid processing of EEG and speech signals, high signal fusion and interaction accuracy of high signal-to-noise ratio are achieved, the problems of signal synchronization and feature distribution drift in the prior art are solved, and the real-time adaptability and adaptability of the hybrid interactive system are enhanced.
Patent Information
- Application Number
- CN202510556249.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art is difficult to achieve precise synchronous acquisition and alignment of EEG and voice signals, and dynamic changes in user cognitive state and environmental conditions lead to drift and loss of feature distribution, affecting the real-time adaptation and online optimization of hybrid interactive systems.
By collecting dual-modal signals of AI headsets, data classification and processing are performed, timing information is output, noise reduction and signal separation is performed, the mapping chain is built to complete modal conversion, identify key timing feature sets, generate association modes, perform path analysis and hierarchical fusion, supplement missing parts, analyze interaction intentions, and generate interaction commands.
It realizes the high signal-to-noise ratio fusion of EEG and voice signals, improves interaction accuracy and robustness, adapts to complex scenarios, and enhances the real-time and adaptability of interaction.
Smart Images

Figure CN120086743B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of brain-computer interface technology, and in particular to a method, device and equipment for hybrid EEG-speech interaction based on brain-computer AI headphones. Background Art
[0002] Brain-computer interfaces (BCIs) and voice interaction, two important paradigms of human-computer interaction, each possesses distinct advantages but also limitations. While voice interaction is natural and intuitive, it can expose privacy when used in public and is susceptible to interference from ambient noise. BCIs enable silent interaction and directly decode user intent, but they suffer from low information transmission rates and are susceptible to physiological interference such as electromyography and electrooculography. Deeply integrating these two interaction methods promises to complement each other's strengths and overcome their respective limitations.
[0003] However, due to significant differences in sampling frequency, signal-to-noise ratio, and time scale between EEG and speech signals, existing systems struggle to achieve accurate dual-modal signal synchronization and alignment. Traditional solutions often rely on independent processing followed by feature splicing, failing to fully exploit the temporal correlations and complementarities between signals. In practical applications, dynamic changes in user cognitive states and environmental conditions not only cause feature distribution drift but also frequently lead to feature loss, posing significant challenges to the real-time adaptation and online optimization of hybrid interactive systems. Furthermore, accurately understanding user intent and generating adaptive interaction strategies has become a key issue hindering the practical application of the system.
[0004] Therefore, a method is urgently needed to solve at least one of the above problems. Summary of the Invention
[0005] The embodiments of the present application provide a method, device and equipment for hybrid EEG-speech interaction based on brain-computer AI headphones. The method aims to solve the problem that the existing system is difficult to achieve accurate dual-modal signal synchronous acquisition and alignment due to significant differences in sampling frequency, signal-to-noise ratio and time scale between EEG and speech signals. Traditional solutions mostly adopt the method of feature splicing after independent processing, which fails to fully explore the temporal correlation and complementarity between signals. In actual applications, the dynamic changes in user cognitive state and environmental conditions not only cause feature distribution to drift, but also often cause feature missing, which brings huge challenges to the real-time adaptation and online optimization of hybrid interactive systems. At the same time, how to accurately understand user intentions and generate adaptive interaction strategies has also become a key issue restricting the practical application of the system.
[0006] In a first aspect, an embodiment of the present application provides an EEG-speech hybrid interaction method based on a brain-computer AI headset, comprising:
[0007] Collect the bimodal signal of the AI headset, classify and process the bimodal signal, and output time series information; perform noise reduction and signal separation on the time series information, optimize and purify the processing results corresponding to the separation process, and output a pure waveform;
[0008] Performing a two-way analysis on the clean waveform and dividing the feature dimensions, completing modal conversion through the constructed mapping chain, and outputting feature identifiers; performing a time series analysis on the feature identifiers to identify a set of key time series features, establishing an aligned feature mapping structure to combine modalities, and generating a correlation pattern;
[0009] Perform path analysis on the association pattern and mark the combined nodes, reconstruct the features through the transformation chain and output the representation structure; perform hierarchical fusion on the representation structure and extract the fusion area, and output the fusion element based on the fusion area; supplement the missing elements and locate the missing parts, reconstruct the data of the missing parts and output the complete content;
[0010] The complete content is intentionally parsed to obtain key elements, the key elements are converted into interaction information and the interaction intention is output; the interaction intention is information integrated to obtain conversion rules, and interaction commands are generated according to the conversion rules to complete the EEG-speech hybrid interaction based on brain-computer AI headphones.
[0011] In a second aspect, the present application further provides an EEG-speech hybrid interaction device, comprising:
[0012] A waveform output unit is used to collect the dual-modal signals of the AI headset, classify and process the dual-modal signals, and output timing information; perform noise reduction and signal separation on the timing information, optimize and purify the processing results corresponding to the separation process, and output a pure waveform;
[0013] A pattern generation unit is configured to perform a two-way analysis on the clean waveform and divide the feature dimensions, complete the modal conversion through the constructed mapping chain, and output a feature identifier; perform a time series analysis on the feature identifier to identify a set of key time series features, establish an aligned feature mapping structure to combine the modalities, and generate a correlation pattern;
[0014] a node labeling unit configured to perform path analysis on the association pattern and label the combined nodes, reconstruct features through a transformation chain and output a representation structure; hierarchically fuse the representation structure and extract fusion regions, and output fusion elements based on the fusion regions; supplement the fusion elements and locate the missing parts, reconstruct the missing parts and output the complete content;
[0015] An element acquisition unit is used to perform intent analysis on the complete content to obtain key elements, convert the key elements into interaction information and output interaction intentions; integrate information on the interaction intentions to obtain conversion rules, generate interaction commands according to the conversion rules, and complete EEG-speech hybrid interaction based on brain-computer AI headphones.
[0016] This method collects electroencephalogram (EEG) and speech signals (bimodal signals), classifies and processes the data, and outputs time series information (including timestamps and signal quality indicators). It then performs noise reduction (e.g., adaptive filtering, wavelet threshold denoising) and signal separation (e.g., independent component analysis, blind source separation) on the time series information to produce clean waveforms. Time-frequency analysis (e.g., wavelet packet decomposition, Mel-frequency cepstral coefficient extraction) is then performed on the EEG and speech signals, dividing the feature dimensions. Modal conversion is performed through a mapping chain, and feature identifiers are output. Key time series feature sets are identified, an aligned feature mapping structure is established, and EEG-speech modalities are combined to generate association patterns. Path analysis (e.g., dynamic programming algorithm, graph attention network) is performed on the association patterns, and combined nodes are labeled. Feature reconstruction is then performed to output a representation structure. Hierarchical fusion is used to extract fusion elements, supplement missing data, and output complete content. Key elements within the complete content are analyzed (e.g., semantic network analysis, fuzzy reasoning), interaction intents are generated, and interaction commands (e.g., voice commands, EEG control commands) are generated after integration.
[0017] By hybrid processing of EEG and speech signals, complementary signal characteristics (high intent accuracy in EEG and semantic richness in speech) are leveraged to enhance interaction accuracy and robustness. Techniques such as adaptive filtering and blind source separation effectively eliminate environmental noise and physiological artifacts, preserving key waveform features (such as event-related potentials and speech formants), and improving the signal-to-noise ratio. A time-aligned mapping chain (such as hierarchical feature selection and tensor decomposition) is constructed to achieve cross-modal correlation of EEG and speech features, resolving the ambiguity inherent in unimodal intent resolution. Reinforcement learning is combined to optimize intent generation strategies, dynamically adapting to complex scenarios (such as noisy environments and user fatigue), and enhancing the real-time and adaptability of interactions. This approach has applications in medical rehabilitation (language assistance for stroke patients) and smart wearables (AR / VR interactive control), expanding the practical value of brain-computer interfaces.
[0018] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of a method for hybrid EEG-speech interaction based on a brain-computer AI headset according to an embodiment of the present application;
[0020] Figure 2 This is a schematic diagram of the structure of the EEG-speech hybrid interaction device shown in an embodiment of the present application;
[0021] Figure 3 This is a schematic diagram of the structure of a computer device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0023] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0024] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0025] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0026] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0027] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0028] The technical solutions of the embodiments of this application are introduced below.
[0029] Brain-computer interfaces (BCIs) and voice interaction, two important paradigms of human-computer interaction, each possesses distinct advantages but also limitations. While voice interaction is natural and intuitive, it can expose privacy when used in public and is susceptible to interference from ambient noise. BCIs enable silent interaction and directly decode user intent, but they suffer from low information transmission rates and are susceptible to physiological interference such as electromyography and electrooculography. Deeply integrating these two interaction methods promises to complement each other's strengths and overcome their respective limitations.
[0030] However, achieving deep fusion of EEG and speech signals faces many technical challenges. Due to significant differences in sampling frequency, signal-to-noise ratio, and time scale between EEG and speech signals, existing systems find it difficult to achieve accurate dual-modal signal synchronous acquisition and alignment. Traditional solutions often use feature splicing after independent processing, which fails to fully explore the temporal correlation and complementarity between signals. In practical applications, the dynamic changes in user cognitive state and environmental conditions not only cause feature distribution drift, but also often lead to feature missing, which poses a huge challenge to the real-time adaptation and online optimization of hybrid interactive systems. At the same time, how to accurately understand user intentions and generate adaptive interaction strategies has also become a key issue restricting the practical application of the system.
[0031] Please refer to Figure 1 , Figure 1 The flowchart of a method for hybrid EEG-speech interaction based on brain-computer AI headset provided in the embodiment of the present application is shown in FIG. The method for hybrid EEG-speech interaction based on brain-computer AI headset provided in the embodiment of the present application can be applied to computer devices, including but not limited to smart phones, laptops, tablet computers, desktop computers, physical servers, cloud servers and other devices. Figure 1 As shown, the EEG-speech hybrid interaction method based on the brain-computer AI headset of this embodiment includes steps S101 to S104, which are detailed as follows:
[0032] In step S101, the dual-modal signal of the AI headset is collected, the dual-modal signal is classified and processed, and the timing information is output; the timing information is subjected to noise reduction and signal separation processing, and the processing result corresponding to the separation processing is optimized and purified to output a pure waveform.
[0033] Specifically, by first collecting the dual-modal signals of the AI headset and locating the time reference.
[0034] In some embodiments, the bimodal signal is subjected to data classification and processing, and timing information is output, including: synchronously collecting EEG signals and voice signals through a high-precision sensor array to form the bimodal signal; establishing a time reference point based on a GPS timing module, and using an FPGA clock synchronization circuit to achieve time alignment of the bimodal signal; segmenting the voice signal, and using an adaptive threshold mechanism to divide the voice activity area and the non-speech area; preprocessing the EEG signal through a sliding window method, and combining it with Hjorth parameters for feature characterization; classifying the bimodal signal into pure speech segments, pure EEG segments, bimodal activity segments, and background noise segments, for constructing the timing information including timestamps and signal quality indicators.
[0035] In practice, the AI headset uses a high-precision sensor array to synchronously collect the user's EEG and voice signals. The EEG sampling frequency is set at 1000Hz to accurately capture rapid changes in brain neural activity. The signal bandwidth ranges from 0.1 to 100Hz, primarily encompassing the delta, theta, alpha, beta, and gamma frequency bands. The voice signal is sampled at 16kHz with 24-bit quantization to ensure full preservation of sound details. The time reference point for the dual-modal signals is established via a hardware trigger mechanism. The system uses a GPS timing module to provide a unified timestamp with a time accuracy of better than 1ms. To address synchronization issues during signal acquisition, a high-precision FPGA-based clock synchronization circuit was designed to ensure that the sampling clock deviation of the two signals is within 100ns. At the signal source, active shielding technology and differential amplification circuits are used to improve the signal-to-noise ratio by at least 45dB, effectively suppressing the effects of power frequency interference and ambient noise. This acquisition and synchronization process creates a dual-modal synchronous sampling data stream, with each sampling point associated with a unique timestamp.
[0036] The acquired bimodal synchronously sampled data stream undergoes data classification and processing. The system first performs time alignment and segmentation of the signals. Taking into account the characteristics of human speech interaction, the signals are segmented into 200ms units, with a 50% overlap between each unit. This ensures signal continuity and facilitates subsequent feature extraction and analysis. An adaptive threshold mechanism is introduced during the classification process, calculating the signal's short-term energy and zero-crossing rate to preliminarily delineate speech activity from non-speech areas. For EEG signals, a sliding window method is used for preprocessing, with a window length of 2 seconds and a step size of 1 second. The Hjorth parameter is also incorporated to characterize signal segments. Using a multi-level data classification strategy, the system classifies the signals into four categories: pure speech segments, pure EEG segments, bimodal activity segments, and background noise segments. The recognition threshold for bimodal activity segments is determined through extensive experimental data analysis and typically requires that the normalized energy values of both signals exceed 85% of the set threshold. After the classification process is completed, a labeled multimodal data sequence is formed, and each data segment contains clear type identification and time range information.
[0037] A time series information output structure is constructed based on the generated labeled multimodal data sequence. The system standardizes the data sequence, with each sequence unit containing key information such as timestamp, signal type identifier, signal quality indicator, and feature vector. The signal quality indicator is calculated by comprehensively evaluating factors such as signal-to-noise ratio, baseline drift, and electromyographic artifact intensity, and is expressed as a normalized numerical value. The feature vector contains multidimensional characteristic parameters in the time and frequency domains, such as the signal's mean, variance, skewness, kurtosis, and dominant frequency component. To facilitate subsequent processing, the system constructs a multi-level index for the time series information: the primary index is based on the timestamp, and the secondary index is based on the signal type and quality indicator. This multi-dimensional index structure significantly improves data retrieval and processing efficiency, enabling the system to locate and extract signals of specific time periods or types within milliseconds. In practical applications, this time series information structure supports subsequent core functions such as feature extraction, pattern recognition, and interaction intent analysis, providing a reliable data foundation for the entire hybrid interaction system. The time-series data generates approximately 2MB of storage per second. The system employs a ring-buffer-based data management strategy, keeping the most recent 10 minutes of data readily accessible while asynchronously writing historical data to storage for persistence. The final output is a standardized time-series data stream that includes complete signal characteristics, classification labels, and time-series correlation information.
[0038] Exemplarily, the timing information is subjected to noise reduction and signal separation processing, and the processing results corresponding to the separation processing are optimized and purified to output a pure waveform, including: using an adaptive filter to eliminate power frequency interference on the EEG signal, and combining independent component analysis to perform electromyography artifact separation; implementing a blind source separation algorithm on the speech signal to eliminate environmental noise, and suppressing sudden interference through morphological filtering; using an improved wavelet threshold method to denoise the EEG signal and retain characteristic waveform details; performing spectral subtraction noise reduction processing on the speech signal, and combining a Wiener filter to enhance speech features; dynamically adjusting the noise reduction parameters of the timing information through signal quality indicators, and outputting the pure waveform with improved signal-to-noise ratio.
[0039] The time series information output from the above steps undergoes noise reduction and signal separation. The system first extracts signal features for each time segment based on the classification labels and quality indicators in the time series information stream. For time series segments with quality indicators below 0.7, the system adds preprocessing steps such as baseline drift correction and outlier detection. Differentiated processing is applied to different signal types identified in the time series information. For EEG signal segments, a five-layer wavelet decomposition structure is used, combining frequency domain feature parameters from the time series correlation information. The db4 wavelet basis function is used for multiresolution analysis, with the decomposition frequency bands corresponding to the δ, θ, α, β, and γ bands. The system dynamically adjusts the wavelet threshold based on the signal quality indicators in the time series information. When the quality indicator is greater than 0.8, a soft thresholding method is used, while when it is below 0.8, a hard thresholding method is used for coefficient processing. For speech signal segments, the system utilizes the short-term energy characteristics in the time series information in conjunction with an improved Wiener filtering algorithm for noise reduction. The filter parameters are dynamically adjusted based on the signal variance characteristics in the time series information. The processing window varies within the range of 16-32ms, with a 50% overlap ratio. For bimodal activity segments, the system establishes a joint processing framework based on the temporal correspondence in the timing information. By calculating the standardized cross-correlation function, it identifies the coupling characteristics between the signals and uses tensor decomposition technology to achieve preliminary signal separation. Throughout the processing process, the system continuously tracks the timestamps in the timing information to ensure that the processed signal segments strictly correspond to the original timing structure. Through this series of targeted noise reduction and separation processes, a separation result containing the main signal components and residual noise is obtained. The main signal components retain more than 90% of the effective information, and the energy of the residual noise part is less than 15% of the original signal.
[0040] The separation results are optimized and purified. The system establishes a multi-level optimization structure based on the characteristics of the main signal components and residual noise. First, the residual aliasing part in the main signal component is finely separated using independent component analysis technology, and a separation model of 15 independent components is constructed. In the separation process, the improved FastICA algorithm is used to project the signal into a high-dimensional feature space, and the optimal projection dimension is determined by analyzing the signal energy distribution characteristics obtained in the first stage. In the feature space, the system optimizes the separation matrix in combination with the negative entropy maximization criterion, adopts an adaptive learning rate strategy, sets the initial value to 0.01, and dynamically adjusts according to the energy level of the residual noise until the convergence threshold reaches 10⁻ 6 . For the separated independent components, the system calculated statistical features such as kurtosis, sample entropy, and Hilbert spectrum, and matched them with the signal features identified in the first stage. At the same time, based on the interference characteristics detected in the first stage, the system established an adaptive notch filter array to accurately suppress power frequency interference and its harmonics. The center frequency of the notch filter tracks the fluctuation of the grid frequency in real time through phase-locked loop technology, and the bandwidth is automatically adjusted in the range of 0.5-2Hz according to the interference characteristics. During the optimization process, the system continuously evaluates the degree of signal improvement, and increases the signal-to-noise ratio of the EEG signal from 15dB in the first stage to 30dB, and the signal-to-noise ratio of the speech signal from 20dB to 35dB, while maintaining the main characteristic structure of the signal. After this series of optimization and purification processes, a highly purified signal result was obtained. Its characteristic structure is consistent with the main signal components separated in the first stage, but the quality indicators are significantly improved.
[0041] A pure waveform is constructed based on the highly purified signal results. The system designs an adaptive reconstruction algorithm, which first selects reconstruction parameters according to the degree of purification of the signal. For signal segments with a signal-to-noise ratio of 30dB or more, the Hilbert transform is used to directly obtain the instantaneous amplitude and phase information; for parts with a lower signal-to-noise ratio, soft threshold processing in the wavelet domain is first performed. During the reconstruction process, the system designs an adaptive zero-phase digital filter group based on the signal characteristics optimized in the second stage, and uses a 0.5-45Hz bandpass filter for the EEG signal and a 300-3400Hz bandpass filter for the speech signal. The order of the filter is dynamically adjusted according to the degree of purification of the signal. In order to deal with the edge effects in the filtering process, the system analyzes the local characteristics of the signal obtained in the second stage, and uses an improved wavelet packet transform technology for adaptive boundary processing. The extension length is dynamically adjusted according to the local change characteristics of the signal, and is maintained between 5% and 10% of the signal length. To ensure the continuity of the reconstructed signal in the time domain, the system, combined with the signal quality assessment results from the second phase, employs an improved cosine smoothing strategy at the junctions of adjacent data segments. The transition interval length is adaptively set based on the signal's rate of change, with shorter transition intervals (approximately 10ms) for high-quality signal segments and longer transition intervals (approximately 20ms) for lower-quality signal segments. The resulting pure waveform achieves a high signal-to-noise ratio (>50dB for EEG signals and >60dB for speech signals), complete time-frequency characteristics, and continuous phase characteristics, surpassing the results of the second phase in all metrics.
[0042] In step S102, a two-way analysis is performed on the pure waveform to divide the feature dimensions, and the modal conversion is completed through the constructed mapping chain to output the feature identifier; a time series analysis is performed on the feature identifier to identify the key time series feature set, and an alignment feature mapping structure is established to combine the modalities and generate the associated pattern.
[0043] Specifically, a two-way feature analysis is performed on the pure waveform output from the above steps. The system fully utilizes the three key characteristics of the processed signals: a high signal-to-noise ratio (EEG > 50dB, speech > 60dB) ensures the accuracy of feature extraction, complete time-frequency characteristics support the reliability of multi-scale analysis, and continuous phase characteristics ensure the consistency of timing features. Taking cross-language communication scenarios as an example, when a user uses AI headphones to converse with a foreign language speaker, the system can accurately capture the EEG characteristics of the user's speech expression and language comprehension based on these high-quality waveform characteristics.
[0044] In some embodiments, the pure waveform is subjected to a two-way analysis and feature dimension division, modal conversion is completed through a constructed mapping chain, and feature identification is output, including: wavelet packet decomposition of the EEG signal to obtain time-frequency features, and intrinsic modal components are extracted in combination with empirical mode decomposition; Mel-frequency cepstral coefficients and acoustic feature parameters are extracted from the speech signal; the feature dimension discrimination is evaluated by the Fisher discriminant criterion, and a hierarchical feature selection strategy is established; and the feature identification is output according to the time-frequency features, intrinsic modal components, Mel-frequency cepstral coefficients, acoustic feature parameters and hierarchical feature selection strategy.
[0045] In the EEG signal channel, wavelet packet decomposition is used to perform time-frequency analysis of the signal. A six-layer decomposition using the sym4 wavelet basis is performed to obtain energy distribution features across 64 frequency bands. For each frequency band, the relative energy ratio, instantaneous frequency center, and band entropy are calculated to construct an initial feature matrix. Simultaneously, empirical mode decomposition is used to extract the signal's intrinsic modal features. The first five intrinsic mode functions are selected and their instantaneous frequency and amplitude envelopes are calculated. In the speech signal channel, Mel-Frequency Cepstral Coefficients (MFCCs) are used for feature extraction, with a frame length of 25ms and a frame shift of 10ms. 13-dimensional basic feature parameters are extracted. Acoustic features such as pitch period, formant frequency, and spectral entropy of the speech are also calculated. This two-way feature analysis fully utilizes the high-quality characteristics of the clean waveform, resulting in a feature set of 128 dimensions across the time, frequency, and time-frequency domains, including 64 dimensions of EEG features and 64 dimensions of speech features.
[0046] The bimodal 128-dimensional feature dimension set is constructed into a feature network. The system first processes the 64-dimensional feature subsets of EEG and speech separately, employing a hierarchical feature selection strategy. The discriminant value of each feature dimension is calculated based on the Fisher discriminant criterion, and an importance ranking is established for the extracted EEG features (such as frequency band energy distribution and modal function features) and speech features (such as MFCC coefficients and acoustic features). Feature dimensions with a discriminant value greater than 0.6 are selected as primary features, including key frequency band energy and typical modal features in EEG signals, as well as core acoustic parameters in speech signals. Principal component analysis is used to reduce the dimensionality of these primary feature dimensions, retaining principal components with a cumulative contribution rate of 85%. For auxiliary feature dimensions with a discriminant value between 0.3 and 0.6, a local linear embedding algorithm is used for nonlinear dimensionality reduction, compressing the dimensions to one-third of the original dimensionality. The system also establishes a feature correlation matrix to identify redundant relationships between features. When the correlation coefficient exceeds 0.8, feature dimensions with high information entropy are retained. Taking cross-language understanding as an example, the system can extract and organize the 45 most discriminative feature nodes from the initial 128-dimensional features, forming a compact feature network structure.
[0047] Modal conversion is performed using the 45-node feature network. The system processes primary feature nodes (discrimination > 0.6) and auxiliary feature nodes (discrimination 0.3-0.6) separately, establishing a hierarchical feature mapping framework. These feature nodes are first normalized, using the z-score method to ensure that the feature distribution has zero mean and unit variance. Based on the normalized features, deep canonical correlation analysis is used to learn the inter-modal mapping relationships for the primary and auxiliary feature groups, respectively. To enhance the robustness of the feature mapping, a regularization constraint is introduced with a regularization coefficient set to 0.01, and the optimal mapping matrix is solved using an alternating optimization algorithm. In practical applications, when users encounter expression barriers during cross-language communication, the system can bidirectionally map the user's EEG cognitive features to their speech expression features based on the established feature network, achieving more accurate cross-language interaction support. After iterative optimization, the system establishes a stable feature conversion mechanism and forms a complete set of conversion parameters, including both primary and auxiliary mapping relationships.
[0048] The system outputs feature identifiers based on the feature conversion parameter set. The system constructs identifier generation networks for the primary and auxiliary mapping relationships, each employing a multi-layer perceptron architecture. The primary mapping channel utilizes a three-layer network structure with 32, 16, and 8 nodes, respectively, using the ReLU activation function and a dropout rate of 0.3. The auxiliary mapping channel adopts a two-layer structure with 16 and 8 nodes. During training, the system performs weighted fusion based on the importance of different mapping relationships and employs batch normalization to improve network convergence. The learning rate is initially set to 0.001 and dynamically adjusted using a cosine annealing strategy. To adapt to the characteristics of different language learners, the system implements a dynamic identifier mapping mechanism. For example, in international business negotiation scenarios, the system can accurately identify and convert the user's language comprehension and expression features based on the mapping relationships provided by the conversion parameter set, providing real-time support for cross-language communication. The final output feature identifier contains eight key dimensions, five from the primary mapping channel and three from the auxiliary mapping channel, forming a stable and reliable feature representation system.
[0049] Temporal analysis is performed on the feature identifiers output from the above steps. The system performs temporal processing based on the 8-dimensional feature identifiers output from the above steps, including the 5 dimensions of the primary mapping channel and the 3 dimensions of the auxiliary mapping channel. In a cross-language communication scenario, when a user uses an AI headset for real-time conversation, the system establishes a time window for each conversation segment, with a window length set to 2 seconds and a sliding step size of 500ms. For the main dimension features (5 dimensions), their temporal stability indicators are calculated, including the smoothness (threshold 0.15) and persistence (at least 3 consecutive windows) of the feature trajectory; for the auxiliary dimension features (3 dimensions), their mutation characteristics (threshold 0.3) and transition patterns are analyzed in detail. At the same time, an improved dynamic time warping algorithm is introduced, using a bidirectional search strategy to compensate for the temporal misalignment between features of different dimensions, and the warping window is set to 100ms. The system also establishes a hierarchical attention mechanism, setting differentiated attention thresholds (0.7 for the main dimension and 0.5 for the auxiliary dimension) based on the stability and distinctiveness of the features, and identifying three typical patterns of feature changes: the gradual change process of stable features, the mutation points of significant features, and the coordinated changes of combined features. These feature change patterns and their time points constitute a key set of temporal features.
[0050] An alignment mechanism is established around the described temporal feature set. Based on the three identified feature change patterns, the system constructs corresponding alignment strategies. For the gradual changes of stable features, sliding correlation analysis is used to calculate the optimal matching interval of the feature gradient, requiring a correlation coefficient greater than 0.75. For the mutation points of significant features, precise timestamp alignment is used, with a maximum time deviation of 30ms allowed. For the coordinated changes of combined features, a dynamic programming algorithm is used to find the optimal temporal mapping relationship, with a mapping quality threshold of 0.8. At the local alignment level, the system fine-tunes the sequence within 200ms before and after each feature change point. The maximum mutual information criterion (threshold 0.7) is used for the primary dimension features, and an improved dynamic programming algorithm (with a warping window of 50ms) is used for the auxiliary dimension features. At the global alignment level, the system constructs a weighted sequence matching graph based on the importance of different feature change patterns, and achieves global optimization using a shortest path algorithm. To account for uncertainty near feature mutation points, the system also dynamically adjusts the alignment window size based on the feature stability index, using smaller alignment windows (±100ms) for highly stable features and larger windows (±200ms) for less stable features. Through this series of targeted alignment processing, an aligned feature mapping structure containing multi-scale temporal relationships is formed.
[0051] Modal combination is performed using the aligned feature mapping structure. The system designs an adaptive feature fusion strategy based on the temporal relationships within the mapping structure. For stable features with gradual changes (alignment quality > 0.9), a smooth fusion mechanism based on gated recurrent units is employed, with a feature retention rate set to 90%. The gating parameters are controlled by the feature stability index. For sudden changes (alignment quality 0.7-0.9), attention-enhanced instantaneous feature combination is used, with an attention weight threshold of 0.6. For co-changing regions (alignment quality < 0.7), nonlinear feature combination is achieved using a multi-layer perceptron (MLP) with a combination threshold of 0.5. The system integrates a feature selection mechanism based on alignment quality. When the confidence level of a local alignment result is low, the fusion weight of the corresponding feature is automatically reduced. In practical applications, this adaptive fusion mechanism based on alignment quality can accurately capture dynamic feature changes during cross-language communication. For example, when the system detects a sudden change in speech features that is highly aligned with a significant change in EEG features, it strengthens the feature combination weight at that moment. For periods of low alignment quality, however, it relies more on highly stable features to maintain the reliability of the modal combination. After this series of fusion processing, a time-aligned dynamic feature combination sequence is formed.
[0052] The system generates association patterns based on the dynamic feature combination sequence. Based on the different types of feature changes in the combination sequence, a hierarchical pattern extraction network is constructed. In the temporal feature extraction layer, dedicated processing units are designed for three types of patterns: gradual, sudden, and coordinated changes. Gradual changes are extracted using a bidirectional long short-term memory network (64 units), sudden changes are captured using a convolutional network with residual connections, and coordinated changes are enhanced using an attention mechanism. In the pattern recognition layer, the system dynamically adjusts the weights of different patterns based on the temporal alignment quality of the feature combination. Feature combinations with high-quality alignment (alignment error <50ms) are given a higher weight (0.5), those with medium-quality alignment (error 50-100ms) are given a medium weight (0.3), and those with low-quality alignment (error >100ms) are given a lower weight (0.2). The association encoding layer integrates different patterns into a unified representation using a variational autoencoder architecture (encoding dimension 32, reconstruction error threshold 0.1). In practical applications, the system can identify complete chains of interaction patterns from continuous user conversations. For example, when identifying a pattern of "difficulty understanding - attempting to understand - breakthrough understanding," the system not only tracks feature changes at each stage but also identifies key moments and triggering conditions for pattern transitions based on the temporal alignment of feature combinations. The resulting correlation pattern contains complete temporal evolution information, preserving the dynamic details of feature changes while embodying the overall patterns of pattern transitions.
[0053] Step S103: perform path analysis on the association pattern and mark the combined nodes, reconstruct the features through the transformation chain and output the representation structure; perform hierarchical fusion on the representation structure and extract the fusion area, and output the fusion elements based on the fusion area; supplement the missing elements and locate the missing parts, reconstruct the data of the missing parts and output the complete content.
[0054] Specifically, path analysis is performed on the association patterns output in the above steps. The system constructs a pattern transition path diagram based on the complete temporal evolution information recorded in the association patterns, including the dynamic details of feature changes (such as the state transition process) and the overall laws governing pattern transitions (such as the transition trigger conditions). For each identified interaction pattern chain, the system first extracts its complete evolution trajectory in the feature space. In intervals dominated by gradual changes (such as the gradual change in cognitive load identified in the above steps), a piecewise linear fitting method is used with a mean square error threshold of 0.1. Path nodes are set when the fitting error exceeds the threshold. In intervals dominated by sudden changes (such as the breakthrough moments of understanding captured in the above steps), the first-order differences of the features are calculated. When the difference value exceeds 2.5 times the standard deviation, the transition point in the feature space is marked as a key node. In intervals of coordinated changes (such as the synchronous changes in the multimodal features identified in the above steps), the marker position is determined by calculating the feature coordination degree (using the Pearson correlation coefficient with a threshold of 0.8). In cross-language communication scenarios, the system can convert the pattern chain identified in the above steps (difficulty in understanding - attempting to understand - breakthrough in understanding) into a specific path node sequence, where each node contains: a timestamp (accuracy 1ms); a feature state vector (32 dimensions); a conversion type identifier (gradual / mutation / cooperation, confidence threshold 0.75); and trigger condition parameters (feature threshold, time window, duration). Through this path analysis based on time-series evolution, the system forms a set of path nodes that include conversion types, trigger conditions, and evolution rules.
[0055] A feature structure is constructed using the set of path nodes described above, which include transition types, trigger conditions, and evolution patterns. Based on the node's transition type, the system designs a corresponding feature organization framework for each node type. For nodes identified in path analysis with a gradual transition type (e.g., cognitive load change points), the system establishes a gradual feature subspace with a dimension of 16 to describe the continuous change of features. Principal component analysis is used to ensure 95% energy retention. For nodes with a sudden transition type (e.g., understanding breakthrough points), a high-dimensional feature subspace with a dimension of 24 is constructed, using kernel function mapping to capture feature transitions. For nodes with a collaborative transition type (e.g., multimodal synchronization points), a collaborative feature subspace with a dimension of 20 is set to express feature interactions. In terms of temporal organization, the system constructs a directed feature graph based on the node's trigger conditions and evolution patterns: nodes with strong trigger conditions (correlation > 0.8) are directly connected, while nodes with weak trigger conditions (correlation 0.5-0.8) are transitionally connected through intermediate states. At the same time, a feature propagation mechanism based on evolutionary laws is introduced, allowing features to flow through the network according to the evolutionary laws determined by path analysis. Through this series of structural processing, a multi-level network structure that reflects the evolutionary laws of features is formed.
[0056] Features are reconstructed using the multi-level network structure. The system employs differentiated reconstruction strategies for different types of node connections within the network. For directly connected node pairs (derived from strong trigger conditions), a high-precision feature reconstruction mechanism is employed, comprising a dual-path encoding and decoding architecture. Each path utilizes a five-layer neural network for feature processing (with layer dimensions configured as 32-64-128-64-32), with residual connections added between key layers to prevent information loss. For transitionally connected node pairs (derived from weak trigger conditions), a smooth reconstruction method based on a bidirectional LSTM is employed, capturing the temporal dependencies of features through hidden layer states (dimension 128). During the reconstruction process, the system dynamically adjusts parameters based on the node's position in the network and its connectivity properties: a large learning rate (0.01) is used for fast convergence of core nodes, while a smaller learning rate (0.005) is used for transition nodes to ensure stability. Furthermore, an attention-based feature enhancement mechanism is implemented, with attention weights dynamically adjusted based on the strength of the feature's connections in the original network. To address the temporal evolution of features, the system designed a forward prediction module, using a causal convolutional network (with kernel sizes of 3, 5, and 7) to capture temporal patterns at different scales. Parameter optimization was performed by comparing the predicted results with the actual feature sequence (requiring a similarity > 0.85). Throughout the reconstruction process, the system simultaneously optimized feature reconstruction accuracy (average error < 0.1) and temporal coherence (correlation between adjacent frames > 0.8). Through multiple rounds of iteration (typically 500-1000), a reconstructed feature set was obtained that both preserves the network's structural characteristics and exhibits good generalization.
[0057] In some embodiments, the path analysis is performed on the association pattern and the combination nodes are labeled, and the representation structure is output after the features are reconstructed through the transformation chain, including: using a dynamic programming algorithm to find the optimal feature path for the association pattern, and labeling the combination nodes of time-series association; identifying the key connection relationships of the combination nodes through a graph attention network, and constructing a feature transformation chain; implementing path optimization based on a Markov model on the feature transformation chain to eliminate redundant feature connections; using a tensor decomposition method to reconstruct the feature space corresponding to the feature transformation chain, maintaining the topological relationship between features, and outputting a multi-dimensional representation structure containing principal component features and time-series relationships.
[0058] The system generates a representation structure based on the reconstructed feature set that preserves the network structure. A hierarchical feature fusion network is constructed to process features reconstructed with different connection types separately. For features reconstructed with direct connections, a priority processing mechanism is implemented, ensuring their dominance in the representation structure through enhanced attention, with attention weights ranging from 0.6 to 0.8. For features reconstructed with transitional connections, a progressive fusion strategy is used with weights ranging from 0.2 to 0.4 to achieve a smooth transition of features. A multi-layer feature verification mechanism based on network topology is designed. The first layer verifies feature similarity by calculating the Euclidean distance between the reconstructed and original features (threshold <0.15). The second layer verifies structural consistency to ensure that the reconstructed features retain key connectivity relationships in the original network (connection retention rate >90%). The third layer verifies temporal coherence, checking the smoothness of feature evolution (feature change rate between adjacent time points <0.2). In cross-language communication scenarios, this network-structure-based representation can accurately describe the evolution of a user's cognitive state. For example, when the system detects the activation of a "comprehension breakthrough" node, it first strengthens the reconstruction features of that node (weight 0.8) while simultaneously supplementing the details of the state transition through adjacent transition nodes (weight 0.3), such as the gradual reduction of cognitive load and the steady improvement of attention levels. Through this multi-level feature fusion and verification mechanism, the system ultimately generates a representation structure that maintains a clear positioning of state transitions while achieving a continuous expression of state evolution.
[0059] The representation structure output from the above steps is hierarchically fused. The system fully utilizes the high-weight features of direct connections (original weight 0.6-0.8) and the low-weight features of transition connections (original weight 0.2-0.4) in the representation structure to construct a multi-level fusion system. For the core feature regions formed by direct connections, a deep feature fusion strategy is adopted, extracting features through three layers of residual blocks. Each residual block consists of two convolutional layers with inherited weights (with an inherited weight ratio of 0.7) and a skip connection. To maintain the clear localization characteristics of the original representation structure, the system introduces a self-attention mechanism in the residual blocks, with the attention weights initialized based on the original weight distribution. For the gradient feature regions formed by transition connections, the original smooth transition characteristics are retained, using a lightweight feature fusion network consisting of a convolutional layer with inherited weights and an adaptive pooling layer. Based on the temporal characteristics of different regions, the system sets a feature aggregation window based on the original temporal granularity. A smaller window (100ms) is used in the core region to preserve detailed features, while a larger window (300ms) is used in the gradient region to capture trend characteristics. In cross-language communication scenarios, when a "comprehension breakthrough" is detected, the system can enhance and smooth out high-weight feature regions (such as sudden changes in cognitive load) and low-weight feature regions (such as gradual changes in attention) in the original representation structure. This hierarchical fusion process yields a set of fused regions with clear hierarchical properties, encompassing the feature combinations and temporal relationships inherited from the original representation structure.
[0060] In some embodiments, the hierarchical fusion of the representation structure and the extraction of the fusion area, and the output of the fusion element based on the fusion area include: constructing a multi-scale feature pyramid structure for hierarchical feature fusion of the representation structure to obtain the fusion area; calculating the feature area weights of the fusion area through the attention mechanism to identify the key fusion area; using the region growing algorithm to expand the fusion boundary of the key fusion area, maintaining feature continuity, and outputting the fusion element containing the EEG-speech collaborative feature.
[0061] The hierarchical fusion regions, along with their contained feature combinations and temporal relationships, are integrated into a feature framework. Based on the hierarchical characteristics of the fusion regions, their inherited feature combinations, and their temporal relationships, the system designs a hierarchical feature organization structure. At the top level of the framework, a feature integration module is deployed for the core regions. A graph neural network is used to handle feature transfer between high-weight nodes, with node updates configured based on the weight distribution within the fusion regions. At the middle level, a feature transfer network is constructed to handle gradually changing regions. Gated recurrent units are used to achieve progressive feature fusion, with gating parameters inheriting the temporal characteristics of the fusion regions. At the bottom level, a feature aggregation module is deployed to perform preliminary integration of features from different hierarchies using inherited weight ratios. For example, when processing the "understanding breakthrough" scenario, the top-level module prioritizes cognitive state features that inherit high weights, the middle-level module processes attention change features that inherit smoothness, and the bottom-level module preliminarily combines these features according to their original weight ratios. The framework maintains the inheritance of the original fusion region's characteristics at all levels of feature processing, forming a comprehensive feature integration system.
[0062] The described complete feature integration system is used to aggregate features. The system implements differentiated feature aggregation strategies for the three layers of the framework. For top-level feature processing, an attention enhancement mechanism based on the original weight distribution is designed. The initial attention matrix dimension is 64×64, and feature selectivity is enhanced through a three-layer feedforward network (with a dimensional configuration of 64-32-16). For mid-level features, an adaptive temporal aggregation method is employed. The sliding window size is dynamically adjusted within the range of 50-200ms based on the feature change rate, and the window overlap ratio is maintained at 40%. For bottom-level feature processing, a multi-level caching mechanism based on feature saliency is implemented, with cache sizes of 32, 16, and 8, respectively, to maintain feature continuity at different time scales. Furthermore, bidirectional consistency constraints are introduced during the feature aggregation process: temporal consistency ensures smooth feature evolution, and spatial consistency maintains the rationality of feature distribution. The system also implements feature checkpoints during the aggregation process, performing feature validation every 200ms of data processing. Verification covers feature integrity, temporal coherence, and the rationality of distribution. Through this series of refined aggregation processing, we finally obtained a set of multi-scale and multi-level feature aggregation results.
[0063] Based on the multi-scale and multi-level feature aggregation results, the system outputs fused features. A three-layer feature integration network is designed, consisting of a feature extraction layer, a mapping optimization layer, and a fusion and reassembly layer. The feature extraction layer uses a depthwise separable convolutional network, with 1×1, 3×3, and 5×5 convolution kernels set for aggregated features at different scales, respectively. The convolutional layer channels are configured to 16, 32, and 64. In the mapping optimization layer, a nonlinear feature transformation mechanism is implemented, using batch normalization and a LeakyReLU activation function (negative slope 0.2) to manipulate the feature distribution. Furthermore, a residual learning module is introduced, with each residual unit consisting of two convolutional layers and a skip connection to maintain feature discriminability. The fusion and reassembly layer employs a dynamic weight allocation mechanism, calculating fusion weights based on local correlation and global consistency of features. For example, when assessing cognitive status, the system dynamically weights feature elements at different levels, such as attention features and comprehension indicators. The weight range for attention features is 0.3-0.5, and for comprehension indicators is 0.4-0.6. The final output fusion elements are represented by a 128-dimensional vector, of which high-level features occupy 48 dimensions, middle-level features occupy 48 dimensions, and bottom-level features occupy 32 dimensions, achieving a balanced proportion of features at each level.
[0064] In some embodiments, the method of supplementing the missing elements and locating the missing parts, and reconstructing the missing parts to output complete content includes: establishing a feature integrity assessment model to obtain the missing parts of the elements corresponding to the fused elements; restoring the temporal continuity of the missing parts of the elements according to a spatiotemporal interpolation algorithm; performing feature distribution alignment processing on the missing parts of the elements to ensure the consistency of the reconstructed data, and outputting the complete content containing complete temporal features and semantic associations.
[0065] The fused features output from the above steps are supplemented for missing features. The system builds a missing feature detection framework based on the fused features generated in the above steps, including their three-layer feature structure (48-dimensional high-level features, 48-dimensional mid-level features, and 32-dimensional low-level features) and a dynamic weighting mechanism (calculated based on local correlation and global consistency of features). The system first analyzes the three-layer feature structure: high-level features (such as attention features, with a weight range of 0.3-0.5) characterize the user's attention and comprehension level; mid-level features (such as comprehension indicators, with a weight range of 0.4-0.6) reflect voice intonation and EEG fluctuations; and low-level features (with a stability indicator threshold of 0.8) provide support and reflect emotional state. In cross-language communication scenarios, the system monitors missing features in real time. For example, when a user is comprehending complex sentences, high-level features may exhibit anomalies that deviate from the original weight distribution (deviation threshold of 0.25); and when responding quickly, mid-level features may deviate from the original dynamic range (exceeding the threshold and being marked as anomaly). By analyzing the sparsity of feature distribution (density threshold 0.7), temporal continuity (interruption tolerance 50ms) and feature correlation (correlation coefficient threshold 0.6), the system ultimately identified three types of deficiencies: abnormal cognitive features, interrupted behavioral features and deviations from basic features.
[0066] Data reconstruction is performed for the three missing features. For cognitive feature anomalies, the system leverages the abnormal distribution features identified in the first stage and reconstructs them using a variational autoencoder (encoding dimension 32, KL divergence constraint 0.01), focusing on preserving the dominant characteristics of the features. For example, when detecting user distraction, the system performs pattern learning based on high-quality feature segments (quality score > 0.85) in historical data to reconstruct features consistent with the dominant distribution. For behavioral feature interruptions, the system uses a bidirectional gated recurrent unit network (hidden dimension 128, forget gate threshold 0.3) with a hidden layer dimension of 128 and a forget gate threshold of 0.3 to predict missing segments based on the complete sequences before and after the interruption (each 200ms). When addressing speech and intonation interruptions, the dynamic features of the context are used to infer the changing trends during the interruption. For deviations from basic features, the system selects adjacent features with high quality assessment scores (> 0.8) as references and reconstructs them to ensure consistency with the overall feature distribution (distribution distance threshold 0.2). Through the joint reconstruction of multiple layers of features, a complete feature supplementation dataset is formed.
[0067] Information infill is performed based on the complete feature supplementation dataset. The system has designed a hierarchical infill strategy, implementing infill based on the primary and secondary relationships of reconstructed features. Different infill strategies are employed in different communication scenarios. For example, in formal meetings, the system prioritizes accurate infilling of cognitive features (accuracy requirement > 0.9). When detecting fluctuations in the user's attention while understanding professional terminology, it prioritizes restoring attention level data and ensures its synchronization with the meeting rhythm (time deviation < 30ms). In everyday conversations, the system prioritizes the natural transition of behavioral features (smoothness threshold 0.8), such as continuous changes in voice intonation. When a user switches from their native language to a foreign language, the system captures feature changes during this transition, including slower speech speed and increased pauses, and implements appropriate feature infill. The system also incorporates a scenario-adaptive mechanism, increasing the infill weight of behavioral features (up to 0.6) in noisy environments and enhancing the accuracy of cognitive features (up to 0.95) in quiet environments. For extended interactions, the system implements a dynamic sliding window filling strategy. The window size adaptively adjusts to the interaction rhythm (50ms-300ms) to ensure real-time and continuous filling. For example, a smaller filling window is used in fast-paced conversations, while a larger window is used during in-depth discussions to capture the complete thought process. This multi-level, multi-scenario filling process forms a clear and complete feature sequence.
[0068] The system outputs complete content based on the complete, hierarchical feature sequence. A feature integration network is constructed to reorganize the populated features. In academic presentations, the system prioritizes cognitive features (increasing them to 0.7) to track the audience's understanding in real time. When the difficulty of understanding certain professional concepts exceeds a threshold (0.8), the system can pinpoint the content segments that require focus. In business negotiations, the system prioritizes a balance between behavioral and cognitive features (weight ratio 0.5:0.5). By analyzing subtle changes in speech expression (change detection threshold 0.15) and fluctuations in cognitive load (fluctuation range ±0.2), the system helps users navigate the negotiation process. For example, if the system detects an increase in cognitive load (exceeding 50% of the baseline) while understanding a foreign party's offer, it focuses on feature integrity during this period to ensure accurate capture of the user's understanding and decision-making process. In everyday social interactions, the system prioritizes the natural expression of emotional features (naturalness threshold 0.85) to ensure smooth interaction. When handling multilingual switching scenarios, the system accurately captures feature changes at language transition points (transition recognition accuracy > 0.9) and maintains continuity in feature distribution before and after transitions (distribution distance < 0.25). When users frequently switch between languages during conversations, the system pays special attention to changes in cognitive load before and after language transitions. By integrating cognitive and behavioral features, it helps users achieve smooth language switching. The resulting complete feature representation maintains the original hierarchical properties while ensuring continuity in feature distribution (temporal correlation > 0.85).
[0069] Step S104: perform intent analysis on the complete content to obtain key elements, convert the key elements into interaction information and output the interaction intention; integrate the information of the interaction intention to obtain conversion rules, generate interaction commands according to the conversion rules, and complete the EEG-speech hybrid interaction based on the brain-computer AI headset.
[0070] Specifically, intent is parsed for the complete content output from the above steps. The system establishes an intent parsing framework based on the feature expressions generated in the above steps, including cognitive comprehension changes, behavioral language switching characteristics, and emotional state fluctuations. When processing high-level cognitive features, the system implements an intent inference mechanism to identify user intent by analyzing the temporal variation patterns of these features (rate of change threshold 0.3). For example, when a user encounters comprehension difficulties in cross-language communication, the system infers the intent of difficulty understanding from sudden changes in attention features (sudden change threshold 0.4) and increased cognitive load (increase > 50%). When processing mid-level behavioral features, the system combines voice intonation changes (fluctuation range ±0.25) and EEG activity patterns (activity threshold 0.7) to capture the user's expressed intent. For low-level emotional features, the system identifies the user's emotional intent by analyzing their fluctuation patterns (fluctuation period 200-500ms). In foreign language conference scenarios, when the system detects frequent fluctuations in attention (fluctuation frequency > 0.5Hz) and slowing speech (slowdown > 30%), it locates these intent change intervals. Through multi-dimensional intention analysis, the system obtains a set of analytical intervals with temporal correlations, each of which contains changes in intention characteristics at the three levels of cognition, behavior, and emotion.
[0071] In some embodiments, the intention parsing of the complete content to obtain key elements, converting the key elements into interaction information and outputting interaction intentions includes: constructing a semantic network analysis model to identify the intention-related key elements of the complete content; obtaining the element association strength corresponding to the intention-related key elements through a fuzzy reasoning algorithm to generate candidate interaction intentions; using a reinforcement learning model to optimize the intention generation strategy corresponding to the candidate interaction intentions, eliminate ambiguous understanding, and output interaction intentions including language assistance, rhythm control and state adaptation intentions.
[0072] The parsing intervals containing the three layers of intent features are established as a comprehension model. The system designs corresponding comprehension structures based on the intent feature type of each interval. For intervals dominated by cognitive intent features (dominance > 0.7), such as those with comprehension difficulties caused by attention fluctuations, a cognitive load-based comprehension assessment module is constructed with an accuracy threshold of 0.85. For intervals dominated by behavioral intent features (dominance 0.5-0.7), such as the language expression preparation phase, a speech-based and EEG-based expression intent recognition mechanism is established with an accuracy threshold of 0.8. For intervals dominated by emotional intent features (dominance < 0.5), an emotional state understanding module is constructed with a state resolution of 0.1. In business negotiation scenarios, when users need to express complex ideas in a foreign language, the system determines the comprehension strategy for the current interval by identifying the primary and secondary intent features (difference in primary and secondary weights > 0.3). For example, during the expression preparation phase, the system increases the comprehension weight of behavioral intent features (to 0.6); during the content organization phase, it increases the comprehension weight of cognitive intent features (to 0.7). Through this understanding structure based on the priority of features, the system forms a complete intention understanding model.
[0073] Information is extracted using the aforementioned feature-based intent understanding model. The system implements hierarchical information extraction based on the weights of features in different intervals specified in the model. In the cognitively dominant interval (weight > 0.6), the system focuses on extracting user comprehension features, such as the difficulty of understanding professional terminology (difficulty score 0-1) and the cognitive load during language switching (load index threshold 0.7). In the behaviorally dominant interval (weight 0.4-0.6), the system prioritizes expression features, including fluency of language organization (fluency index > 0.8) and accuracy of expression (accuracy > 0.75). In multi-person video conferencing scenarios, the system dynamically adjusts the focus of feature extraction based on user roles. When the user is in the audience, the system primarily extracts comprehension-related cognitive features (weight 0.7). When the user transitions to the speaker role, the system quickly switches to an expression-focused extraction mode (weight 0.6). In group discussion settings, the system monitors the transition from listening to speaking, extracting changes in cognitive load (change threshold 0.3), expression preparation (preparedness threshold 0.6), and emotional regulation (regulation amplitude ±0.2). For example, as a user gradually adjusts from a tense state to a foreign language environment, the system can detect a relaxation trend in EEG features (trend slope <-0.1), an improvement in the naturalness of voice and intonation (improvement >20%), and an increase in cognitive processing fluency (increase >30%). When communicating in professional fields, the system pays special attention to the comprehension and expression of terminology, extracting changes in the user's cognitive features as they process professional concepts (sensitivity to change 0.8). Through this information extraction process that considers feature weights, the system captures a complete sequence of user intent features during cross-language interaction.
[0074] The complete intent feature sequence is used to output interaction intent. Based on the extracted feature sequence, the system devises an intent integration mechanism to convert feature information from different intervals into specific interaction intents. In formal meeting scenarios, the system integrates features from the cognitively dominant interval (weight 0.8) to identify critical moments requiring language assistance (assistance trigger threshold 0.7). For example, when the system detects that a user's difficulty understanding the other party's speech exceeds a threshold (0.75), it interprets this as an intent to refine content. In cross-language negotiation scenarios, the system combines features from the behaviorally dominant interval (weight 0.6) and the emotionally dominant interval (weight 0.4) to identify an intent to adjust the conversational rhythm. The system employs different intent recognition strategies for scenarios with varying language difficulty levels. For entry-level language proficiency levels, the system prioritizes identifying basic comprehension and expression intent (basic comprehension threshold 0.6); for advanced language proficiency levels, it focuses on identifying linguistic details and contextual understanding intent (detailed comprehension threshold 0.8). In remote collaboration scenarios, the system can identify interaction barriers caused by network latency (latency threshold >100ms) and distinguish between technical and linguistic factors contributing to comprehension difficulties. When users need to quickly switch between languages, the system predicts their language readiness by analyzing cognitive and behavioral changes in feature sequences (rate of change > 0.4). In informal social scenarios, the system adjusts the sensitivity of intent recognition based on the formality of the situation (0.8 for formal situations, 0.6 for informal situations), ensuring effective interaction while maintaining the naturalness of the conversation. Furthermore, the system can identify user fatigue (fatigue index > 0.7) and adjust intent recognition strategies appropriately during extended cross-language interactions. Through this multi-scenario, multi-level integration of intents, the system ultimately outputs a set of interaction intents that includes interaction type, priority, and execution sequence.
[0075] The system integrates information from the interaction intents generated in the above steps. Based on the interaction intent set generated in the above steps, which includes language assistance intents (such as comprehension of professional terminology and expression suggestions, with a priority weight of 0.8), pacing intents (such as speech rate adjustment and interaction timing, with a priority weight of 0.6), and state adaptation intents (such as fatigue adjustment and attention supplementation, with a priority weight of 0.4), the system establishes an information integration framework. When processing high-priority language assistance interaction intents, the system implements a rapid response mechanism. For example, if it detects that a user's difficulty understanding professional terminology in an international conference exceeds a threshold (0.75), the system prioritizes the term explanation intent and simultaneously schedules related context understanding intents (with a relevance greater than 0.6) for subsequent processing. When processing regular-priority pacing interaction intents, the system determines the assistance strategy based on scenario characteristics and user status. In a video conferencing scenario, when the user transitions from listener to speaker, the system can identify a 30% increase in the priority of the speech rate adjustment intent. For low-priority state adaptation interaction intents, the system analyzes their timing characteristics (duration greater than 2 minutes) to plan an appropriate response time. For example, if fatigue adjustment intent (fatigue index > 0.7) is detected during a long meeting, it will be processed when the meeting node switches. Through this hierarchical integration, the system forms a set of conversion rules based on intent type and scenario characteristics.
[0076] The conversion rules based on intent type and scenario characteristics are applied to information processing. For language assistance rules, the system activates a real-time processing channel to ensure timely interpretation of professional terminology and expression suggestions (response latency <100ms). For example, in business negotiations, when a rule triggers terminology comprehension assistance, the system prioritizes core vocabulary information (importance > 0.8) and simultaneously builds a terminology association network to facilitate subsequent comprehension. For rhythm control rules, the system implements a dynamic batch processing mechanism (batch processing window 200-500ms). In academic communication scenarios, the system adjusts the user's expression rhythm (rhythm coordination > 0.75) based on speech rate adjustment rules and adjusts the information input rate based on attention level rules (attention threshold 0.7). For state adaptation rules, the system employs a predictive processing strategy, pre-planning break prompts and attention recovery strategies (prediction window 5-10 minutes). In multi-person discussion scenarios, the system can simultaneously handle the different rule requirements of multiple users, for example, coordinating the speaker's expression assistance rules (priority 0.8) with the audience's comprehension assistance rules (priority 0.6). Through this multi-level rule processing, a structured processing result containing core information, auxiliary information and status information is formed.
[0077] Information is organized using the structured processing results containing these three types of information. The system implements a three-dimensional organization of information based on the hierarchical attributes of the processing results. For the core information layer, the system prioritizes key content (importance > 0.85) within the language assistance results by highlighting key points. For example, in real-time translation scenarios, the system prominently annotates explanations of professional terminology and key concept prompts, and establishes connections with other auxiliary information (association strength > 0.7). For the auxiliary information layer, the system constructs a contextual association network (network density > 0.6), logically organizing content such as pacing tips and expression suggestions. In technical discussion scenarios, the system systematically organizes auxiliary content such as connections between professional concepts (association threshold 0.65) and application scenario descriptions. For the status information layer, the system implements a dynamic update mechanism (update cycle 500ms) to adjust information organization strategies based on real-time changes in user status. For example, if the system detects a decrease in user attention (a decrease > 20%), it increases the prominence of key annotations (by 30%) and simplifies the presentation of non-core information. The system also incorporates a scenario-adaptive mechanism to adjust the organization of information in different interactive environments (environmental adaptability > 0.8). Through this multi-dimensional information organization, the system forms a complete information architecture encompassing core commands, auxiliary commands, and status feedback.
[0078] Interaction commands are output according to the complete information architecture. Based on the hierarchical organization of information, the system employs an adaptive command generation mechanism. When generating core commands, the system prioritizes timeliness and accuracy, transforming language-supporting information into intuitive interactive instructions (clarity index > 0.9). For example, when a user encounters expression difficulties (difficulty index > 0.7) in a cross-language discussion, the system generates a main command sequence consisting of expression templates and keyword prompts based on the core information layer. When generating auxiliary commands, the system employs a progressive strategy, transforming pacing control information into guiding instructions (guidance effectiveness index > 0.8). In team collaboration scenarios, the system can transform the content of the auxiliary information layer into personalized interaction prompts (personalization matching > 0.75) based on the needs and language proficiency of different roles. For example, it generates speech rate control suggestions for speakers (control accuracy ±10%) and key point tracking guides for listeners (tracking accuracy > 0.85). When generating status feedback commands, the system implements proactive command planning, transforming the content of the status information layer into regulatory instructions (regulatory effectiveness index > 0.7). For example, during long meetings, the system generates timely rest reminders and attention adjustment suggestions based on changes in user fatigue (trend slope > 0.1). During dynamic interactions, the system can adjust its command generation strategy in real time based on changing scenarios (strategy fitness > 0.8), ensuring the coherence and practicality of command sequences. Through this multi-layered command generation process, the system ultimately outputs a complete interactive command system consisting of a primary command sequence, auxiliary instruction sets, and state adjustment commands.
[0079] The method has at least the following beneficial effects:
[0080] 1. By constructing a multi-level feature extraction and fusion mechanism, the system achieves precise synchronous acquisition, noise reduction, and feature reconstruction of EEG and speech signals. The system utilizes high-precision clock synchronization circuits and dynamic timebase calibration to significantly reduce the sampling time difference between the two modal signals. Adaptive noise reduction and signal separation algorithms effectively improve the signal-to-noise ratio of the two signals. Through multidimensional feature mapping and modal conversion, accurate correspondence and complementary enhancement of bimodal features are achieved.
[0081] 2. A feature supplementation and optimization framework based on temporal correlation has been established, effectively addressing the issues of missing features and distribution drift. The system precisely locates missing regions by analyzing the temporal evolution of features; utilizes a multi-level reconstruction network to achieve high-quality feature restoration; and adaptively adjusts feature distribution based on a dynamic feature fusion mechanism. In practical applications, the system can promptly detect and supplement missing features caused by factors such as environmental noise and attention fluctuations, effectively improving the system's adaptability in complex environments.
[0082] 3. A hierarchical intent understanding and command generation mechanism has been designed to achieve scenario-adaptive interaction strategies. The system constructs a multi-level intent parsing framework based on interaction needs of varying priorities. Feature integration and rule conversion improve the accuracy of intent recognition. Based on scenario characteristics and user status, personalized interaction commands are dynamically generated. This adaptive interaction mechanism enables rapid response and accurate interaction in complex scenarios such as multi-person meetings and cross-language communication.
[0083] In order to implement the EEG-speech hybrid interaction method based on brain-computer AI headset corresponding to the above method embodiment, to achieve the corresponding functions and technical effects. Figure 2 , Figure 2 The following is a block diagram of a brain-wave and speech hybrid interaction device 200 provided in an embodiment of the present application. For ease of explanation, only the parts related to this embodiment are shown. The brain-wave and speech hybrid interaction device 200 provided in an embodiment of the present application includes:
[0084] The waveform output unit 201 is used to collect the bimodal signal of the AI headset, classify and process the bimodal signal, and output time series information; perform noise reduction and signal separation on the time series information, optimize and purify the processing results corresponding to the separation process, and output a pure waveform;
[0085] The pattern generation unit 202 is configured to perform a two-way analysis on the clean waveform and divide the feature dimensions, complete the modal conversion through the constructed mapping chain, and output a feature identifier; perform a time series analysis on the feature identifier to identify a set of key time series features, establish an aligned feature mapping structure to combine the modalities, and generate a correlation pattern;
[0086] The node labeling unit 203 is configured to perform path analysis on the association pattern and label the combined nodes, reconstruct features through a transformation chain, and output a representation structure; perform hierarchical fusion on the representation structure and extract fusion regions, and output fusion elements based on the fusion regions; supplement the fusion elements and locate the missing parts, and reconstruct the missing parts to output the complete content;
[0087] The element acquisition unit 204 is used to perform intent analysis on the complete content to obtain key elements, convert the key elements into interaction information and output interaction intentions; integrate information on the interaction intentions to obtain conversion rules, generate interaction commands according to the conversion rules, and complete the EEG-speech hybrid interaction based on the brain-computer AI headset.
[0088] Figure 3 This is a schematic diagram of the structure of a computer device provided in one embodiment of the present application. Figure 3 As shown, the computer device 3 of this embodiment includes: at least one processor 30 ( Figure 3Only one is shown in the figure), a memory 31 and a computer program 32 stored in the memory 31 and executable on the at least one processor 30, wherein the processor 30 implements the steps of any of the above method embodiments when executing the computer program 32.
[0089] The computer device 3 may be a computing device such as a smart phone, a tablet computer, a desktop computer, or a cloud server. The computer device may include but is not limited to a processor 30 and a memory 31. It will be understood by those skilled in the art that Figure 3 This is merely an example of the computer device 3 and does not constitute a limitation on the computer device 3 . The computer device 3 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device 3 may also include input and output devices, network access devices, etc.
[0090] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0091] In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as a hard drive or memory of the computer device 3. In other embodiments, the memory 31 may also be an external storage device of the computer device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device 3. Furthermore, the memory 31 may include both an internal storage unit of the computer device 3 and an external storage device. The memory 31 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 31 may also be used to temporarily store data that has been output or is about to be output.
[0092] In addition, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0093] An embodiment of the present application provides a computer program product. When the computer program product is run on a computer device, the computer device implements the steps in the above-mentioned various method embodiments when executing the computer program product.
[0094] In several embodiments provided in the present application, it is understood that each box in the flow chart or block diagram can represent a part of a module, program segment or code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which depends on the functions involved.
[0095] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0096] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application by those skilled in the art should be included within the scope of protection of this application.
Claims
1. A brain-computer AI headset-based EEG-speech hybrid interaction method, characterized in that: include: Collect the dual-modal signal of the AI headset, classify and process the dual-modal signal, and output timing information, including: synchronously collect EEG signals and voice signals through a high-precision sensor array to form the dual-modal signal; establish a time reference point based on the GPS timing module, and use FPGA clock synchronization circuit to achieve time alignment of the dual-modal signal; segment the voice signal and use an adaptive threshold mechanism to divide the voice activity area and the non-voice area; pre-process the EEG signal through the sliding window method and combine it with the Hjorth parameter for feature characterization; classify the dual-modal signal into pure voice segment, pure EEG segment, dual-modal activity segment and background noise segment, which is used to construct a time stamp and signal segment. The timing information of the quality index; performing noise reduction and signal separation processing on the timing information, optimizing and purifying the processing results corresponding to the separation processing and then outputting a pure waveform, including: using an adaptive filter to eliminate power frequency interference on the EEG signal, and combining independent component analysis to perform electromyographic artifact separation; implementing a blind source separation algorithm on the speech signal to eliminate environmental noise, and suppressing sudden interference through morphological filtering; using an improved wavelet threshold method to denoise the EEG signal and retain characteristic waveform details; performing spectral subtraction noise reduction processing on the speech signal, and combining a Wiener filter to enhance speech features; dynamically adjusting the noise reduction parameters of the timing information according to the signal quality index, and outputting the pure waveform with improved signal-to-noise ratio; The clean waveform is subjected to a two-way analysis and feature dimension division, and modal conversion is completed through a constructed mapping chain to output a feature identifier, including: performing wavelet packet decomposition on the EEG signal to obtain time-frequency features, and extracting intrinsic modal components in combination with empirical mode decomposition; extracting Mel-frequency cepstral coefficients and acoustic feature parameters from the speech signal; evaluating the feature dimension discrimination through the Fisher discriminant criterion and establishing a hierarchical feature selection strategy; outputting the feature identifier based on the time-frequency features, intrinsic modal components, Mel-frequency cepstral coefficients, acoustic feature parameters and the hierarchical feature selection strategy; performing time series analysis on the feature identifier to identify a key time series feature set, establishing an aligned feature mapping structure to combine modalities and generate a correlation pattern; Perform path analysis on the association pattern and mark the combination nodes, reconstruct the features through the transformation chain and output the representation structure; perform hierarchical fusion on the representation structure and extract the fusion area, and output the fusion element according to the fusion area; supplement the missing parts of the fusion element and locate the missing parts, reconstruct the data of the missing parts and output the complete content, including: establishing a feature integrity assessment model to obtain the missing parts of the elements corresponding to the fusion elements; the three types of missing parts corresponding to the missing parts include cognitive feature abnormalities, behavioral feature interruptions and basic feature deviations. For cognitive feature abnormalities, a variational autoencoder is used for reconstruction. For behavioral feature interruptions, a bidirectional gated recurrent unit network is used to predict the missing fragments based on the complete sequence before and after the interruption position. For basic feature deviations, adjacent features with a quality assessment score > 0.8 are selected as references, and their consistency with the overall feature distribution is ensured through feature reconstruction; a complete feature supplementation data set is formed through the joint reconstruction of multiple layers of features; the temporal continuity of the missing parts of the elements is restored according to the spatiotemporal interpolation algorithm; feature distribution alignment processing is performed on the missing parts of the elements to ensure the consistency of the reconstructed data, and the complete content containing complete temporal features and semantic associations is output; The complete content is intentionally parsed to obtain key elements, the key elements are converted into interaction information and the interaction intention is output; the interaction intention is information integrated to obtain conversion rules, and interaction commands are generated according to the conversion rules to complete the EEG-speech hybrid interaction based on brain-computer AI headphones.
2. The method according to claim 1, characterized in that The performing path analysis on the association pattern and marking the combined nodes, reconstructing the features through the transformation chain and outputting the representation structure, includes: A dynamic programming algorithm is used to find the optimal characteristic path for the association pattern, and the combination nodes of the temporal association are marked; Identify the key connection relationships of the combined nodes through a graph attention network and construct a feature transformation chain; Implementing a Markov model-based path optimization on the feature conversion chain to eliminate redundant feature connections; The tensor decomposition method is used to reconstruct the feature space corresponding to the feature transformation chain, maintain the topological relationship between features, and output a multi-dimensional representation structure containing principal component features and temporal relationships.
3. The method according to claim 1, characterized in that The step of hierarchically fusing the representation structures and extracting fusion regions, and outputting fusion elements according to the fusion regions, includes: Constructing a multi-scale feature pyramid structure for performing hierarchical feature fusion on the representation structure to obtain the fusion area; Calculating the feature area weights of the fusion area through the attention mechanism to identify the key fusion area; A region growing algorithm is used to expand the fusion boundary of the key fusion area, maintain feature continuity, and output fusion elements containing EEG-speech collaborative features.
4. The method according to claim 1, wherein The performing intent parsing on the complete content to obtain key elements, converting the key elements into interaction information and outputting interaction intent includes: Constructing a semantic network analysis model to identify key elements related to the intent of the complete content; Obtain the element association strength corresponding to the key elements related to the intention through a fuzzy reasoning algorithm to generate candidate interaction intentions; A reinforcement learning model is used to optimize the intention generation strategy corresponding to the candidate interaction intention, eliminate ambiguous understanding, and output interaction intentions that include language assistance, rhythm control, and state adaptation intentions.
5. A brain wave-speech hybrid interaction device, characterized in that: include: The waveform output unit is used to collect the dual-modal signal of the AI headset, classify and process the dual-modal signal, and output timing information, including: synchronously collecting EEG signals and voice signals through a high-precision sensor array to form the dual-modal signal; establishing a time reference point based on the GPS timing module, and using an FPGA clock synchronization circuit to achieve time alignment of the dual-modal signal; segmenting the voice signal, and using an adaptive threshold mechanism to divide the voice activity area and the non-speech area; pre-processing the EEG signal through a sliding window method, and combining it with the Hjorth parameter for feature characterization; classifying the dual-modal signal into pure speech segment, pure EEG segment, dual-modal activity segment and background noise segment, for constructing a time-domain waveform generator containing time information. The timing information of the time stamp and signal quality index; performing noise reduction and signal separation processing on the timing information, optimizing and purifying the processing results corresponding to the separation processing and outputting a pure waveform, including: using an adaptive filter to eliminate power frequency interference on the EEG signal, and combining independent component analysis to perform electromyographic artifact separation; implementing a blind source separation algorithm on the speech signal to eliminate environmental noise, and suppressing sudden interference through morphological filtering; using an improved wavelet threshold method to denoise the EEG signal and retain characteristic waveform details; performing spectral subtraction noise reduction processing on the speech signal, and combining a Wiener filter to enhance speech features; dynamically adjusting the noise reduction parameters of the timing information according to the signal quality index, and outputting the pure waveform with improved signal-to-noise ratio; A pattern generation unit is used to perform a two-way analysis on the pure waveform and divide the feature dimensions, complete the modal conversion through the constructed mapping chain, and output the feature identifier, including: performing wavelet packet decomposition on the EEG signal to obtain time-frequency features, and extracting the intrinsic modal components in combination with empirical mode decomposition; extracting Mel-frequency cepstral coefficients and acoustic feature parameters from the speech signal; evaluating the feature dimension discrimination through the Fisher discriminant criterion and establishing a hierarchical feature selection strategy; outputting the feature identifier based on the time-frequency features, intrinsic modal components, Mel-frequency cepstral coefficients, acoustic feature parameters and the hierarchical feature selection strategy; performing time series analysis on the feature identifier to identify a key time series feature set, establishing an aligned feature mapping structure to combine modalities and generate a correlation pattern; a node annotation unit, configured to perform path analysis on the association pattern and annotate the combined nodes, output a representation structure after reconstructing the features through a transformation chain; hierarchically fuse the representation structure and extract the fusion region, output the fusion element according to the fusion region; supplement the missing elements and locate the missing parts, reconstruct the data of the missing parts and output the complete content, including: establishing a feature integrity assessment model to obtain the missing parts of the elements corresponding to the fusion elements; the three types of missing parts corresponding to the missing parts include cognitive feature abnormalities, behavioral feature interruptions and basic feature deviations; for cognitive feature abnormalities, a variational autoencoder is used for reconstruction; for behavioral feature interruptions, a bidirectional gated recurrent unit network is used to predict the missing segments based on the complete sequence before and after the interruption position; for basic feature deviations, adjacent features with a quality assessment score > 0.8 are selected as references, and their consistency with the overall feature distribution is ensured through feature reconstruction; a complete feature supplementation data set is formed through the joint reconstruction of multiple layers of features; the temporal continuity of the missing parts of the elements is restored according to the spatiotemporal interpolation algorithm; feature distribution alignment processing is performed on the missing parts of the elements to ensure the consistency of the reconstructed data, and the complete content containing complete temporal features and semantic associations is output; An element acquisition unit is used to perform intent analysis on the complete content to obtain key elements, convert the key elements into interaction information and output interaction intentions; integrate information on the interaction intentions to obtain conversion rules, generate interaction commands according to the conversion rules, and complete EEG-speech hybrid interaction based on brain-computer AI headphones.
6. A computer device, characterized in that: The method comprises a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the method according to any one of claims 1 to 4 when executing the computer program.
Citation Information
Patent Citations
Electroencephalogram signal voice decoding method based on generative adversarial network
CN116364096A
Emotion perception interaction method and system of brain-computer interface earphone
CN119576121A