Electroencephalogram-voice mixed interaction method, device and equipment based on brain-computer AI earphone

By collecting and processing EEG and voice signals in brain-computer AI headsets, the precise synchronization and feature complementarity of dual-mode signals are achieved, solving the problem of feature distribution drift and missing in existing systems, and improving the real-time adaptation and interaction accuracy of hybrid interactive systems.

CN120086743AActive Publication Date: 2025-06-03XIAOZHOU TECH CO LTD

Patent Information

Application Number
CN202510556249.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-03
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

It is difficult for existing systems to achieve precise synchronous acquisition and alignment of EEG and voice signals, resulting in the drift and lack of feature distributions of hybrid interactive systems when changes in user cognitive state and environmental conditions, affecting real-time adaptation and online optimization.

Method used

Through a brain-computer AI headset-based method, dual-modal signals are collected, data classification and processing are performed, timing information is output, and pure waveforms are generated through noise reduction and signal separation processing. Then, dual-channel analysis and feature dimension division are performed, modal transformation is completed by building a mapping chain, association mode is generated, and output representation structure is output through path analysis and feature reconstruction, and interactive commands are finally generated.

Benefits of technology

It realizes the precise synchronization processing of EEG and speech signals and feature complementarity, improves interaction accuracy and robustness, solves the problems of feature distribution drift and missing, and enhances the system's real-time adaptability and interactive practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086743A_ABST
    Figure CN120086743A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of brain-computer interfaces, and discloses an electroencephalogram-voice mixed interaction method, device and equipment based on a brain-computer AI earphone. The method comprises the following steps: acquiring a bimodal signal of the AI earphone, performing data classification and processing on the bimodal signal, and outputting a pure waveform after output time sequence information is purified; performing two-way analysis on the pure waveform and dividing feature dimensions, outputting a feature identifier to perform time sequence analysis so as to identify a key time sequence feature set, establishing an alignment feature mapping structure combination mode, then generating an association mode to execute path analysis, and marking a combination node output representation structure; performing hierarchical fusion on the representation structure, extracting a fusion region, and outputting fusion elements according to the fusion region; and carrying out missing supplementation on the fusion elements, positioning missing parts, carrying out data reconstruction, outputting complete contents, carrying out intention analysis to obtain key elements, converting the key elements into interaction information, outputting an interaction intention, generating an interaction command, and completing electroencephalogram-voice mixed interaction based on the brain-computer AI earphone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of brain-computer interfaces, and particularly to an electroencephalogram (EEG)-voice hybrid interaction method, device, and equipment based on a brain-computer AI headset. Background Art

[0002] Brain-computer interfaces and voice interaction, as two important paradigms of human-computer interaction, each have their own characteristics but also have limitations. Although voice interaction is natural and intuitive, using it in public places will expose privacy and is easily interfered by environmental noise. Brain-computer interfaces can achieve silent interaction and directly decode user intentions, but the information transmission rate is low and it is easily affected by physiological interferences such as myoelectricity and electrooculogram. Deeply integrating these two interaction methods is expected to achieve complementary advantages and overcome their respective limitations.

[0003] However, due to significant differences in sampling frequency, signal-to-noise ratio, and time scale between EEG and voice signals, existing systems are difficult to achieve precise synchronous acquisition and alignment of bimodal signals. Traditional solutions mostly adopt the method of feature splicing after independent processing, and fail to fully explore the temporal correlation and complementarity between signals. In practical applications, the dynamic changes in the user's cognitive state and environmental conditions not only cause the feature distribution to drift, but also often result in feature loss, which poses a huge challenge to the real-time adaptation and online optimization of the hybrid interaction system. At the same time, how to accurately understand user intentions and generate adaptive interaction strategies has also become a key issue restricting the practical application of the system.

[0004] Therefore, there is an urgent need for a method to solve at least one of the above problems. Summary of the Invention

[0005] The embodiments of this application provide an electroencephalogram (EEG)-voice hybrid interaction method, device, and equipment based on a brain-computer AI headset. The method aims to solve the problems that due to significant differences in sampling frequency, signal-to-noise ratio, and time scale between EEG and voice signals, existing systems are difficult to achieve precise synchronous acquisition and alignment of bimodal signals. Traditional solutions mostly adopt the method of feature splicing after independent processing, and fail to fully explore the temporal correlation and complementarity between signals. In practical applications, the dynamic changes in the user's cognitive state and environmental conditions not only cause the feature distribution to drift, but also often result in feature loss, which poses a huge challenge to the real-time adaptation and online optimization of the hybrid interaction system. At the same time, how to accurately understand user intentions and generate adaptive interaction strategies has also become a key issue restricting the practical application of the system.

[0006] In the first aspect, the embodiments of this application provide an electroencephalogram (EEG)-voice hybrid interaction method based on a brain-computer AI headset, including:

[0007] Collect the bimodal signals of the AI headset, classify and process the bimodal signals, and output timing information; perform noise reduction and signal separation processing on the timing information, and optimize and purify the processing results corresponding to the separation processing and then output a pure waveform;

[0008] Perform two-way analysis on the pure waveform and divide the feature dimensions, complete modal conversion through the constructed mapping chain, and output feature identifiers; perform timing analysis on the feature identifiers to identify a set of key timing features, and generate an associated pattern after establishing an aligned feature mapping structure to combine the modes;

[0009] Perform path analysis on the associated pattern and label the combined nodes, reconstruct the features through the conversion chain and then output a representation structure; perform hierarchical fusion on the representation structure and extract the fusion region, and output fusion elements according to the fusion region; supplement the missing parts of the fusion elements and locate the missing parts, and perform data reconstruction on the missing parts to output the complete content;

[0010] Perform intention parsing on the complete content to obtain key elements, convert the key elements into interaction information and output an interaction intention; perform information integration on the interaction intention to obtain a conversion rule, and generate an interaction command according to the conversion rule to complete the electroencephalogram-voice hybrid interaction based on the brain-computer AI headset.

[0011] In a second aspect, the present application further provides an electroencephalogram-voice hybrid interaction device, including:

[0012] A waveform output unit, configured to collect the bimodal signals of the AI headset, classify and process the bimodal signals, and output timing information; perform noise reduction and signal separation processing on the timing information, and optimize and purify the processing results corresponding to the separation processing and then output a pure waveform;

[0013] A mode generation unit, configured to perform two-way analysis on the pure waveform and divide the feature dimensions, complete modal conversion through the constructed mapping chain, and output feature identifiers; perform timing analysis on the feature identifiers to identify a set of key timing features, and generate an associated pattern after establishing an aligned feature mapping structure to combine the modes;

[0014] A node annotation unit, configured to perform path analysis on the associated pattern and label the combined nodes, reconstruct the features through the conversion chain and then output a representation structure; perform hierarchical fusion on the representation structure and extract the fusion region, and output fusion elements according to the fusion region; supplement the missing parts of the fusion elements and locate the missing parts, and perform data reconstruction on the missing parts to output the complete content;

[0015] The element acquisition unit is used to perform intention parsing on the complete content to obtain key elements, convert the key elements into interaction information and output an interaction intention; perform information integration on the interaction intention to obtain a conversion rule, generate an interaction command according to the conversion rule, and complete the electroencephalogram (EEG)-voice hybrid interaction based on the brain-computer AI headset.

[0016] This method collects electroencephalogram (EEG) signals and voice signals (bimodal signals), performs data classification and processing, and outputs timing information (including timestamps and signal quality indicators). By performing noise reduction (such as adaptive filtering, wavelet threshold denoising) and signal separation (such as independent component analysis, blind source separation) on the timing information, a pure waveform is output. Perform time-frequency analysis (wavelet packet decomposition, Mel-frequency cepstral coefficient extraction) on the EEG and voice signals respectively, divide the feature dimensions, and complete modal conversion through a mapping chain to output feature identifiers. Identify the key timing feature set, establish an aligned feature mapping structure, and combine the EEG-voice modality to generate an association pattern. Perform path analysis (dynamic programming algorithm, graph attention network) on the association pattern, label the combined nodes, and output a representation structure after reconstructing the features; extract fusion elements through hierarchical fusion, supplement missing data, and output the complete content. Parse the key elements (semantic network analysis, fuzzy reasoning) in the complete content, generate an interaction intention, and generate an interaction command (such as a voice command, an EEG control command) after integration.

[0017] Through the hybrid processing of EEG and voice signals, the complementary signal characteristics (high intention accuracy of EEG, semantic richness of voice) are utilized to improve the interaction accuracy and robustness. Technologies such as adaptive filters and blind source separation effectively eliminate environmental noise and physiological artifacts, retain key feature waveforms (such as event-related potentials, voice formants), and improve the signal-to-noise ratio. Construct a timing-aligned mapping chain (such as hierarchical feature selection, tensor decomposition) to achieve cross-modal association of EEG-voice features and solve the ambiguity problem of single-modal intention parsing. Combine reinforcement learning to optimize the intention generation strategy, dynamically adapt to complex scenarios (such as noisy environments, user fatigue states), and enhance the real-time performance and adaptability of the interaction. It can be applied to fields such as medical rehabilitation (language assistance for stroke patients) and intelligent wearables (AR / VR interaction control), expanding the practical value of the brain-computer interface.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit this application. Brief Description of the Drawings

[0019] Figure 1 It is a schematic flowchart of the EEG-voice hybrid interaction method based on the brain-computer AI headset shown in the embodiments of this application;

[0020] Figure 2 It is a schematic structural diagram of the EEG-voice hybrid interaction device shown in the embodiments of this application;

[0021] Figure 3 This is a schematic structural diagram of the computer device shown in the embodiments of the present application. Detailed implementation manners

[0022] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0023] It should be understood that when used in the specification and claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0024] It should also be understood that the term "and / or" as used in the specification and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0025] As used in the specification and claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" depending on the context.

[0026] In addition, in the description of the specification and claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0027] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0028] The technical solutions of the embodiments of this application will be introduced below.

[0029] Brain-computer interface and speech interaction, as two important paradigms of human-computer interaction, each have their own characteristics but also have limitations. Although speech interaction is natural and intuitive, using it in public places will expose privacy and is easily interfered by environmental noise. The brain-computer interface can achieve silent interaction and directly decode the user's intention, but the information transmission rate is low and it is easily interfered by physiological signals such as electromyogram and electrooculogram. Deeply integrating these two interaction methods is expected to achieve complementary advantages and overcome their respective limitations.

[0030] However, achieving deep integration of electroencephalogram-speech signals faces many technical challenges. Due to significant differences in sampling frequency, signal-to-noise ratio, and time scale between electroencephalogram and speech signals, it is difficult for existing systems to achieve precise synchronous acquisition and alignment of bimodal signals. Traditional solutions mostly adopt the method of feature splicing after independent processing, and fail to fully explore the temporal correlation and complementarity between signals. In practical applications, the dynamic changes of the user's cognitive state and environmental conditions not only cause the feature distribution to drift, but also often result in feature loss, which brings great challenges to the real-time adaptation and online optimization of the hybrid interaction system. At the same time, how to accurately understand the user's intention and generate adaptive interaction strategies has also become a key issue restricting the practical application of the system.

[0031] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for electroencephalogram-speech hybrid interaction based on a brain-computer AI headset provided by an embodiment of this application. The method for electroencephalogram-speech hybrid interaction based on a brain-computer AI headset in the embodiments of this application can be applied to computer devices, including but not limited to devices such as smart phones, laptop computers, tablet computers, desktop computers, physical servers, and cloud servers. As Figure 1 shown, the method for electroencephalogram-speech hybrid interaction based on a brain-computer AI headset in this embodiment includes steps S101 to S104, which are described in detail as follows:

[0032] Step S101, collect the bimodal signals of the AI headset, classify and process the bimodal signals, and output timing information; perform noise reduction and signal separation processing on the timing information, and optimize and purify the processing results corresponding to the separation processing and then output the pure waveform.

[0033] Specifically, first, locate the time reference by collecting the bimodal signals of the AI headset.

[0034] In some embodiments, the classifying and processing the bimodal signals to output timing information includes: synchronously collecting electroencephalogram (EEG) signals and voice signals through a high-precision sensor array to form the bimodal signals; establishing a time reference point based on a GPS timing module, and using an FPGA clock synchronization circuit to achieve time alignment of the bimodal signals; performing segmentation processing on the voice signals, and using an adaptive threshold mechanism to divide the voice activity region and the non-voice region; preprocessing the EEG signals by a sliding window method, and performing feature characterization in combination with Hjorth parameters; classifying the bimodal signals into pure voice segments, pure EEG segments, bimodal activity segments, and background noise segments for constructing the timing information including timestamps and signal quality indicators.

[0035] In practical applications, the AI headset synchronously collects the user's EEG signals and voice signals through a high-precision sensor array. The sampling frequency of the EEG signals is set at 1000 Hz to ensure accurate capture of the rapid changes in brain nerve activities, and the signal bandwidth range is 0.1 - 100 Hz, mainly including frequency bands such as δ, θ, α, β, and γ. The sampling frequency of the voice signals is 16 kHz, and 24-bit quantization accuracy is adopted to ensure the complete retention of sound details. The time reference point of the bimodal signals is established through a hardware trigger mechanism, and the system uses a GPS timing module to provide a unified timestamp, with a time accuracy better than 1 ms. To solve the synchronization problem in the signal acquisition process, a high-precision clock synchronization circuit based on FPGA is designed to ensure that the sampling clock deviation of the two channels of signals is within 100 ns. At the signal source, active shielding technology and differential amplification circuits are used to increase the signal-to-noise ratio by at least 45 dB, effectively suppressing the influence of power frequency interference and environmental noise. After the above acquisition and synchronization processes, a bimodal synchronous sampling data stream is formed, and each sampling point corresponds to a unique timestamp mark.

[0036] Classify and process the acquired dual-modal synchronous sampling data stream. The system first aligns and segments the signals in time. Considering the characteristics of human speech interaction, the signals are segmented with a basic unit of 200 ms, and there is a 50% overlap rate between each unit, which not only ensures the continuity of the signals but also facilitates subsequent feature extraction and analysis. During the classification process, an adaptive threshold mechanism is introduced. By calculating the short-time energy and zero-crossing rate of the signals, the speech activity region and the non-speech region are initially divided. For electroencephalogram (EEG) signals, the sliding window method is used for preprocessing. The window length is set to 2 seconds, and the step size is 1 second. The Hjorth parameters are combined to characterize the signal segments. Through a multi-level data classification strategy, the system classifies the signals into four categories: pure speech segments, pure EEG segments, dual-modal activity segments, and background noise segments. Among them, the recognition threshold for dual-modal activity segments is obtained through statistical analysis of a large amount of experimental data. Generally, it is required that the normalized energy values of both signals exceed 85% of the set threshold. After the classification process, a labeled multi-modal data sequence is formed, and each data segment contains clear type identification and time range information.

[0037] Construct a temporal information output structure based on the formed labeled multi-modal data sequence. The system normalizes the data sequence. Each sequence unit contains key information such as time stamps, signal type identification, signal quality indicators, and feature vectors. The signal quality indicators are calculated by comprehensively evaluating factors such as signal-to-noise ratio, baseline drift degree, and electromyogram (EMG) artifact intensity, and are represented by normalized values. The feature vectors contain multi-dimensional feature parameters in the time domain and frequency domain, such as the mean, variance, skewness, kurtosis, and main frequency components of the signals. To facilitate subsequent processing, the system constructs a multi-level index for the temporal information. The main index is established based on time stamps, and the secondary indexes are established based on signal types and quality indicators. This multi-dimensional index structure significantly improves the data retrieval and processing efficiency, enabling the system to locate and extract signals within a specific time period or of a specific type within milliseconds. In practical applications, this temporal information structure supports core functions such as subsequent feature extraction, pattern recognition, and interaction intention analysis, providing a reliable data foundation for the entire hybrid interaction system. The data organized in time sequence generates a storage volume of approximately 2 MB per second. The system adopts a data management strategy based on a circular buffer to keep the data of the most recent 10 minutes accessible at any time, while historical data is written to storage asynchronously for persistent storage. Finally, a normalized temporal data stream is output, which contains complete signal features, classification labels, and temporal correlation information.

[0038] Exemplarily, the denoising and signal separation processing of the timing information, and optimizing and purifying the processing result corresponding to the separation processing and then outputting a pure waveform, includes: using an adaptive filter to eliminate power frequency interference from the electroencephalogram (EEG) signal, and combining independent component analysis to separate electromyogram (EMG) artifacts; implementing a blind source separation algorithm to eliminate ambient noise from the speech signal, and suppressing sudden interference through morphological filtering; using an improved wavelet threshold method to denoise the EEG signal and retaining the details of the characteristic waveform; performing spectral subtraction denoising processing on the speech signal, and combining a Wiener filter to enhance speech features; dynamically adjusting the denoising parameters of the timing information through a signal quality index, and outputting the pure waveform with an improved signal-to-noise ratio.

[0039] Perform denoising and signal separation processing on the timing information output from the above steps. The system first extracts the signal features of each time period based on the classification labels and quality indicators in the timing information flow. For the timing segments with a quality indicator lower than 0.7, the system adds preprocessing steps such as baseline drift correction and outlier detection. Differentiated processing is adopted for different signal types identified in the timing information: for the EEG signal segment, combined with the frequency domain characteristic parameters in the timing correlation information, a 5-layer wavelet decomposition structure is used, and the db4 wavelet basis function is selected for multi-resolution analysis. The decomposed frequency bands correspond to the δ, θ, α, β, and γ bands respectively. The system dynamically adjusts the wavelet threshold according to the signal quality indicator in the timing information. When the quality indicator is greater than 0.8, the soft threshold method is used, and when it is lower than 0.8, the hard threshold method is used for coefficient processing. For the speech signal segment, the system uses the short-time energy feature in the timing information and cooperates with an improved Wiener filtering algorithm for denoising. The filter parameters are dynamically adjusted according to the signal variance feature in the timing information, and the processing window varies within the range of 16 - 32 ms, and the overlap rate is maintained at 50%. For the bimodal activity segment, the system establishes a joint processing framework according to the time correspondence relationship in the timing information, identifies the coupling features between signals by calculating the normalized cross-correlation function, and uses tensor decomposition technology to achieve the preliminary separation of signals. During the entire processing process, the system continuously tracks the timestamps in the timing information to ensure that the processed signal segments strictly correspond to the original timing structure. Through this series of targeted denoising and separation processes, a separation result containing the main signal components and residual noise is obtained, where more than 90% of the effective information of the main signal components is retained, and the energy of the residual noise part is lower than 15% of the original signal.

[0040] Optimize and purify the separation results. The system establishes a multi-level optimization structure according to the characteristics of the main signal components and residual noise. First, for the aliased part remaining in the main signal components, independent component analysis technology is used for fine separation, and a separation model with 15 independent components is constructed. During the separation process, an improved FastICA algorithm is adopted to project the signal into a high-dimensional feature space. By analyzing the signal energy distribution characteristics obtained in the first stage, the optimal projection dimension is determined. In the feature space, the system combines the negative entropy maximization criterion to optimize the separation matrix, adopts an adaptive learning rate strategy with an initial value set to 0.01, and dynamically adjusts it according to the energy level of the residual noise until the convergence threshold reaches 10⁻ 6 . For the separated independent components, the system calculates statistical features such as kurtosis, sample entropy, and Hilbert spectrum, and matches them with the signal features identified in the first stage. At the same time, according to the interference features detected in the first stage, the system establishes an adaptive notch filter array to accurately suppress power frequency interference and its harmonics. The center frequency of the notch filter tracks the fluctuations of the power grid frequency in real time through phase-locked loop technology, and the bandwidth is automatically adjusted within the range of 0.5 - 2 Hz according to the interference characteristics. During the optimization process, the system continuously evaluates the improvement degree of the signal, improves the signal-to-noise ratio of the EEG signal from 15 dB in the first stage to 30 dB, and improves the signal-to-noise ratio of the voice signal from 20 dB to 35 dB, while maintaining the main feature structure of the signal. After this series of optimization and purification processes, a highly purified signal result is obtained, whose feature structure is consistent with the main signal components separated in the first stage, but the quality index is significantly improved.

[0041] Construct a pure waveform based on the highly purified signal results. The system designs an adaptive reconstruction algorithm. First, reconstruction parameters are selected according to the purification degree of the signal. For signal segments with a signal-to-noise ratio above 30 dB, Hilbert transform is used to directly obtain instantaneous amplitude and phase information; for parts with a lower signal-to-noise ratio, soft threshold processing in the wavelet domain is performed first. During the reconstruction process, according to the signal characteristics optimized in the second stage, the system designs an adaptive zero-phase digital filter bank. For EEG signals, band-pass filtering of 0.5 - 45 Hz is performed, and for speech signals, band-pass filtering of 300 - 3400 Hz is performed. The order of the filter is dynamically adjusted according to the purification degree of the signal. To handle the edge effect during the filtering process, the system analyzes the local characteristics of the signal obtained in the second stage and uses an improved wavelet packet transform technology for adaptive boundary processing. The extension length is dynamically adjusted according to the local variation characteristics of the signal, maintaining between 5% and 10% of the signal length. To ensure the continuity of the reconstructed signal in the time domain, the system combines the signal quality assessment results in the second stage and adopts an improved cosine function smoothing strategy at the connection of adjacent data segments. The length of the transition interval is adaptively set according to the change rate of the signal. For high-quality signal segments, a shorter transition interval (about 10 ms) is used, and for lower-quality signal segments, a longer transition interval (about 20 ms) is used. The finally output pure waveform achieves a high signal-to-noise ratio (EEG signal > 50 dB, speech signal > 60 dB), has complete time-frequency characteristics, continuous phase characteristics, and all indicators are better than the processing results in the second stage.

[0042] Step S102, perform dual-channel analysis on the pure waveform and divide the feature dimensions, complete the modal conversion through the constructed mapping chain, and output the feature identifier; perform timing analysis on the feature identifier to identify the key timing feature set, and generate the association pattern after establishing the alignment feature mapping structure combination mode.

[0043] Specifically, perform dual-channel feature analysis on the pure waveform output in the above steps. The system makes full use of three key characteristics of the signal processed in the above steps: high signal-to-noise ratio (EEG > 50 dB, speech > 60 dB) ensures the accuracy of feature extraction, complete time-frequency characteristics support the reliability of multi-scale analysis, and continuous phase characteristics ensure the consistency of timing features. Taking the cross-language communication scenario as an example, when the user uses an AI headset to talk with a foreign language user, the system can accurately capture the EEG features during the user's speech expression and language understanding process based on these high-quality waveform characteristics.

[0044] In some embodiments, performing two-way analysis on the pure waveform and dividing the feature dimensions, and completing modal conversion through the constructed mapping chain to output feature identifiers, including: performing wavelet packet decomposition on the electroencephalogram (EEG) signal to obtain time-frequency features, and extracting intrinsic mode components by combining empirical mode decomposition; extracting Mel-frequency cepstral coefficients and acoustic feature parameters from the speech signal; evaluating the discriminability of feature dimensions through the Fisher discriminant criterion to establish a hierarchical feature selection strategy; and outputting the feature identifiers according to the time-frequency features, intrinsic mode components, Mel-frequency cepstral coefficients, acoustic feature parameters, and hierarchical feature selection strategy.

[0045] In the EEG signal channel, the wavelet packet decomposition technique is used to perform time-frequency analysis on the signal. The sym4 wavelet basis is selected for 6-layer decomposition to obtain the energy distribution characteristics of 64 frequency bands. For each frequency band, the relative energy ratio, instantaneous frequency center, and band entropy value are calculated to construct an initial feature matrix. At the same time, the empirical mode decomposition method is used to extract the intrinsic mode features of the signal. The first 5 intrinsic mode functions are selected, and their instantaneous frequencies and amplitude envelopes are calculated. In the speech signal channel, the system uses Mel-frequency cepstral coefficients (MFCCs) for feature extraction, sets the frame length to 25 ms and the frame shift to 10 ms, and extracts 13-dimensional basic feature parameters. At the same time, acoustic features such as the fundamental period, formant frequency, and spectral entropy of the speech are calculated. This two-way feature analysis makes full use of the high-quality characteristics of the pure waveform, and obtains a feature dimension set of 128 dimensions in total, including the time domain, frequency domain, and time-frequency domain, with 64 EEG features and 64 speech features.

[0046] Construct the 128-dimensional feature dimension set of the dual modality into a feature network. The system first processes the 64-dimensional feature subsets of EEG and speech respectively, using a hierarchical feature selection strategy. Based on the Fisher discriminant criterion, the discriminability of each feature dimension is calculated, and importance rankings are established for the extracted EEG features (such as frequency band energy distribution, modal function features) and speech features (such as MFCC coefficients, acoustic features) respectively. The feature dimensions with a discriminability greater than 0.6 are selected as the main features, including the key frequency band energy and typical modal features in the EEG signal, and the core acoustic parameters in the speech signal. For these main feature dimensions, the principal component analysis method is used for dimensionality reduction, and the principal components with a cumulative contribution rate reaching 85% are retained. For the auxiliary feature dimensions with a discriminability between 0.3 and 0.6, the locally linear embedding algorithm is used for non-linear dimensionality reduction, and the dimension is compressed to 1 / 3 of the original dimension. At the same time, the system establishes a feature correlation matrix to identify the redundant relationships between features. When the correlation coefficient exceeds 0.8, the feature dimension with a larger information entropy is retained. Taking cross-language understanding as an example, the system can extract and organize 45 most discriminative feature nodes from the initial 128-dimensional features to form a compact feature network structure.

[0047] Perform modal conversion through the 45-node feature network. The system processes the main feature nodes (discrimination degree > 0.6) and auxiliary feature nodes (discrimination degree 0.3 - 0.6) separately, and establishes a hierarchical feature mapping framework. First, standardize these feature nodes, and use the z-score method to make the feature distribution satisfy zero mean and unit variance. Based on the standardized features, use the deep canonical correlation analysis method to learn the inter-modal mapping relationships of the main feature group and the auxiliary feature group respectively. To enhance the robustness of the feature mapping, introduce a regularization constraint, set the regularization coefficient to 0.01, and solve the optimal mapping matrix through an alternating optimization algorithm. In practical applications, when users encounter expression obstacles during cross-language communication, the system can, based on the established feature network, perform bidirectional mapping between the users' electroencephalogram cognitive features and speech expression features to achieve more accurate cross-language interaction support. After optimization and iteration, the system has established a stable feature conversion mechanism and formed a complete set of conversion parameters including main mapping relationships and auxiliary mapping relationships.

[0048] Output feature identifiers according to the set of feature conversion parameters. The system constructs identifier generation networks for the main mapping relationship and the auxiliary mapping relationship respectively, using a multi-layer perceptron structure. The main mapping channel uses a three-layer network structure with 32, 16, and 8 nodes respectively, uses the ReLU activation function, and sets the dropout rate to 0.3; the auxiliary mapping channel adopts a two-layer structure with 16 and 8 nodes. During the training process, the system performs weighted fusion according to the importance of different mapping relationships, uses batch normalization technology to improve the convergence performance of the network, sets the initial learning rate to 0.001, and dynamically adjusts it using the cosine annealing strategy. To adapt to the characteristics of different language learners, the system constructs a dynamic identifier mapping mechanism. For example, in the scenario of international business talks, the system can accurately identify and convert the users' language understanding and expression features according to the mapping relationships provided by the set of conversion parameters to provide real-time support for cross-language communication. The finally output feature identifiers contain 8 key dimensions, of which 5 come from the main mapping channel and 3 come from the auxiliary mapping channel, jointly constituting a stable and reliable feature representation system.

[0049] Perform temporal analysis on the feature identifiers output from the above steps. The system performs temporal processing based on the 8-dimensional feature identifiers output from the above steps, including 5 dimensions of the main mapping channel and 3 dimensions of the auxiliary mapping channel. In a cross-language communication scenario, when the user uses an AI headset for real-time conversations, the system creates a time window for each conversation segment, with the window length set to 2 seconds and the sliding step size set to 500 ms. For the main dimension features (5 dimensions), calculate their temporal stability indicators, including the smoothness (threshold 0.15) and persistence (at least 3 consecutive windows) of the feature trajectories; for the auxiliary dimension features (3 dimensions), focus on analyzing their mutation characteristics (threshold 0.3) and transformation patterns. At the same time, introduce an improved dynamic time warping algorithm, and adopt a two-way search strategy to compensate for the temporal misalignment between features of different dimensions, with the warping window set to 100 ms. The system also establishes a hierarchical attention mechanism, and sets different attention thresholds based on the stability and distinctiveness of the features (0.7 for the main dimension, 0.5 for the auxiliary dimension), and identifies three typical patterns of feature changes: the gradual change process of stable features, the mutation points of significant features, and the co-variation of combined features. These feature change patterns and their time points constitute the key temporal feature set.

[0050] Establish an alignment mechanism around the temporal feature set. The system constructs corresponding alignment strategies based on the three identified feature change patterns. For the gradual change process of stable features, use sliding correlation analysis to calculate the optimal matching interval for feature gradual change, requiring the correlation coefficient to be greater than 0.75; for the mutation points of significant features, use precise timestamp alignment, allowing a maximum time deviation of 30 ms; for the co-variation of combined features, use the dynamic programming algorithm to find the optimal temporal mapping relationship, with the mapping quality threshold set to 0.8. At the local alignment level, the system performs fine alignment on the sequences within 200 ms before and after each feature change point. For the main dimension features, use the maximum mutual information criterion (threshold 0.7), and for the auxiliary dimension features, use the improved dynamic programming algorithm (warping window 50 ms). At the global alignment level, the system constructs a weighted sequence matching graph according to the importance of different feature change patterns, and achieves global optimization through the shortest path algorithm. To handle the uncertainty near the feature mutation points, the system also dynamically adjusts the alignment window size according to the stability indicators of the features. Features with high stability use a smaller alignment window (±100 ms), while features with lower stability use a larger window (±200 ms). Through this series of targeted alignment processes, an alignment feature mapping structure containing multi-scale temporal relationships is formed.

[0051] Perform modal combination using the described alignment feature mapping structure. The system designs an adaptive feature fusion strategy according to the temporal relationship in the mapping structure. For the gradual change interval of stable features (alignment quality > 0.9), a smoothing fusion mechanism based on gated recurrent units is adopted, with the feature retention rate set at 90%, and the gating parameter is controlled by the stability index of the features; for mutation feature points (alignment quality 0.7 - 0.9), an attention-enhanced instantaneous feature combination is used, with the attention weight threshold at 0.6; for co-varying regions (alignment quality < 0.7), non-linear combination of features is achieved through a multi-layer perceptron, with the combination threshold at 0.5. The system integrates a feature selection mechanism based on alignment quality, and when the confidence of the local alignment result is low, the fusion weight of the corresponding feature will be automatically reduced. In practical applications, this alignment-quality-based adaptive fusion mechanism can accurately capture the dynamic feature changes of users in cross-language communication. For example, when the system detects that the mutation point of speech features is highly aligned with the significant change of EEG features, it will strengthen the feature combination weight at this moment; while for time periods with low alignment quality, it relies more on features with high stability to maintain the reliability of modal combination. After this series of fusion processes, a sequence of dynamically aligned feature combinations is formed.

[0052] Generate association patterns following the described sequence of dynamically aligned feature combinations. The system constructs a hierarchical pattern extraction network based on different types of feature changes in the combination sequence. In the temporal feature extraction layer, dedicated processing units are designed for the three patterns of gradual change, mutation, and co-variation respectively. For gradual change features, a bidirectional long short-term memory network (with 64 units) is used to extract long-term trends, for mutation features, a convolutional network with residual connections is used to capture local features, and for co-varying features, an attention mechanism is used for feature enhancement. In the pattern recognition layer, the system dynamically adjusts the weight contributions of different patterns according to the temporal alignment quality of the feature combinations. A higher weight (0.5) is assigned to the feature combinations in the high-quality alignment region (alignment error < 50ms), a medium weight (0.3) is assigned to the medium-quality region (error 50 - 100ms), and the weight is reduced (0.2) for regions with low alignment quality (error > 100ms). The association coding layer integrates different types of patterns into a unified representation form through a variational autoencoder structure (coding dimension 32, reconstruction error threshold 0.1). In practical applications, the system can identify a complete chain of interaction patterns from the continuous conversations of users. For example, when the pattern evolution of "difficult to understand - attempt to understand - breakthrough in understanding" is recognized, the system can not only track the feature changes at each stage, but also determine the critical moments and triggering conditions of pattern conversion based on the temporal alignment relationship of the feature combinations. The finally output association patterns contain complete temporal evolution information, retaining both the dynamic details of feature changes and reflecting the overall laws of pattern conversion.

[0053] Step S103: Perform path analysis on the association pattern and label the combined nodes. After reconstructing the features through the conversion chain, output the representation structure; perform hierarchical fusion on the representation structure and extract the fusion region, and output the fusion elements according to the fusion region; supplement the missing parts of the fusion elements and locate the missing parts, and perform data reconstruction on the missing parts to output the complete content.

[0054] Specifically, perform path analysis on the association pattern output in the above step. The system constructs a pattern conversion path diagram based on the complete temporal evolution information recorded in the association pattern, including the dynamic details of feature changes (such as the state transition process) and the overall law of pattern conversion (such as the conversion trigger condition). For each identified interaction pattern chain, the system first extracts its complete evolution trajectory in the feature space. In the interval dominated by gradual features (such as the gradually changing cognitive load identified in the above step), the piecewise linear fitting method is adopted, and the fitting mean square error threshold is set to 0.1. When the fitting error exceeds the threshold, path nodes are set; in the interval dominated by abrupt features (such as the moment of understanding breakthrough captured in the above step), calculate the first-order difference of the feature. When the difference value exceeds 2.5 times the standard deviation, mark the jump point in the feature space as a key node; in the co-variation interval (such as the synchronous change of multi-modal features located in the above step), determine the marking position by calculating the feature co-degree (using the Pearson correlation coefficient, the threshold is set to 0.8). In the cross-lingual communication scenario, the system can convert the pattern chain "difficult to understand - attempt to understand - understanding breakthrough" identified in the above step into a specific sequence of path nodes, where each node includes: timestamp (precision 1ms); feature state vector (32 dimensions); conversion type identifier (gradual / abrupt / co-variation, confidence threshold 0.75); trigger condition parameters (feature threshold, time window, duration); through this path analysis based on temporal evolution, the system forms a set of path node collections including conversion types, trigger conditions, and evolution laws.

[0055] Construct a feature structure based on the set of path nodes including conversion types, trigger conditions, and evolution laws. The system designs corresponding feature organization frameworks for each type of node according to the conversion type of the node. For nodes with a gradual change conversion type identified in path analysis (such as cognitive load change points), the system establishes a gradual change feature subspace to describe the continuous change process of features. The subspace dimension is 16, and principal component analysis is used to ensure an energy retention rate of 95%; for nodes with a mutation conversion type (such as understanding breakthrough points), a high-dimensional feature subspace is constructed with a dimension of 24, and a kernel function mapping is used to capture the jump information of features; for nodes with a collaborative conversion type (such as multimodal synchronization points), a collaborative feature subspace is set with a dimension of 20 to express the interaction relationship of features. In terms of temporal organization, the system constructs a directed feature graph according to the trigger conditions and evolution laws of the nodes: direct connections are established between nodes with a strong trigger condition (correlation degree > 0.8), and transitional connections are established through intermediate states between nodes with a weak trigger condition (correlation degree 0.5 - 0.8). At the same time, a feature propagation mechanism based on the evolution law is introduced to enable features to flow in the network according to the evolution law determined by path analysis. Through this series of structured processing, a multi-level network structure reflecting the feature evolution law is formed.

[0056] Reconstruct features through the multi-level network structure. The system designs differentiated reconstruction strategies for different types of node connection relationships in the network. For directly connected node pairs (derived from strong trigger conditions), a high-precision feature reconstruction mechanism is adopted, including an encoding-decoding dual-path structure. Each path uses a five-layer neural network for feature processing (layer dimension configuration: 32 - 64 - 128 - 64 - 32), and residual connections are added between key layers to prevent information loss. For transitional connected node pairs (derived from weak trigger conditions), a smoothing reconstruction method based on bidirectional LSTM is used to capture the temporal dependence relationship of features through the hidden layer state (dimension 128). During the reconstruction process, the system dynamically adjusts parameters according to the position and connection attributes of the nodes in the network: core nodes use a larger learning rate (0.01) for rapid convergence, and transitional nodes use a smaller learning rate (0.005) to ensure stability. At the same time, an attention-based feature enhancement mechanism is implemented, and the attention weights are dynamically adjusted according to the connection strength of the features in the original network. For the temporal evolution of features, the system designs a forward prediction module, which uses a causal convolutional network (convolution kernel sizes are 3, 5, 7 respectively) to capture temporal patterns at different scales, and optimizes parameters by comparing the consistency between the prediction results and the actual feature sequence (requirement similarity > 0.85). During the entire reconstruction process, the system simultaneously optimizes the feature reconstruction accuracy (average error < 0.1) and temporal coherence (adjacent frame correlation degree > 0.8), and obtains a reconstructed feature set that not only retains the characteristics of the network structure but also has good generalization ability through multiple rounds of iteration (typical iteration times 500 - 1000).

[0057] In some embodiments, performing path analysis on the association pattern and annotating combined nodes, and outputting a representation structure after reconstructing features through a transformation chain, includes: using a dynamic programming algorithm for the association pattern to find an optimal feature path and annotating combined nodes with temporal associations; identifying key connection relationships of the combined nodes through a graph attention network to construct a feature transformation chain; performing path optimization based on a Markov model on the feature transformation chain to eliminate redundant feature connections; using a tensor decomposition method to reconstruct the feature space corresponding to the feature transformation chain, maintaining the topological relationship between features, and outputting a multi-dimensional representation structure including principal component features and temporal relationships.

[0058] Generating a representation structure according to the reconstructed feature set that retains the network structure. The system constructs a hierarchical feature fusion network to process features reconstructed from different connection types respectively. For features reconstructed from direct connections, a priority processing mechanism is adopted, and their dominant position in the representation structure is ensured through attention strengthening, with the attention weight range being 0.6 - 0.8; for features reconstructed from transitional connections, a progressive fusion strategy is used, with the weight range being 0.2 - 0.4, to achieve smooth feature transition. The system designs a multi-layer feature verification mechanism based on network topology. The first layer performs feature similarity verification by calculating the Euclidean distance between the reconstructed feature and the original feature (threshold < 0.15); the second layer performs structural consistency verification to ensure that the reconstructed feature retains the key connection relationships in the original network (connection retention rate > 90%); the third layer performs temporal coherence verification to check the smoothness of feature evolution (change rate of features at adjacent time points < 0.2). In a cross-language communication scenario, this network structure-based representation method can accurately describe the evolution of the user's cognitive state. For example, when it is detected that the "understanding breakthrough" node is activated, the system first intensively processes the reconstructed feature of this node (weight 0.8), and at the same time supplements details of state transition through adjacent transitional connection nodes (weight 0.3), such as the gradual decrease of cognitive load and the steady increase of attention level. Through this multi-level feature fusion and verification mechanism, the representation structure finally generated by the system not only maintains the clear positioning of state transition but also realizes the continuous expression of state evolution.

[0059] Hierarchically fuse the representation structure output by the above steps. The system makes full use of the high-weight features directly connected (original weight 0.6 - 0.8) and the low-weight features connected through transitions (original weight 0.2 - 0.4) in the representation structure to construct a multi-level fusion system. For the core feature region formed by direct connection, a deep feature fusion strategy is adopted, and feature extraction is performed through three residual blocks. Each residual block contains two convolutional layers with weight inheritance (inheritance weight ratio 0.7) and a skip connection. To maintain the clear localization characteristics of the original representation structure, the system introduces a self-attention mechanism into the residual block, and the attention weights are initialized based on the original weight distribution. For the gradual feature region formed by transitional connection, the original smooth transition characteristics are inherited, and a lightweight feature fusion network is used, which includes a convolutional layer with inheritance weight and an adaptive pooling layer. The system sets a feature aggregation window based on the original temporal granularity for different regions' temporal characteristics. A smaller window (100ms) is used in the core region to retain detailed features, and a larger window (300ms) is used in the gradual change region to obtain trend features. In a cross-language communication scenario, when the "understanding breakthrough" state is recognized, the system can enhance and smooth the high-weight feature regions (such as sudden changes in cognitive load) and low-weight feature regions (such as gradual changes in attention) in the original representation structure respectively. Through this hierarchical fusion process, a set of fusion regions with clear hierarchical attributes is obtained, including the feature combinations and temporal relationships inherited from the original representation structure.

[0060] In some embodiments, hierarchically fusing the representation structure and extracting the fusion region, and outputting the fusion elements according to the fusion region includes: constructing a multi-scale feature pyramid structure for hierarchically fusing the features of the representation structure to obtain the fusion region; calculating the feature region weights of the fusion region through an attention mechanism to identify the key fusion region; using a region growing algorithm to expand the fusion boundary of the key fusion region to maintain feature continuity, and outputting the fusion elements including the electroencephalogram-speech collaboration features.

[0061] Integrate the fusion region with hierarchical attributes, its included feature combinations, and temporal relationships into a feature framework. The system designs a hierarchical feature organizational structure based on the hierarchical characteristics of the fusion region, the inherited feature combinations, and temporal relationships. At the top layer of the framework, a feature integration module for the core region is deployed, which uses a graph neural network to process feature transmission between high-weight nodes, and the node update method is configured according to the weight distribution in the fusion region. At the middle layer of the framework, a feature transmission network for processing the gradual change region is constructed, and a gated recurrent unit is used to achieve progressive fusion of features, and the gating parameters inherit the temporal characteristics in the fusion region. At the bottom layer of the framework, a feature aggregation module is set up to perform preliminary integration of features at different levels through the inherited weight ratio. For example, when processing the "understanding breakthrough" scenario, the top-layer module of the framework preferentially processes the cognitive state features that inherit high weights, the middle-layer module processes the attention change features that inherit smooth characteristics, and the bottom-layer module initially combines these features according to the original weight ratio. The framework maintains the inheritance of the characteristics of the original fusion region in the feature processing at each level, forming a structurally complete feature integration system.

[0062] Aggregate features using the structurally complete feature integration system. The system implements differentiated feature aggregation strategies for the three levels of the framework. In the top-layer feature processing, an attention enhancement mechanism based on the original weight distribution is designed. The initial dimension of the attention matrix is 64×64, and feature selective enhancement is performed through a three-layer feed-forward network (with dimension configuration of 64-32-16). For the middle-layer features, an adaptive temporal aggregation method is adopted, and the sliding window size is dynamically adjusted within the range of 50-200 ms according to the feature change rate, and the window overlap rate is maintained at 40%. In the bottom-layer feature processing, a multi-level caching mechanism based on feature saliency is implemented, and the cache sizes are 32, 16, and 8 respectively, to maintain the feature continuity at different time scales. At the same time, a bidirectional consistency constraint is introduced in the feature aggregation process: temporal consistency ensures the smoothness of feature evolution, and spatial consistency maintains the rationality of feature distribution. The system also sets feature checkpoints during the aggregation process, and performs feature validity verification once every 200 ms of data processing. The verification content includes the integrity of features, the coherence of timings, and the rationality of distributions. Through this series of refined aggregation processes, a set of multi-scale and multi-level feature aggregation results are finally obtained.

[0063] Output the fused elements based on the multi-scale and multi-level feature aggregation results. The system designs a three-level feature integration network, including a feature extraction layer, a mapping optimization layer, and a fusion and recombination layer. The feature extraction layer uses a depthwise separable convolutional network, and sets convolutional kernels of 1×1, 3×3, and 5×5 respectively for the aggregated features of different scales. The number of channels in the convolutional layer is configured as 16, 32, and 64. In the mapping optimization layer, a non-linear feature transformation mechanism is implemented, and the feature distribution is processed through batch normalization and the LeakyReLU activation function (negative slope 0.2). At the same time, a residual learning module is introduced, and each residual unit contains two convolutional layers and a skip connection to maintain the discriminability of the features. The fusion and recombination layer adopts a dynamic weight allocation mechanism to calculate the fusion weights based on the local correlation and global consistency of the features. For example, when performing cognitive state assessment, the system dynamically combines the feature elements of different levels such as attention features and understanding degree indicators. The weight range of the attention features is 0.3-0.5, and the weight range of the understanding degree indicators is 0.4-0.6. The finally output fused elements are represented by a 128-dimensional vector, where the high-level features occupy 48 dimensions, the middle-level features occupy 48 dimensions, and the low-level features occupy 32 dimensions, achieving a proportional balance of the features at each level.

[0064] In some embodiments, the method of supplementing the missing parts of the fused elements and locating the missing parts, and reconstructing the data of the missing parts to output the complete content includes: establishing a feature integrity evaluation model for obtaining the missing parts of the elements corresponding to the fused elements; restoring the temporal continuity of the missing parts of the elements according to the spatio-temporal interpolation algorithm; performing feature distribution alignment processing on the missing parts of the elements to ensure the consistency of the reconstructed data, and outputting the complete content including the complete temporal features and semantic associations.

[0065] Perform missing supplementation on the fused elements output from the above steps. Based on the fused elements formed in the above steps, including its three-layer feature structure (48 dimensions for high-level features, 48 dimensions for middle-level features, and 32 dimensions for low-level features) and dynamic weight allocation mechanism (calculated based on local correlation and global consistency of features), the system constructs a missing detection framework. The system first analyzes the three-layer feature structure: high-level features (such as attention features, weight range 0.3 - 0.5) are used to characterize the user's attention concentration and understanding level; middle-level features (such as understanding degree indicators, weight range 0.4 - 0.6) are reflected in speech intonation changes and EEG fluctuation features; low-level features continue to play a supporting role (stability index threshold 0.8), reflecting the emotional state index. In cross-language communication scenarios, the system monitors the feature missing situation in real time. For example, when the user is understanding complex sentences, the high-level features may show anomalies that do not match the original weight distribution (deviation threshold 0.25); when answering quickly, the middle-level features may deviate from the original dynamic change range (marked as abnormal when exceeding the threshold). By analyzing the sparsity of feature distribution (density threshold 0.7), temporal continuity (intermittent tolerance 50ms), and feature correlation (correlation coefficient threshold 0.6), the system finally locates three types of missing: cognitive feature anomalies, behavioral feature interruptions, and basic feature deviations.

[0066] Perform data reconstruction on the above three types of feature missing. For cognitive feature anomalies, the system uses the abnormal distribution features located in the first stage and adopts a variational autoencoder for reconstruction (encoding dimension 32, KL divergence constraint 0.01), focusing on maintaining the dominant feature characteristics. For example, when it is detected that the user's attention is scattered, the system performs pattern learning based on high-quality feature segments in historical data (quality score > 0.85) and reconstructs features that conform to the dominant distribution. For behavioral feature interruptions, the system predicts the missing segment using a bidirectional gated recurrent unit network (hidden layer dimension 128, forgetting gate threshold 0.3) according to the complete sequences before and after the interruption position (each taking 200ms). When dealing with speech intonation interruptions, the changing trend during the interruption is inferred through the dynamic features of the context. For basic feature deviations, the system selects adjacent features with high quality assessment scores (> 0.8) as references and ensures its consistency with the overall feature distribution through feature reconstruction (distribution distance threshold 0.2). Through the joint reconstruction of multi-layer features, a complete feature supplementation data set is formed.

[0067] Perform information filling relying on the complete feature supplement dataset. The system designs a hierarchical filling strategy and conducts filling according to the primary and secondary relationships of the reconstructed features. In different communication scenarios, the system adopts differentiated filling schemes. For example, in a formal meeting scenario, the system pays more attention to the precise filling of cognitive features (accuracy requirement > 0.9). When detecting fluctuations in the user's attention during the understanding of professional terms, it will first restore the attention level data and ensure its synchronization with the meeting rhythm (time deviation < 30ms). In the daily conversation scenario, the system is more concerned about the natural transition of behavioral features (smoothness threshold 0.8), such as the continuous change of speech intonation. When the user switches from the mother tongue to a foreign language for communication, the system can capture the feature changes during this conversion process, including behavioral features such as slower speech speed and increased pauses, and perform corresponding feature filling. At the same time, the system also establishes a scenario adaptation mechanism, increasing the filling weight of behavioral features (weight increased to 0.6) in a noisy environment and strengthening the filling accuracy of cognitive features (accuracy increased to 0.95) in a quiet environment. For a long-term interaction process, the system implements a dynamic filling strategy with a sliding window, and the window size is adaptively adjusted according to the interaction rhythm (50ms - 300ms) to ensure the real-time and continuous nature of the filling process. For example, a smaller filling window is used in a fast-paced conversation, while a larger window is adopted during in-depth discussions to capture the complete thinking process. Through this multi-level and multi-scenario filling process, a well-defined complete feature sequence is formed.

[0068] Output the complete content according to the well - structured complete feature sequence. The system constructs a feature integration network to reorganize the filled features. In the academic report scenario, the system strengthens the weight of cognitive features (increased to 0.7), real - time tracks the understanding state of the audience, and when it is found that the difficulty of understanding some professional concepts exceeds the threshold (0.8), the system can accurately locate the content segments that need to be focused on. In the business negotiation scenario, the system pays attention to the balance between behavioral features and cognitive features (weight ratio 0.5:0.5), and by analyzing the subtle changes in speech expression (change detection threshold 0.15) and the fluctuations in cognitive load (fluctuation range ±0.2), it helps users grasp the negotiation rhythm. For example, when it is detected that the user's cognitive load increases (more than 50% of the baseline value) when understanding the foreign party's quotation plan, the system will focus on the feature integrity during this period to ensure accurate capture of the user's understanding and decision - making process. In the daily social scenario, the system pays more attention to the natural expression of emotional features (naturalness threshold 0.85) to ensure the fluency of the interaction process. When dealing with the multi - language switching scenario, the system can accurately capture the feature changes at the language conversion point (conversion recognition accuracy > 0.9) and maintain the continuity of the feature distribution before and after the conversion (distribution distance < 0.25). When the user needs to frequently switch between different languages in the conversation, the system will pay special attention to the changes in cognitive load before and after the language conversion, and by integrating cognitive and behavioral features, it assists the user to achieve smooth language switching. The finally formed complete feature expression not only maintains the original hierarchical attributes but also realizes the continuity of the feature distribution (temporal correlation > 0.85).

[0069] In step S104, perform intention parsing on the complete content to obtain key elements, convert the key elements into interaction information and output the interaction intention; perform information integration on the interaction intention to obtain conversion rules, and generate interaction commands according to the conversion rules to complete the electroencephalogram - speech hybrid interaction based on the brain - computer AI headset.

[0070] Specifically, intention parsing is performed on the complete content output by the above steps. Based on the feature expressions formed in the above steps, including the change in the degree of understanding at the cognitive level, the language switching feature at the behavioral level, and the state fluctuation at the emotional level, the system establishes an intention parsing framework. When processing high-level cognitive features, the system constructs an intention reasoning mechanism to identify the user's intention by analyzing the temporal change pattern of the features (change rate threshold 0.3). For example, when the user encounters an understanding obstacle in cross-language communication, the system infers the intention of understanding difficulty from the sudden change in attention features (sudden change threshold 0.4) and the increase in cognitive load (increase > 50%). When processing middle-level behavioral features, the system combines the change in speech intonation (fluctuation range ±0.25) and the electroencephalogram activity pattern (activity threshold 0.7) to capture the user's expression intention. For low-level emotional features, the system identifies the user's emotional intention by analyzing its fluctuation law (fluctuation period 200 - 500 ms). In the foreign language conference scenario, when it is detected that the user frequently has attention fluctuations (fluctuation frequency > 0.5 Hz) and a slowdown in speech speed (slowdown amplitude > 30%), the system locates these intention change intervals. Through multi-dimensional intention analysis, the system obtains a set of parsing intervals with temporal correlation, and each interval contains the intention feature changes at the cognitive, behavioral, and emotional levels.

[0071] In some embodiments, the intention parsing of the complete content to obtain key elements, converting the key elements into interaction information and outputting an interaction intention includes: constructing a semantic network analysis model to identify the key elements related to the intention of the complete content; obtaining the element association strength corresponding to the key elements related to the intention through a fuzzy inference algorithm to generate candidate interaction intentions; using a reinforcement learning model to optimize the intention generation strategy corresponding to the candidate interaction intentions, eliminating ambiguous understanding, and outputting an interaction intention including language assistance, rhythm regulation, and state adaptation intentions.

[0072] The parsing interval containing the three-layer intention features is established as an understanding model. The system designs corresponding understanding structures according to the intention feature types of each interval. For the interval dominated by cognitive intention features (dominance > 0.7), such as the understanding difficulty interval caused by attention fluctuations, a comprehension evaluation module based on cognitive load is constructed, and the evaluation accuracy threshold is set to 0.85; for the interval dominated by behavioral intention features (dominance 0.5 - 0.7), such as the language expression preparation stage, an expression intention recognition mechanism for speech-electroencephalogram features is established, and the recognition accuracy threshold is set to 0.8; for the interval dominated by emotional intention features (dominance < 0.5), an emotional state understanding module is constructed, and the state resolution is set to 0.1. In a business negotiation scenario, when the user needs to elaborate complex viewpoints in a foreign language, the system determines the understanding strategy for the current interval by identifying the primary and secondary relationships of intention features (primary-secondary ratio difference > 0.3). For example, in the expression preparation stage, the system strengthens the understanding weight of behavioral intention features (increased to 0.6); in the content organization stage, the understanding proportion of cognitive intention features is increased (increased to 0.7). Through this understanding structure based on the primary and secondary features, the system forms a complete intention understanding model.

[0073] Extract information through the intention understanding model based on the primary and secondary features. The system realizes hierarchical extraction of information according to the different interval feature weights set in the model. In the cognition-dominated interval (weight > 0.6), the system focuses on extracting the user's understanding features, such as the understanding difficulty of professional terms (difficulty score 0 - 1), the cognitive burden during language switching (burden index threshold 0.7), etc. In the behavior-dominated interval (weight 0.4 - 0.6), the system focuses on extracting expression features, including the fluency of language organization (fluency index > 0.8), the accuracy of expression (accuracy rate > 0.75), etc. In a multi-person video conference scenario, the system can dynamically adjust the focus of feature extraction according to the different user roles. When the user is an audience, the system mainly extracts cognitive features related to understanding (weight 0.7); when switching to the speaker role, it quickly switches to an extraction mode mainly based on expression features (weight 0.6). In a group discussion environment, the system will pay attention to the conversion process of the user from listening to speaking, and extract the changes in cognitive load (change threshold 0.3), expression preparation features (preparation threshold 0.6), and emotion regulation features (regulation amplitude ±0.2) during this process. For example, when the user gradually adapts to the foreign language environment from a tense state, the system can capture the relaxation trend of electroencephalogram features (trend slope < -0.1), the improvement in the naturalness of speech intonation (improvement amplitude > 20%), and the increase in the fluency of cognitive processing (increase amplitude > 30%). When conducting professional field communication, the system will pay special attention to the term understanding and expression links, and extract the changes in cognitive features of the user when dealing with professional concepts (change sensitivity 0.8). Through this information extraction considering feature weights, the system obtains the complete intention feature sequence of the user during cross-language interaction.

[0074] Output the interaction intention using the described complete intention feature sequence. The system designs an intention integration mechanism based on the extracted feature sequence, converting the feature information in different intervals into specific interaction intentions. In the formal meeting scenario, the system integrates the features in the cognition-dominated interval (weight 0.8) and identifies the critical moments requiring language assistance (assistance trigger threshold 0.7). For example, when it is detected that the difficulty level of the user in understanding the foreign party's speech exceeds the threshold (0.75), the system parses it as the intention of content refinement. In the cross-language negotiation scenario, the system combines the features of the behavior-dominated interval (weight 0.6) and the emotion-dominated interval (weight 0.4) to identify the intention of adjusting the conversation rhythm. In scenarios with different language difficulties, the system will adopt different intention recognition strategies. For the entry-level language proficiency, the system pays more attention to identifying the intentions of basic understanding and expression (basic understanding threshold 0.6); for the advanced language proficiency, it focuses on identifying the intentions of language details and context understanding (detail understanding threshold 0.8). In the remote collaboration scenario, the system can identify the interaction obstacles caused by network latency (latency threshold > 100ms) and distinguish the understanding difficulties caused by technical factors and language factors. When the user needs to quickly switch between different languages, the system anticipates their language preparation status by analyzing the cognitive and behavioral changes in the feature sequence (change rate > 0.4). In the informal social scenario, the system adjusts the sensitivity of intention recognition according to the formality of the occasion (0.8 for formal occasions, 0.6 for informal occasions), maintaining the naturalness of the conversation while ensuring the interaction effect. At the same time, the system can identify the fatigue state of the user (fatigue index > 0.7) and adjust the intention recognition strategy in a timely manner during long-term cross-language communication. Through this multi-scenario and multi-level intention integration, the system finally outputs a set of interaction intentions including interaction types, priorities, and execution timings.

[0075] Integrate the information of the interaction intentions output from the above steps. Based on the set of interaction intentions formed in the above steps, including language assistance intentions (such as professional term understanding, expression suggestions, priority weight 0.8), rhythm regulation intentions (such as speech rate adjustment, interaction timing, priority weight 0.6), and state adaptation intentions (such as fatigue adjustment, attention replenishment, priority weight 0.4), the system establishes an information integration framework. When processing high-priority interaction intentions of the language assistance type, the system constructs a fast response mechanism. For example, when it is detected that the difficulty of the user's understanding of professional terms in an international conference exceeds the threshold (0.75), the system preferentially processes the term explanation intention, and at the same time sets the relevant context understanding intention (correlation degree > 0.6) for subsequent processing. When processing regular-priority interaction intentions of the rhythm regulation type, the system combines the scene characteristics and the user's state to determine the assistance strategy. In the video conference scene, when the user's role changes from listener to speaker, the system can recognize that the priority of the speech rate adjustment intention is increased (the increase range is 30%). For low-priority interaction intentions of the state adaptation type, the system analyzes their temporal characteristics (duration > 2 minutes) to plan appropriate response times. For example, the fatigue adjustment intention (fatigue index > 0.7) detected in a long meeting will be arranged for processing when the meeting node is switched. Through this hierarchical integration, the system forms a set of conversion rules based on intention types and scene characteristics.

[0076] Apply the conversion rules based on intention types and scene characteristics to information processing. For the language assistance type of rules, the system starts a real-time processing channel to ensure the timeliness of professional term explanation and expression suggestions (response delay < 100ms). For example, in a business negotiation, when the rule triggers term understanding assistance, the system preferentially processes the core vocabulary information (importance > 0.8), and at the same time establishes a term association network to prepare for subsequent understanding. For the rhythm regulation type of rules, the system establishes a dynamic batch processing mechanism (batch processing window 200 - 500ms). In the academic exchange scene, the system processes the user's expression rhythm according to the speech rate adjustment rule (rhythm coordination degree > 0.75), and at the same time adjusts the information input rate in combination with the attention level rule (attention threshold 0.7). For the state adaptation type of rules, the system adopts a predictive processing strategy to plan in advance the rest reminder and attention recovery plan (prediction window 5 - 10 minutes). In the multi-person seminar scene, the system can simultaneously process the different rule requirements of multiple users, such as coordinating the expression assistance rule of the speaker (priority 0.8) and the understanding assistance rule of the listener (priority 0.6). Through this multi-level rule processing, a structured processing result including core information, auxiliary information, and state information is formed.

[0077] Organize information using the structured processing results containing three types of information. The system realizes the three-dimensional organization of information according to the hierarchical attributes of the processing results. For the core information layer, the system adopts an organization method that highlights key points, and prioritizes the key content (importance > 0.85) in the language assistance results. For example, in the real-time translation scenario, the system prominently marks the explanations of professional terms and core idea prompts obtained through processing, and establishes associations (association strength > 0.7) with other auxiliary information. For the auxiliary information layer, the system constructs a context association network (network density > 0.6), and organizes content such as rhythm control prompts and expression suggestions according to logical relationships. In the technical discussion scenario, the system can systematically organize auxiliary content such as the association relationships between professional concepts (association degree threshold 0.65) and application scenario descriptions. For the status information layer, the system establishes a dynamic update mechanism (update cycle 500ms), and adjusts the information organization strategy according to the real-time changes in the user's status. For example, when it detects a decrease in the user's attention (decrease amplitude > 20%), the system will increase the salience of key markings (salience increase 30%) and simplify the presentation form of non-core information. At the same time, the system also establishes a scenario adaptation mechanism to adjust the information organization form in different interaction environments (environment adaptability > 0.8). Through this multi-dimensional information organization, the system forms a complete information architecture including core commands, auxiliary commands, and status feedback.

[0078] Output interaction commands according to the described complete information architecture. Based on the hierarchical organization of information, the system designs an adaptive command generation mechanism. When generating core commands, the system focuses on timeliness and accuracy, and transforms language assistance information into intuitive interaction instructions (clarity index > 0.9). For example, when users encounter expression difficulties (difficulty index > 0.7) in cross - language discussions, the system can generate a main command sequence containing expression templates and keyword prompts based on the content of the core information layer. When generating auxiliary commands, the system adopts a progressive strategy and transforms rhythm control information into guiding instructions (guidance effect index > 0.8). In a team collaboration scenario, the system can transform the content of the auxiliary information layer into personalized interaction prompts (personalized matching degree > 0.75) according to the needs and language levels of different roles, such as generating speech rate control suggestions (control accuracy ± 10%) for the speaker and generating key point tracking guides (tracking accuracy > 0.85) for the audience. When generating status feedback commands, the system realizes preventive command planning and transforms the content of the status information layer into regulatory instructions (regulation effect index > 0.7). For example, in a long - term meeting, the system will generate timely rest reminders and attention adjustment suggestions according to the change of user fatigue (change trend slope > 0.1). During the dynamic interaction process, the system can adjust the command generation strategy in real - time according to the scene change (strategy adaptability > 0.8) to ensure the coherence and practicality of the command sequence. Through this multi - level command generation process, the system finally outputs a complete interaction command system including a main command sequence, an auxiliary instruction set, and status adjustment commands.

[0079] The method has at least the following beneficial effects:

[0080] 1. By constructing a multi - level feature extraction and fusion mechanism, precise synchronous acquisition, noise reduction processing, and feature reconstruction of EEG signals and speech signals are achieved. The system uses a high - precision clock synchronization circuit and dynamic time - base calibration to significantly reduce the sampling time difference of bimodal signals; adopts an adaptive noise reduction and signal separation algorithm to effectively improve the signal - to - noise ratio of the two signals; and realizes the accurate correspondence and complementary enhancement of bimodal features through multi - dimensional feature mapping and modal conversion.

[0081] 2. A feature supplement and optimization framework based on temporal correlation is established, effectively solving the problems of feature loss and distribution drift. The system accurately locates the missing areas by analyzing the temporal evolution law of features; uses a multi - level reconstruction network to achieve high - quality repair of features; and adaptively adjusts the feature distribution based on a dynamic feature fusion mechanism. In practical applications, the system can timely detect and supplement feature losses caused by factors such as environmental noise and attention fluctuations, effectively improving the adaptability of the system in complex environments.

[0082] 3. A hierarchical intention understanding and command generation mechanism is designed to achieve the scenario adaptation of interaction strategies. The system constructs a multi-level intention parsing framework according to different priorities of interaction requirements; improves the accuracy of intention recognition through feature integration and rule conversion; and dynamically generates personalized interaction commands based on scenario features and user states. This adaptive interaction mechanism enables the system to achieve fast response and accurate interaction in complex scenarios such as multi-person meetings and cross-language communications.

[0083] To execute the EEG-voice hybrid interaction method corresponding to the above method embodiments of the brain-computer AI headset to achieve the corresponding functions and technical effects. Refer to Figure 2 , Figure 2 The block diagram of a EEG-voice hybrid interaction device 200 provided by an embodiment of the present application is shown. For the sake of convenience of description, only the parts related to this embodiment are shown. The EEG-voice hybrid interaction device 200 provided by the embodiment of the present application includes:

[0084] A waveform output unit 201, configured to collect the bimodal signals of the AI headset, classify and process the bimodal signals, and output timing information; perform noise reduction and signal separation processing on the timing information, and output a pure waveform after optimizing and purifying the processing results corresponding to the separation processing;

[0085] A mode generation unit 202, configured to perform dual-channel analysis on the pure waveform and divide the feature dimensions, complete mode conversion through the constructed mapping chain, and output feature identifiers; perform timing analysis on the feature identifiers to identify a set of key timing features, and generate an associated mode after establishing an aligned feature mapping structure and combining modes;

[0086] A node annotation unit 203, configured to perform path analysis on the associated mode and annotate combined nodes, reconstruct features through a conversion chain and then output a representation structure; perform hierarchical fusion on the representation structure and extract a fusion region, and output fusion elements according to the fusion region; perform missing supplement on the fusion elements and locate the missing parts, and perform data reconstruction on the missing parts to output complete content;

[0087] An element acquisition unit 204, configured to perform intention parsing on the complete content to obtain key elements, convert the key elements into interaction information and output an interaction intention; perform information integration on the interaction intention to obtain a conversion rule, and generate an interaction command according to the conversion rule to complete the EEG-voice hybrid interaction based on the brain-computer AI headset.

[0088] Figure 3 The structural schematic diagram of a computer device provided by an embodiment of the present application. As Figure 3 shown, the computer device 3 of this embodiment includes: at least one processor 30 ( Figure 3Only one is shown in the figure), a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30. When the processor 30 executes the computer program 32, the steps in any of the above method embodiments are implemented.

[0089] The computer device 3 may be a computing device such as a smart phone, a tablet computer, a desktop computer, and a cloud server. The computer device may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art can understand that Figure 3 merely examples of the computer device 3, which do not constitute a limitation on the computer device 3, and may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0090] The so-called processor 30 may be a central processing unit (CPU), and the processor 30 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0091] In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In other embodiments, the memory 31 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 3. Further, the memory 31 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 31 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program, etc. The memory 31 may also be used to temporarily store data that has been output or is to be output.

[0092] In addition, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0093] An embodiment of the present application provides a computer program product. When the computer program product runs on a computer device, the computer device is caused to implement the steps in each of the above method embodiments when executed.

[0094] In several embodiments provided by the present application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.

[0095] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0096] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A brain-computer AI headset-based EEG-speech hybrid interaction method, characterized in that: include: Collecting bimodal signals from the AI ​​headset, classifying and processing the bimodal signals, and outputting timing information; Performing noise reduction and signal separation processing on the timing information, optimizing and purifying the processing result corresponding to the separation processing, and then outputting a pure waveform; Performing two-way analysis on the pure waveform and dividing the feature dimensions, completing modal conversion through the constructed mapping chain, and outputting feature identification; Performing time series analysis on the feature identifiers to identify a set of key time series features, and generating a correlation pattern after establishing an alignment feature mapping structure combination mode; Perform path analysis on the association pattern and mark the combined nodes, reconstruct the features through the transformation chain and output the representation structure; perform hierarchical fusion on the representation structure and extract the fusion area, and output the fusion element according to the fusion area; supplement the missing of the fusion element and locate the missing part, reconstruct the data of the missing part and output the complete content; Performing intent analysis on the complete content to obtain key elements, converting the key elements into interaction information and outputting interaction intent; The interaction intention is integrated to obtain conversion rules, and interaction commands are generated according to the conversion rules to complete the EEG-speech hybrid interaction based on the brain-computer AI headset.

2. The method according to claim 1, characterized in that The step of classifying and processing the bimodal signal and outputting timing information includes: Synchronously collecting EEG signals and speech signals through a high-precision sensor array to form the dual-modal signal; Establish a time reference point based on the GPS timing module and use FPGA clock synchronization circuit to achieve time alignment of dual-mode signals; The speech signal is processed in segments, and an adaptive threshold mechanism is used to divide the speech activity area and the non-speech area; The EEG signals were preprocessed by the sliding window method and characterized by combining the Hjorth parameter; The bimodal signal is classified into a pure speech segment, a pure EEG segment, a bimodal activity segment and a background noise segment, and is used to construct the timing information including a timestamp and a signal quality indicator.

3. The method according to claim 2, characterized in that The step of performing noise reduction and signal separation processing on the time series information, optimizing and purifying the processing result corresponding to the separation processing and outputting a pure waveform includes: Adopting an adaptive filter to eliminate power frequency interference on the EEG signal, and combining independent component analysis to separate electromyographic artifacts; Implementing a blind source separation algorithm on the speech signal to eliminate environmental noise and suppressing sudden interference through morphological filtering; Using an improved wavelet threshold method to denoise the EEG signal and retain characteristic waveform details; Performing spectral subtraction noise reduction processing on the speech signal, and combining with a Wiener filter to enhance speech features; The noise reduction parameters of the timing information are dynamically adjusted according to the signal quality index, and the pure waveform with improved signal-to-noise ratio is output.

4. The method according to claim 2, characterized in that: The two-way analysis is performed on the pure waveform and the feature dimensions are divided, the modal conversion is completed through the constructed mapping chain, and the feature identification is output, including: Performing wavelet packet decomposition on the EEG signal to obtain time-frequency features, and extracting intrinsic modal components in combination with empirical mode decomposition; Extracting Mel-frequency cepstral coefficients and acoustic feature parameters from the speech signal; The Fisher discriminant criterion is used to evaluate the discriminability of feature dimensions and establish a hierarchical feature selection strategy; The feature identifier is output according to the time-frequency features, the intrinsic modal components, the Mel-frequency cepstral coefficients, the acoustic feature parameters and the hierarchical feature selection strategy.

5. The method according to claim 1, characterized in that The performing path analysis on the association pattern and marking the combined nodes, reconstructing the features through the transformation chain and then outputting the representation structure, comprises: A dynamic programming algorithm is used to find the optimal characteristic path for the association pattern, and the combination nodes associated with the time series are marked; Identify the key connection relationship of the combined nodes through a graph attention network and construct a feature conversion chain; Implementing a Markov model-based path optimization on the feature conversion chain to eliminate redundant feature connections; The tensor decomposition method is used to reconstruct the feature space corresponding to the feature transformation chain, maintain the topological relationship between features, and output a multi-dimensional representation structure including principal component features and temporal relationships.

6. The method according to claim 1, characterized in that The step of performing hierarchical fusion on the representation structure and extracting a fusion region, and outputting a fusion element according to the fusion region, comprises: Constructing a multi-scale feature pyramid structure for performing hierarchical feature fusion on the representation structure to obtain the fusion region; Calculating feature region weights of the fusion region through an attention mechanism to identify key fusion regions; A region growing algorithm is used to expand the fusion boundary of the key fusion area, maintain feature continuity, and output fusion elements containing EEG-speech collaborative features.

7. The method according to claim 1, characterized in that The missing elements are supplemented and the missing parts are located, and the missing parts are reconstructed and the complete content is output. include: Establishing a feature integrity assessment model for obtaining missing elements corresponding to the fused elements; Restoring the temporal continuity of the missing parts of the elements according to a spatiotemporal interpolation algorithm; The feature distribution alignment processing is performed on the missing parts of the elements to ensure the consistency of the reconstructed data, and the complete content including complete temporal features and semantic associations is output.

8. The method according to claim 1, characterized in that The performing intent parsing on the complete content to obtain key elements, converting the key elements into interaction information and outputting interaction intent includes: Constructing a semantic network analysis model to identify key elements related to the intent of the complete content; Obtain the element association strength corresponding to the key elements related to the intention through a fuzzy reasoning algorithm to generate a candidate interaction intention; A reinforcement learning model is used to optimize the intention generation strategy corresponding to the candidate interaction intention, eliminate ambiguous understanding, and output the interaction intention including language assistance, rhythm regulation and state adaptation intention.

9. A brain wave-speech hybrid interaction device, characterized in that: include: A waveform output unit, used to collect the bimodal signal of the AI ​​headset, classify and process the bimodal signal, and output timing information; Performing noise reduction and signal separation processing on the timing information, optimizing and purifying the processing result corresponding to the separation processing, and then outputting a pure waveform; A mode generation unit, used for performing a two-way analysis on the pure waveform and dividing the feature dimensions, completing the mode conversion through the constructed mapping chain, and outputting the feature identification; Performing time series analysis on the feature identifiers to identify a set of key time series features, and generating a correlation pattern after establishing an alignment feature mapping structure combination mode; A node labeling unit is used to perform path analysis on the association pattern and label the combined nodes, reconstruct the features through the transformation chain and then output the representation structure; hierarchically fuse the representation structure and extract the fusion area, and output the fusion element according to the fusion area; supplement the missing of the fusion element and locate the missing part, reconstruct the data of the missing part and output the complete content; An element acquisition unit, configured to perform intent analysis on the complete content to acquire key elements, convert the key elements into interaction information and output interaction intent; The interaction intention is integrated to obtain conversion rules, and interaction commands are generated according to the conversion rules to complete the EEG-speech hybrid interaction based on the brain-computer AI headset.

10. A computer device, characterized in that: The method comprises a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the method according to any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Voice signal analysis sub-system based on multi-modal emotion identification system

    CN108899050A

  • Music emotion label generation method based on multi-modal fusion

    CN115757860A

  • Data processing method, system, device and equipment and storage medium

    CN116048282A

  • Electroencephalogram signal voice decoding method based on generative adversarial network

    CN116364096A

  • Multi-mode fusion depression identification auxiliary decision-making system based on EEG and voice signals

    CN117116468A

Cited By

  • Emotion evaluation method and device based on depth time sequence modeling

    CN120959742A

  • Robot control method based on voice recognition and brain waves

    CN121200043A

  • Brain-computer interface-based self-adaptive music regulation and control upper limb movement method and system

    CN122331772A

  • A method and system for regulating upper limb movement based on brain-computer interface adaptive music

    CN122331772B