Brainwave audio auditory perception synchronous feedback method, device, system and electronic equipment
By acquiring EEG signals online in real time and generating corresponding audio representations, the problem of lack of real-time feedback and subjective dependence in existing technologies for brainwave audio has been solved, realizing the organic unity of neural modulation and audio generation and personalized audio generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WEIZHINAO DATA SERVICE (TIANJIN) CO LTD
- Filing Date
- 2025-05-20
- Publication Date
- 2026-04-17
AI Technical Summary
Existing brainwave audio technology lacks real-time feedback capabilities in neural modulation and audio generation applications. It cannot adjust stimulation signals according to the dynamic changes in individual neural activity, and the generated audio depends on subjective parameters and preset templates, failing to achieve objective mapping and real-time synchronization of neural activity.
By acquiring EEG signals online in real time, extracting EEG features in a specified frequency band, segmenting the EEG signals into multiple segments based on audio representation and generation methods, generating audio representations corresponding to the segments, and using a MIDI synthesizer or audio synthesis library to generate brainwave audio, the direct mapping and real-time feedback between neural activity and audio features are achieved.
It achieves an organic unity between neural regulation and audio generation, enhancing physiological adaptability and artistic expression. Through closed-loop feedback interaction, it realizes real-time dynamic regulation of the nervous system and personalized audio generation.
Smart Images

Figure CN120803248B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of electroencephalogram (EEG) signal processing technology, specifically to a method, device, system, and electronic device for synchronous feedback of proprioceptive brainwave audio auditory perception. Background Technology
[0002] In recent years, the development of brain-computer interface (BCI) technology has provided new possibilities for neural modulation and audio generation. Among them, brain-wave audio (BWA), as an auditory stimulus generated based on brain neural activity signals (such as EEG, MEG, and fNIRS), has shown potential application value in neuroscience, psychotherapy, and music computing. Brain-wave audio achieves the mapping between brainwave features and audio features (BwA()) through a brain-computer audio interface (BCAI), thereby generating personalized auditory stimulus signals. Brainwave music (BWM), as a subclass of BWA, needs to further meet the synthesis requirements of musical features and instrument timbres.
[0003] Currently, brainwave audio technology is mainly applied in two areas: neuromodulation and audio generation. In neuromodulation applications, brainwave audio affects the central nervous system by stimulating the auditory cortex, and has been used to improve sleep quality and regulate emotional state, showing promise for the regulation or treatment of neurological disorders. In audio generation applications, as a means of music generation, brainwave audio uses BCAI technology to establish a pathway from neural activity to music output, using brainwave signals to control music generation. It can be used as a means of music creation and entertainment, and applied in accessible music creation.
[0004] However, existing technologies have certain functional limitations in both neural modulation and audio generation applications, mainly reflected in the following aspects:
[0005] For neuromodulation applications, non-real-time methods are primarily used to generate brainwave music with musical characteristics. This involves synthesizing fixed audio with musical features based on offline EEG data, and then repeatedly playing it for extended periods to stimulate the brain. This approach is only suitable for offline applications, lacks real-time feedback capabilities, and cannot respond to dynamic changes in brainwaves in real time or adjust stimulation signals according to the dynamic changes in individual neural activity, resulting in a mismatch between the stimulation signal and the current neural state.
[0006] In audio generation applications, EEG signals generated during specific tasks (such as attention or motor imagery) are typically used as control signals to instruct music generation software to produce brainwave audio. This involves using a limited set of selectable musical combinations or patterns (such as pre-defined melody fragments, harmonic progressions, and rhythmic patterns). The EEG signals are used only to select or adjust among these options to generate the music desired by the operator, rather than a direct audio representation of brain activity. Compared to the individual EEG signals of the subject, the operator's subjective selection of preset parameters has a greater impact on the audio synthesis effect. The brainwave audio generated using this method relies on preset templates and the operator's active intention. The final output reflects the operator's subjective preferences or intentions; it is the result of the operator's "selection" through EEG signals, rather than an objective mapping of spontaneous brain activity. Its output is merely for music creation and cannot achieve the regulation of neural activity, thus reducing the neural representativeness of the brainwave audio.
[0007] In summary, current brainwave audio generation methods lack effective EEG signal processing techniques. The generation mapping relationship and effect of brainwave audio rely on subjective parameter settings and subjective scale evaluations, lacking an objective mapping mechanism. This results in the generation effect failing to stably reproduce personalized characteristics and failing to establish an essential connection between neural activity characteristics and audio characteristics. The real-time nature, diversity, and adaptive control possibilities of brainwave audio generation have not been fully explored. It cannot simultaneously satisfy the physiological adaptability of neural modulation and the artistic expressiveness of audio generation, and fails to achieve the organic unity of neural modulation and audio generation processes. Consequently, existing brainwave audio technology cannot fully realize the therapeutic potential of neural modulation, nor can it achieve creative audio expression based on real neural activity. Its application effect is fundamentally limited by artificial intervention and static methods.
[0008] How to simultaneously satisfy the physiological adaptability of neural regulation and the artistic expressiveness of audio generation, and achieve the organic unity of neural regulation and audio generation processes, is an urgent problem to be solved. Summary of the Invention
[0009] To address the problems in the related technologies, this disclosure provides a method, apparatus, system, and electronic device for synchronous feedback of proprioceptive brainwave audio auditory perception.
[0010] In a first aspect, this disclosure provides a method for synchronous feedback of proprioceptive brainwave audio auditory perception. The method is applied to a host computer, which is connected to an EEG acquisition device via an EEG data transmission interface. The method includes:
[0011] The EEG data transmission interface receives the EEG signals of the user being collected online from the EEG acquisition device. The EEG signals are acquired by the EEG acquisition device through one or more electrode channels, wherein each electrode channel corresponds to one electrode or a combination of multiple electrodes.
[0012] One or more specified frequency band EEG signals are acquired from the EEG signals, and EEG features corresponding to the specified frequency band EEG signals are extracted based on the one or more specified frequency band EEG signals. The specified frequency band EEG signals refer to EEG signals located in the specified frequency band, and the EEG features include EEG time domain features and / or EEG frequency domain features.
[0013] For the specified frequency band EEG signal, based on a specified audio representation and a specified audio generation method, the specified frequency band EEG signal is segmented into multiple EEG segments according to the EEG characteristics or preset rhythm parameters, and multiple audio representations corresponding to the multiple EEG segments are generated. The feature parameters of the audio representation are determined according to the EEG characteristics of the corresponding EEG segment. The audio representation includes: note form and / or waveform form, wherein the audio representation corresponding to the note form is a note, and the audio representation corresponding to the waveform form is a waveform. The audio generation method includes: direct mapping method and / or indirect control method using control parameters.
[0014] Brainwave audio representation data is generated based on the audio representations corresponding to one or more specified frequency band EEG signals.
[0015] Brainwave audio is obtained based on the brainwave audio representation data.
[0016] According to embodiments of this disclosure, when the specified audio representation is in the form of musical notes and the specified audio generation method is a direct mapping method, the characteristic parameters of the audio representation include: note duration, pitch, and intensity; the step of segmenting the specified frequency band EEG signal into multiple EEG segments according to preset rhythm parameters and generating multiple audio representations corresponding to the multiple EEG segments respectively includes:
[0017] For a predetermined duration of EEG signal in the specified frequency band of the specified electrode channel:
[0018] The EEG signal of the preset duration is divided according to the note duration indicated by the preset rhythm parameters to obtain a set of EEG segments. The set of EEG segments contains multiple EEG segments, and each EEG segment corresponds to a note.
[0019] For any EEG segment in the set of EEG segments, the pitch and intensity of the note corresponding to the any EEG segment are determined based on the frequency domain characteristics of the EEG segment, until the pitch and intensity of the note corresponding to each EEG segment in the set of EEG segments are obtained. This includes: selecting the peak frequency or weighted average frequency with the largest amplitude in the power spectrum characteristics of the power spectrum characteristics as the feature frequency; obtaining the corresponding specified pitch range based on the frequency range of the specified frequency band based on a preset mapping method; obtaining the pitch corresponding to the any EEG segment based on the feature frequency and the specified pitch range; generating the intensity corresponding to the any EEG segment based on the total energy or peak amplitude of the power spectrum characteristics; the preset mapping method is used to describe the mapping relationship between the frequency range of the specified frequency band and the pitch range.
[0020] According to embodiments of this disclosure, generating brainwave audio representation data based on the audio representations corresponding to the one or more specified frequency band EEG signals includes:
[0021] The notes corresponding to each EEG segment in the set of EEG segments of a preset duration in the specified frequency band of each electrode channel are written into the musical score corresponding to the specified frequency band.
[0022] The musical score corresponding to each specified frequency band in each electrode channel is written into a complete musical score set corresponding to the preset duration, and the complete musical score set is used as the brainwave audio representation data.
[0023] The step of obtaining brainwave audio based on the brainwave audio representation data includes:
[0024] According to a preset frequency band-timbre mapping table or configuration parameter-timbre mapping table, the timbre is configured for the musical score corresponding to each specified frequency band in the brainwave audio representation data. The frequency band-timbre mapping table is used to describe the correspondence between frequency bands and timbres. The configuration parameter-timbre mapping table is used to describe the correspondence between specified timbre configuration parameters and timbres. The specified timbre configuration parameters include: non-brain-derived control parameters and / or non-brain-derived event occurrence frequency.
[0025] The non-brain-derived control parameter is obtained as follows: when the EEG signal of the preset length in the specified electrode channel contains both non-brain-derived signals and denoised EEG signals, the non-brain-derived control parameter is obtained by taking the ratio of the absolute values of the non-brain-derived signals and the denoised EEG signals in the preset length of the EEG signal in the specified electrode channel to the preset fractional part.
[0026] The non-brain-derived event frequency is obtained as follows: a non-brain-derived event threshold is obtained by taking the absolute value of the denoised EEG signal and dividing it by a preset fractional number; positions in the non-brain-derived signal that are higher than the non-brain-derived event threshold are counted to obtain the number of non-brain-derived events in the non-brain-derived signal; and the non-brain-derived event frequency is obtained based on the number of non-brain-derived events and the preset duration.
[0027] The brainwave audio is generated using a MIDI synthesizer or audio synthesis library based on the musical score corresponding to each specified frequency band in the brainwave audio representation data.
[0028] According to embodiments of this disclosure, when the specified audio representation is in the form of musical notes and the specified audio generation method is an indirect control method using control parameters, the characteristic parameters of the audio representation include: note duration, pitch, and intensity; the step of segmenting the specified frequency band EEG signal into multiple EEG segments based on the EEG characteristics and generating multiple audio representations corresponding to the multiple EEG segments respectively includes:
[0029] The EEG signal of a preset duration in the specified frequency band is used as input, and a pre-trained user state discrimination model is used to generate control parameters, which include: arousal level and valence.
[0030] The duration of a single note in the note sequence corresponding to the EEG signal of the preset duration is calculated based on the arousal level and / or the valence.
[0031] The EEG signal of the preset duration is divided according to the duration of the note to obtain a set of EEG segments. The set of EEG segments contains multiple EEG segments, and each EEG segment corresponds to a note.
[0032] The probability of occurrence of a note corresponding to each EEG segment in the set of EEG segments is calculated based on the arousal level and / or the valence. When the occurrence probability meets the occurrence condition, the note corresponding to the corresponding EEG segment appears; otherwise, the note corresponding to the corresponding EEG segment does not appear. The note sequence corresponding to the EEG signal of the preset duration is determined based on the calculation results.
[0033] The intensity of the notes in the note sequence is calculated based on the arousal level, the valence, and the power spectrum energy characteristics of the EEG segments corresponding to the notes in the note sequence.
[0034] Calculate the tonality parameters corresponding to the notes in the note sequence based on the arousal level and / or the valence;
[0035] The pitch of the notes in the note sequence is calculated based on the tonality parameter and the power spectrum frequency characteristics of the EEG segments corresponding to the notes in the note sequence;
[0036] Configure the timbre of the notes in the note sequence according to the arousal level and / or the valence; or,
[0037] According to the preset configuration parameter-timbre mapping table, configure the timbre of all notes in the note sequence; the configuration parameter-timbre mapping table is used to describe the correspondence between the specified timbre configuration parameters and the timbre, and the specified timbre configuration parameters include: non-brain-derived control parameters and / or non-brain-derived event occurrence frequency;
[0038] The non-brain-derived control parameter is obtained as follows: when the EEG signal of the preset length in the specified electrode channel contains both non-brain-derived signals and denoised EEG signals, the ratio of the preset fractional part of the absolute values of the non-brain-derived signals and the denoised EEG signals in the preset length of the EEG signal in the specified electrode channel is taken as the non-brain-derived control parameter.
[0039] The non-brain-derived event frequency is obtained as follows: a non-brain-derived event threshold is obtained by taking the absolute value of the denoised EEG signal and dividing it by a preset fractional number; positions in the non-brain-derived signal that are higher than the non-brain-derived event threshold are counted to obtain the number of non-brain-derived events occurring in the non-brain-derived signal; and the non-brain-derived event frequency is obtained based on the number of non-brain-derived events and the preset duration.
[0040] According to embodiments of this disclosure, generating brainwave audio representation data based on the audio representations corresponding to the one or more specified frequency band EEG signals includes:
[0041] Based on the note sequence of the specified frequency band in each electrode channel, and the intensity, pitch, and timbre of the notes in the note sequence, a musical score corresponding to the specified frequency band is generated;
[0042] The musical score corresponding to each specified frequency band in each electrode channel is written into a complete musical score set corresponding to the preset duration, and the complete musical score set is used as the brainwave audio representation data.
[0043] The step of obtaining brainwave audio based on the brainwave audio representation data includes:
[0044] The brainwave audio is generated using a MIDI synthesizer or audio synthesis library based on the musical score corresponding to each specified frequency band in the brainwave audio representation data.
[0045] According to embodiments of this disclosure, when the specified audio representation is a waveform and the specified audio generation method is a direct mapping method, the characteristic parameters of the audio representation include: fundamental frequency and harmonic frequency, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency; the step of segmenting the specified frequency band EEG signal into multiple EEG segments based on the EEG characteristics and generating multiple audio representations corresponding to the multiple EEG segments respectively includes:
[0046] For a predetermined duration of EEG signal in the specified frequency band of the specified electrode channel:
[0047] Based on the temporal envelope corresponding to the EEG signal of the preset duration, the EEG signal of the preset duration is segmented to obtain a set of EEG segments, including: determining an effective amplitude threshold based on the mean of the temporal envelope; determining a set of troughs in the temporal envelope based on the effective amplitude threshold, wherein the amplitude of the troughs in the set of troughs is not greater than the effective amplitude threshold; segmenting the EEG signal of the preset duration based on a preset effective time threshold and the set of troughs in the temporal envelope to obtain the set of EEG segments, wherein the set of EEG segments contains one or more EEG segments, and the time value of each EEG segment is greater than the effective time threshold;
[0048] For any EEG segment in the set of EEG segments, the characteristic parameters of the waveform corresponding to any EEG segment are determined based on the frequency domain characteristics of the EEG segment, until the characteristic parameters of the waveform corresponding to each EEG segment in the set of EEG segments are obtained. This includes: based on the power spectrum characteristics of any EEG segment, determining the fundamental frequency and harmonic frequency of the waveform corresponding to any EEG segment, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency, based on the peak frequencies and amplitudes of the top preset number of peak frequencies ranked by amplitude in the power spectrum characteristics.
[0049] According to embodiments of this disclosure, generating brainwave audio representation data based on the audio representations corresponding to the one or more specified frequency band EEG signals includes:
[0050] For any EEG segment in the set of EEG segments, the audio waveform corresponding to any EEG segment is determined based on the characteristic parameters and time-domain envelope of the waveform corresponding to any EEG segment, until the audio waveforms corresponding to each EEG segment in the set of EEG segments are obtained. This includes: generating sine waves of corresponding frequencies based on the fundamental frequency and harmonic frequency of the waveform corresponding to any EEG segment, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency, and mixing them to obtain a mixed sine wave; applying the time-domain envelope corresponding to any EEG segment to the mixed sine wave for modulation to obtain the audio waveform corresponding to any EEG segment.
[0051] The audio waveforms corresponding to each EEG segment in the set of EEG segments in each specified frequency band of the specified electrode channel are used as the audio waveforms corresponding to the EEG signals of the specified electrode channel with the preset duration.
[0052] The audio waveforms corresponding to the EEG signals of each electrode channel for the preset duration are used as the EEG audio representation data.
[0053] The step of obtaining brainwave audio based on the brainwave audio representation data includes:
[0054] The brainwave audio representation data is directly used as brainwave audio.
[0055] Secondly, this disclosure provides a proprioceptive brainwave audio-visual perception synchronous feedback device. The device is mounted on a host computer, which is connected to an EEG acquisition device via an EEG data transmission interface. The device includes: an EEG signal receiving module, an EEG feature extraction module, and an EEG-audio generation module.
[0056] The EEG signal receiving module is configured to receive EEG signals of the user being collected online from the EEG acquisition device through the EEG data transmission interface. The EEG signals are acquired by the EEG acquisition device through one or more electrode channels, wherein each electrode channel corresponds to one electrode or a combination of multiple electrodes.
[0057] The EEG feature extraction module is configured to acquire one or more EEG signals of a specified frequency band from the EEG signals, and extract EEG features corresponding to the EEG signals of the specified frequency bands based on the one or more EEG signals of the specified frequency bands. The EEG signals of the specified frequency bands refer to EEG signals located in the specified frequency bands. The EEG features include EEG time-domain features and / or EEG frequency-domain features.
[0058] The EEG-audio generation module is configured to, for the specified frequency band EEG signal, based on a specified audio representation format and a specified audio generation method, segment the specified frequency band EEG signal into multiple EEG segments according to the EEG characteristics or preset rhythm parameters, and generate multiple audio representations corresponding to the multiple EEG segments respectively. The feature parameters of the audio representation are determined based on the EEG characteristics of the corresponding EEG segment. The audio representation format includes: note form and / or waveform form, wherein the audio representation corresponding to the note form is a note, and the audio representation corresponding to the waveform form is a waveform. The audio generation method includes: direct mapping method and / or indirect control method using control parameters. Brainwave audio representation data is generated based on the audio representations corresponding to one or more specified frequency band EEG signals. Brainwave audio is obtained based on the brainwave audio representation data.
[0059] Thirdly, this disclosure provides an electronic device that includes the apparatus described in any one of the second aspects, or includes a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method described in any one of the first aspects.
[0060] Fourthly, this disclosure provides a proprioceptive brainwave audio auditory perception synchronous feedback system, the system comprising: an EEG acquisition device and the electronic device described in the third aspect, wherein the EEG acquisition device and the electronic device are connected via an EEG data transmission interface, wherein:
[0061] The EEG acquisition device is configured to acquire the EEG signals of the user through one or more electrode channels, and transmit the EEG signals of the one or more electrode channels online to the electronic device based on the EEG data transmission interface.
[0062] Fifthly, this disclosure provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the method described in any one of the first aspects.
[0063] In a sixth aspect, this disclosure provides a computer program product, including a computer program, characterized in that the computer program, when executed by a processor, implements the method described in any one of the first aspects.
[0064] According to the technical solution provided in this disclosure, an upper computer receives the EEG signals of a user online from an EEG acquisition device, acquires one or more EEG signals of a specified frequency band, extracts the EEG features corresponding to the specified frequency band EEG signals, and, based on a specified audio representation and a specified audio generation method, segments the specified frequency band EEG signals into multiple EEG segments according to the EEG features or preset rhythm parameters, generating multiple audio representations corresponding to each of the multiple EEG segments. The feature parameters of the audio representations are determined based on the EEG features of the corresponding EEG segments. Brainwave audio representation data is generated based on the audio representations corresponding to the one or more specified frequency band EEG signals, and brainwave audio is obtained based on the brainwave audio representation data. When the brainwave audio is synchronously fed back to the user, the user's nervous system activity can be dynamically controlled in real time through auditory perception, achieving precise auditory representation of physiological signals. Furthermore, by determining the characteristic parameters of the corresponding audio performance based on the EEG characteristics of each EEG segment, the brainwave audio generated based on the audio performance can accurately match the EEG characteristics of the user being sampled, enhancing the physiological adaptability of neural feedback and surpassing the limitations of traditional preset music. At the same time, the EEG characteristics are generated based on EEG signals acquired online in real time. Thus, through closed-loop EEG-audio real-time feedback interaction, a high degree of synchronization between neural activity and auditory feedback is achieved, while simultaneously satisfying the physiological adaptability of neural regulation and the artistic expressiveness of audio generation, realizing the organic unity of neural regulation and audio generation.
[0065] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0066] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken in conjunction with the accompanying drawings. In the drawings:
[0067] Figure 1 A flowchart is shown for a proprioceptive brainwave audio auditory perception synchronous feedback method according to an embodiment of the present disclosure;
[0068] Figure 2 A flowchart is shown for another ontological brainwave audio auditory perception synchronous feedback method according to an embodiment of the present disclosure;
[0069] Figure 3 A schematic diagram of a network model for separating brain-derived signals and non-brain-derived signals according to an embodiment of the present disclosure is shown.
[0070] Figure 4 A flowchart of a mapping-symbol audio synthesis method according to an embodiment of the present disclosure is shown;
[0071] Figure 5A flowchart of a control-symbol audio synthesis method according to an embodiment of the present disclosure is shown;
[0072] Figure 6 A flowchart of a mapping-waveform audio synthesis method according to an embodiment of the present disclosure is shown;
[0073] Figure 7 A schematic diagram showing the segmentation results of EEG segments in the mapping-waveform audio synthesis method according to an embodiment of the present disclosure is illustrated.
[0074] Figure 8 A flowchart is shown for yet another proprioceptive brainwave audio auditory perception synchronous feedback method according to an embodiment of the present disclosure;
[0075] Figure 9 A schematic diagram showing the sound intensity location of brainwave audio in a specified frequency band according to an embodiment of the present disclosure is provided.
[0076] Figure 10 A schematic diagram of the structure of a proprioceptive brainwave audio auditory perception synchronous feedback device according to an embodiment of the present disclosure is shown.
[0077] Figure 11 A schematic diagram of the structure of another proprioceptive brainwave audio auditory perception synchronous feedback device according to an embodiment of the present disclosure is shown.
[0078] Figure 12 A schematic diagram of the structure of another proprioceptive brainwave audio auditory perception synchronous feedback device according to an embodiment of the present disclosure is shown.
[0079] Figure 13 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown;
[0080] Figure 14 A structural block diagram of another electronic device according to an embodiment of the present disclosure is shown;
[0081] Figure 15 A structural block diagram of another electronic device according to an embodiment of the present disclosure is shown;
[0082] Figure 16 A structural block diagram of another electronic device according to an embodiment of the present disclosure is shown;
[0083] Figure 17 A structural block diagram of a proprioceptive brainwave audio auditory perception synchronous feedback system according to an embodiment of the present disclosure is shown. Detailed Implementation
[0084] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of exemplary embodiments have been omitted from the drawings.
[0085] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.
[0086] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0087] As mentioned above, in traditional technologies, the application of brainwave audio is often separated into two independent directions: one is offline neuromodulation based on fixed templates, lacking real-time performance and individual adaptability; the other is music generation relying on subjective presets, which is difficult to truly reflect spontaneous brain activity. Therefore, it is impossible to simultaneously satisfy the physiological adaptability of neuromodulation and the artistic expressiveness of audio generation, failing to achieve the organic unity of neuromodulation and audio generation processes. Consequently, existing brainwave audio technologies cannot fully realize the therapeutic potential of neuromodulation, nor can they achieve creative audio expression based on real neural activity.
[0088] To overcome the functional limitations of existing technologies in the field of brainwave audio applications, this disclosure proposes a proprioceptive brainwave audio auditory perception synchronous feedback method. Through real-time online EEG signal acquisition, EEG feature extraction, dynamic brainwave audio generation, and feedback, it achieves an organic unity between neural modulation and audio generation. Here, "proprioception" refers to the user's own EEG signal, emphasizing the direct acquisition of the user's own EEG signals, rather than analog signals or those of others. Based on the user's unique neural activity patterns and real-time EEG characteristics, corresponding dynamic brainwave audio is generated, rather than a standardized "one-size-fits-all" approach. Then, through real-time "synchronous feedback" and "auditory perception," the brainwave audio feedback is directly applied to the signal provider (the user themselves), helping them perceive and regulate their own brain state (such as focus or relaxation), forming a "self-regulation" closed loop, rather than a one-way output (such as brain-controlled music composition), thus strengthening the role of "proprioception" in neural modulation.
[0089] In this disclosure, real-time brainwave audio of a specific individual directly stimulates the auditory cortex located in the proprioceptive brain through the auditory system. This triggers auditory perception and auditory cognitive responses in the central nervous system via auditory neural circuits, forming a synchronous feedback loop between proprioceptive brainwave audio and auditory perception. The difference between proprioceptive brainwave audio and non-proprioceptive brainwave audio lies in the fact that changes in brain network connectivity triggered by auditory perception and cognition cause changes in proprioceptive brainwave audio through the brainwave-audio feature mapping relationship, forming a closed-loop process of brainwave-audio-stimulation-brainwave. This fulfills the need for closed-loop brainwave audio regulation in real-time neuromodulation applications.
[0090] Figure 1 A flowchart illustrating a method for synchronous feedback of proprioceptive brainwave audio auditory perception according to an embodiment of the present disclosure is shown. The method is applied to a host computer, which is connected to an EEG acquisition device via an EEG data transmission interface.
[0091] like Figure 1 As shown, the proprioceptive brainwave audio auditory perception synchronous feedback method includes the following steps S110 to S150:
[0092] In step S110, the EEG signal of the user being collected is received online from the EEG acquisition device through the EEG data transmission interface. The EEG signal is acquired by the EEG acquisition device through one or more electrode channels, wherein each electrode channel corresponds to one electrode or a combination of multiple electrodes.
[0093] The EEG acquisition device disclosed herein is a specialized device for receiving the electrical activity of neurons on the scalp. It collects weak EEG signals from the scalp surface through electrodes, and then processes the raw signals by amplification, filtering, analog-to-digital conversion, etc. The processed EEG signals are then transmitted to a host computer for analysis and processing through an EEG data transmission interface.
[0094] The EEG acquisition device disclosed herein can be a portable / wearable EEG acquisition device (such as a dry electrode headband device) or a large fixed device, depending on the specific design and application scenario. In a specific example, the EEG acquisition device of this disclosure adopts a portable EEG acquisition device and a real-time EEG data transmission interface to realize real-time acquisition and data acquisition of scalp EEG. This portable EEG acquisition device is a multimodal human signal acquisition system, designed specifically for real-time, lightweight neural signal monitoring. It can support the simultaneous acquisition of multiple modalities of human electrical signals and adopts wireless transmission (such as Bluetooth / Wi-Fi) or a lightweight wired design. The EEG signal uses five electrode channels on a single frontal lobe electrode, corresponding to the specific electrode positions {F7, Fp1, Fpz, Fp2, F8}, and the electrode positions meet the international standard 10 / 20 electrode arrangement, with a sampling rate of 1000Hz. Based on the device's accompanying software development kit (SDK), low-latency and high-reliability transmission of EEG signals from the acquisition device to the host computer is achieved, enabling real-time transmission of EEG signals from the acquisition device to the host computer. The transmission interval is set to 30ms when transmitting EEG signals.
[0095] In this disclosure, "electrode channel" refers to an independent signal acquisition pathway in an EEG acquisition system. Each electrode channel (e.g., F7, Fp1, etc.) records the comprehensive electrical signal of neuronal activity at that location. The correspondence between electrodes can be one-to-one, i.e., a single electrode corresponds to an independent electrode channel (e.g., Fpz electrode → Channel 1), directly recording the potential at that point (relative to the reference electrode); or it can be many-to-one, i.e., multiple electrodes combine to form an electrode channel (e.g., F7+F8 → bipolar lead Channel X), recording the potential difference between two points. In the specific example above, each recording electrode is an independent channel (a total of 5 electrode channels), and the reference electrode is typically the earlobe (A1 / A2) or an average reference. Although electrodes at different locations may be more sensitive to certain frequency bands, generally, the EEG signals acquired by each electrode channel cover all frequency bands (Delta, Theta, Alpha, Beta, Gamma).
[0096] When the host computer in this disclosure is connected to the EEG acquisition device through the EEG data transmission interface, it can be directly connected to the EEG acquisition device through the EEG data transmission interface. In this case, the host computer is the corresponding accessory device for the EEG acquisition device and has the processing capability to generate EEG audio based on the acquired EEG signals by executing steps S110 to S150. Alternatively, it can be indirectly connected to the EEG acquisition device through the EEG data transmission interface based on the accessory device corresponding to the EEG acquisition device. That is, the EEG acquisition device is connected to its accessory device through the EEG data transmission interface, and then its accessory device is connected to the host computer through a wireless or wired communication connection so as to transmit the acquired EEG signals to the host computer for analysis and processing. In this case, the host computer corresponds to the EEG signal processing terminal.
[0097] When acquiring a user's EEG signals using the EEG acquisition device in the specific example above, the device collects EEG signals through five electrode channels at a frequency of 1000 data points per second. It then sends a data packet to the host computer every 30ms. Each transmission contains a continuous signal of 30ms × 1000Hz = 30 time points. The data format can be a numerical matrix, such as a 5×30 matrix, with each row corresponding to the voltage sequence of one electrode channel. When acquiring the EEG signal, the host computer extracts the EEG signals from each electrode channel by parsing the received numerical matrix.
[0098] Typically, EEG signals collected by EEG acquisition devices are highly susceptible to contamination from artifacts such as electrooculography (EOG) and electromyography (EMG). If EOG is not removed, blinking may cause sudden high-pitched noise in the audio, while failing to remove EMG may cause the generated audio rhythm to be interfered with by jaw muscle activity. Therefore, in cases where EEG acquisition devices and their SDKs do not include real-time denoising modules for EOG and EMG, it is necessary to denoise the EEG signals before subsequent analysis and processing to ensure that the signals reflect genuine neural activity rather than physiological artifacts, and to avoid misjudgment or misoperation in scenarios such as medical diagnosis and BCI control, thus guaranteeing the effectiveness of neural modulation and audio generation. This denoising process yields brain-derived signals free of EOG and EMG components.
[0099] Figure 2 A flowchart illustrating another ontological brainwave audio auditory perception synchronization feedback method according to an embodiment of this disclosure is shown. Figure 2 As shown, before performing step S120, that is, before acquiring one or more specified frequency band EEG signals from the EEG signals, the following step S111 is also included:
[0100] In step S111, the EEG signals of one or more electrode channels are preprocessed using a pre-trained brain-derived signal and non-brain-derived signal separation network model to obtain denoised EEG signals.
[0101] Current technologies typically use Independent Component Analysis (ICA) based on blind source separation to decompose mixed signals into independent components, and then manually or automatically remove electrooculography (EOG) / electromyography (EMG) components. While ICA can separate complex mixed signals, it requires offline analysis and has high computational complexity. Furthermore, it has certain requirements regarding the number of independent electrodes and signal length (usually requiring 1-5 minutes of data segments), resulting in high latency. This makes it difficult to meet the requirements of real-time EEG processing and unsuitable for real-time automated applications.
[0102] This disclosure utilizes deep learning technology to provide a network model for separating brain-derived signals from non-brain-derived signals. The pre-trained network model can remove noise components such as electrooculography (EOG) and electromyography (EMG) from EEG signals in real time, automatically, and effectively.
[0103] Figure 3 This diagram illustrates the structure of a network model for separating brain-derived signals from non-brain-derived signals according to an embodiment of the present disclosure. Figure 3 As shown, the brain-derived signal and non-brain-derived signal separation network model includes: a denoising module and a skip connection module. The denoising module includes: a decomposition module, a channel spatiotemporal attention processing module, and a reconstruction module.
[0104] According to embodiments of this disclosure, the preprocessing of the EEG signals from one or more electrode channels using a pre-trained brain-derived signal and non-brain-derived signal separation network model to obtain denoised EEG signals includes:
[0105] The noise reduction module removes non-brain-derived signals from the EEG signals of one or more electrode channels. The specific process includes the following steps:
[0106] First, the electroencephalogram (EEG) signals of the one or more electrode channels are decomposed into multidimensional embedding vectors corresponding to multiple signal channels by the decomposition module. The multiple signal channels include brain-derived signal channels and non-brain-derived signal channels.
[0107] The non-brain-derived signals include, but are not limited to, electrooculography (EOG) signals and / or electromyography (EMG) signals.
[0108] Specifically, the EEG signals from one or more electrode channels are first mapped to a multidimensional space using a set of learnable convolutional kernels, as follows:
[0109] V = Conv1D(X, W) d )∈R D×T ;
[0110] Where X is the EEG signal of the one or more electrode channels, X∈R C×TC is the number of electrode channels, T is the number of time points, and W d W is the convolution kernel weight matrix. d ∈R D×C×K D represents the number of signal channels, and K represents the convolution kernel duration. The convolution operation stride defaults to 1, and padding is set to "same" to maintain the time dimension T. Conv1D is a one-dimensional convolution operation, a local feature extraction operation for time-series signals, used to map EEG signals to a multi-dimensional embedding space (D-dimensional), implicitly separating different signal components (such as brain-derived / non-brain-derived), where V is the multi-dimensional embedding vector.
[0111] Then, the multidimensional embedding vector is separated into embedding vectors V of brain-derived signal channels according to a preset ratio. b Embedding vector V of non-brain-derived signaling channels n ,in, D b D represents the number of brain-derived signal channels. n The number of non-brain-derived signal channels, D = D b +D n V = V b +V n .
[0112] Subsequently, the multidimensional embedding vector is processed by channel attention and spatiotemporal attention through the channel spatiotemporal attention processing module to generate channel attention weights and temporal attention weights. The generated channel attention weights and temporal attention weights are then fused with the multidimensional embedding vectors of the corresponding multiple signal channels to obtain the processed multidimensional embedding vector.
[0113] In this disclosure, the channel spatiotemporal attention processing module enhances the weights of brain-derived signal channels through attention mechanisms and captures time dependence. Specifically, it includes channel attention processing and spatiotemporal attention processing.
[0114] When performing channel attention processing, the channel attention weight α is calculated using the following formula:
[0115] α=σ(W2·ReLU(W1·GAP(V)));
[0116] Where GAP(·) represents global average pooling, and W1 and W2 represent the first and second fully connected layer weights, respectively. The first fully connected layer weights are used to reduce the dimensionality of the channel descriptors after global average pooling (GAP) to a lower-dimensional space. The design purpose is to reduce computational cost and extract nonlinear interaction information between channels. The second fully connected layer weights are used to remap the dimensionality-reduced features back to the original dimension. The design purpose is to reconstruct the dependencies between signal channels and generate channel-wise attention weights, where W1∈R. D / r×D W2∈R D×D / r r represents the reduction ratio, ReLU is the corrected linear unit used to output the nonlinear features after dimensionality reduction, and σ represents the Sigmoid activation function.
[0117] The role of introducing ReLU in channel attention includes:
[0118] 1. Introducing nonlinearity: This enables the model to learn complex interactions between signal channels, rather than simple linear combinations.
[0119] 2. Sparse activation: Suppresses negative features and highlights the contribution of important signal channels.
[0120] 3. Prevent gradient vanishing: Compared to Sigmoid / Tanh, ReLU's gradient is always 1 in the positive interval, alleviating the training problems of deep networks.
[0121] When performing spatiotemporal attention processing, linear projection is first performed to generate a query matrix Q, a key matrix K, and a value matrix V. value Specifically, the following formula is used:
[0122] Q = VW q ;
[0123] K = VW k ;
[0124] V value =VW v ;
[0125] Among them, W q W k and W v These are the projected weights, which are automatically learned through the training process (such as backpropagation and gradient descent). d k For the attention dimension, usually d k =D / 2.
[0126] Then, the time attention weight β is calculated using the following formula:
[0127]
[0128] Then, the signal weights and temporal attention weights of each generated signal channel are fused with the corresponding multidimensional embedding vector using the following formula to obtain the processed multidimensional embedding vector V. out :
[0129] V channel =α⊙V;
[0130] V time =βV value ∈R D×T ;
[0131] V out =V+V channel +V time ;
[0132] Here, ⊙ represents element-wise dot product.
[0133] Next, the reconstruction module generates the reconstructed EEG signal X based on the processed multidimensional embedding vector using the following formula. rec :
[0134] X rec =ConvTranspose1D(V out W r )∈R C×T ;
[0135] Wherein, ConvTranspose1D is a one-dimensional deconvolution operation used to reconstruct the original EEG signal space from a multi-dimensional embedding vector, W r For deconvolution kernel, W r ∈R C×D×K .
[0136] After obtaining the reconstructed EEG signal through the above steps, the processed multidimensional embedding vector is obtained through the skip connection module. Based on the processed multidimensional embedding vector and the EEG signal of one or more electrode channels, dynamic fusion parameters are generated collaboratively using the gated linear unit (GLU) and the noise perception module (NAM). The dynamic fusion parameters include: preliminary weights and noise level.
[0137] Specifically, the initial weights γ are first generated using a gated linear unit (GLU) according to the following formula:
[0138] γ=σ(W g ·[V out ;X])+b g ;
[0139] Among them, W g W is the weight matrix of GLU. g ∈RC×(D+C) ;[V out ;X] is used to convert V out spliced with the EEG signal X from one or more electrode channels, [V out ;X]∈R (D+C)×T b g For the bias term, b g ∈R C σ is the Sigmoid function, used to compress the initial weights to [0, 1].
[0140] Then, the noise level η is estimated by analyzing the residual signal using the noise sensing module NAM with the following formula:
[0141] η = MLP(Std(XX) rec ));
[0142] Where, η∈R G×T Std(·) represents the calculation of the noise standard deviation of each signal channel, and MLP stands for Multilayer Perceptron. The larger the value of η for quantizing the noise level, the more significant the noise in that signal channel, and the more necessary it is to reduce the fusion ratio of the EEG signals from one or more electrode channels.
[0143] Finally, the skip connection module fuses the EEG signals of the one or more electrode channels with the reconstructed EEG signals according to the dynamic fusion parameters to obtain denoised EEG signals. The dynamic fusion parameters are used to control the fusion ratio of the EEG signals of the one or more electrode channels with the reconstructed EEG signals.
[0144] Specifically, the denoised EEG signal X can be obtained using the following formula based on the initial weight γ and the noise level η. final :
[0145] X final =γ⊙X rec +(1-γ-ρη)⊙X;
[0146] Where ρ represents the noise suppression coefficient, which is used to control the intensity of noise perception.
[0147] The brain-derived and non-brain-derived signal separation network model disclosed herein automatically identifies noise / clean signal segments through a collaborative mechanism of deep decomposition, channel attention, and dynamic fusion. This eliminates the need for pre-labeling of noise segments, achieving end-to-end real-time adaptive processing. Compared to traditional ICA, it offers advantages in computational efficiency (real-time performance), adaptive capability (automation), and signal-to-noise ratio improvement (effectiveness), making it particularly suitable for online EEG processing systems. Furthermore, the skip connection module employs a technique that uses GLU (Gated Linear Unit) and NAM (Noise Awareness Module) to collaboratively generate dynamic fusion parameters. This not only allows noise-free EEG signal segments to pass directly through the network model, avoiding distortion caused by unnecessary processing, but also exhibits significant advantages in multiple dimensions compared to traditional methods based on average amplitude and fully connected layers. Regarding dynamic adaptability, GLU generates weights time-by-time by fusing spatiotemporal attention features and the original signal, accurately addressing transient noise (such as blink artifacts), while traditional methods only output static channel weights and cannot handle time-varying interference. Furthermore, traditional methods rely solely on amplitude mean, making it difficult to distinguish between valid signals and noise. In contrast, the present invention utilizes NAM to explicitly quantify noise intensity by analyzing the local standard deviation of the residual signal, thereby specifically suppressing high-noise regions and improving noise perception capabilities. Additionally, traditional methods are susceptible to failure due to abrupt amplitude changes. However, the present invention employs residual analysis to effectively resist sudden interference such as electrode loosening, enhancing its anti-interference robustness.
[0148] After obtaining the denoised EEG signal based on step S111, it is used in subsequent steps to analyze the denoised EEG signal to extract various features (such as EEG features, brain network features, and source space features), and is also used to dynamically display the waveform of the denoised EEG signal.
[0149] According to embodiments of this disclosure, the method further includes: dynamically displaying the waveform of the denoised EEG signal in real time on a visualization interface, so as to provide testers or experts with an interactive feedback interface for intuitive visual observation of changes in brain waves at different frequency bands, which is beneficial for testers to observe changes in brain waves at different frequency bands of the subject, thereby providing a basis for generating targeted stimulation programs.
[0150] In a specific example, a web page dynamically displays the waveform of a five-channel EEG signal after separating noise from electrooculography (EOG), electromyography (EMG), and other signals. The waveform refresh rate is f. r Determined by the EEG acquisition device, the selectable range is [2Hz, 8Hz], and the display front-end feedback delay t delay The range is
[0151] In step S120, one or more specified frequency band EEG signals are acquired from the EEG signals, and EEG features corresponding to the specified frequency band EEG signals are extracted based on the one or more specified frequency band EEG signals. The specified frequency band EEG signals refer to EEG signals located in the specified frequency band, and the EEG features include EEG time domain features and / or EEG frequency domain features.
[0152] Typically, EEG signals comprise five typical frequency bands: Delta (1-4 Hz), Theta (4-8 Hz), Alpha (8-13 Hz), Beta (13-30 Hz), and Gamma (30-48 Hz). In this disclosure, the EEG signal corresponding to each electrode channel may encompass a mixture of multiple or all frequency bands. When acquiring one or more specified frequency band EEG signals, a bandpass filter can be used to separate the signals in each band to obtain the specified frequency band EEG signal.
[0153] The EEG time-domain features include, but are not limited to, time-domain envelope. Unlike traditional EEG signal time-domain feature extraction methods, this disclosure does not focus on related EEG amplitude changes, but rather on the EEG signal itself, extracting the EEG time-domain envelope corresponding to the EEG signal in a specified frequency band. Since the EEG time-domain envelope reflects the change in EEG signal intensity over time in that frequency band, the EEG audio generated using this EEG feature has a more distinctive "brainwave color".
[0154] When extracting the time-domain envelope corresponding to the EEG signal of the specified frequency band based on one or more specified frequency band EEG signals, an envelope extraction algorithm (such as Hilbert transform) can be used to extract the time-domain envelope corresponding to the EEG signal of the specified frequency band.
[0155] The EEG frequency domain features include, but are not limited to, EEG power spectrum features. In this disclosure, when extracting EEG power spectrum features corresponding to the EEG signals of the specified frequency bands based on the one or more specified frequency band EEG signals, a refined EEG power spectrum feature corresponding to the EEG signals of the specified frequency bands is extracted by a band-selective spectrum transformation algorithm. The band-selective spectrum transformation algorithm includes, but is not limited to, any of the following transformation algorithms: Chirp-Z transform, Zoom-FFT, Goertzel algorithm, and Continuous Wavelet Transform (CWT).
[0156] In this disclosure, considering that the effective frequency band of scalp EEG signals is between 0 and 50 Hz, while the frequency range of a specified band is even narrower and much smaller than the device's sampling rate (e.g., 1000 Hz), using uniform sampling Discrete Fourier Transform (DFT) would result in a large number of out-of-band results, leading to wasted computation and reduced frequency domain resolution. Therefore, this disclosure uses a band-selective spectral transform algorithm (e.g., Chirp-z transform) instead of the conventionally used DFT in the prior art to obtain a refined power spectral density in a specified frequency band. By focusing on the target frequency band, the frequency resolution is improved, making it more suitable for analyzing the narrowband oscillation characteristics of EEG. Refining the EEG power spectral characteristics refers to performing high-resolution spectral estimation for a specified frequency band (e.g., the Alpha band 8-13 Hz) based on traditional power spectral analysis, in order to more accurately capture the subtle frequency components of the EEG signal.
[0157] According to embodiments of this disclosure, the step of extracting refined EEG power spectrum features corresponding to the specified frequency band EEG signal using a frequency band-selective spectrum transformation algorithm includes:
[0158] First, the rotation factor and starting point are calculated based on the start and end frequencies of the specified frequency band and the target resolution.
[0159] Specifically, the rotation factor W is calculated using the following formula:
[0160]
[0161] Calculate the starting point A using the following formula:
[0162]
[0163] Among them, f s Where f1 is the sampling rate, f2 is the start frequency of the specified frequency band, M is the number of output points within the specified frequency band, and the target resolution is expressed as:
[0164] Then, based on the rotation factor and the starting point, the frequency band selective spectrum transformation algorithm is performed on the specified frequency band EEG signal to obtain the spectrum corresponding to the specified frequency band EEG signal.
[0165] Taking the Chirp-z transform (CZT) as an example, it is specifically implemented through the following steps:
[0166] 1. Construct the preprocessing sequence g[n]:
[0167]
[0168] 2. Construct the convolution kernel h[n]:
[0169]
[0170] 3. Perform FFT operations on g[n] and h[n] respectively:
[0171] G[k] = FFT(g[n]);
[0172] H[k] = FFT(h[n]);
[0173] 4. Frequency domain multiplication:
[0174] Y[k]=G[k]·H[k];
[0175] 5. Obtain the convolution result y[n] by performing an inverse FFT:
[0176] y[n] = IFFT(Y[k]);
[0177] 6. Correct the output and take the first M points:
[0178]
[0179] Wherein, s[n] is the specified frequency band EEG signal, and N is the length of the specified frequency band EEG signal.
[0180] Thus, the CZT transformation result of the specified frequency band EEG signal s[n] is obtained, namely: the spectrum S[k], where S[k] is a complex number representing the frequency domain representation of the EEG signal at the refined frequency point k.
[0181] Finally, the refined EEG power spectrum feature P[k] corresponding to the specified frequency band EEG signal is calculated based on the spectrum using the following formula:
[0182] P[k]=A[k] 2 ;
[0183]
[0184] Where A[k] is the amplitude spectrum corresponding to the specified frequency band EEG signal, and the refined EEG power spectrum feature P[k] can be directly used for: frequency band energy statistics (e.g., calculating the total energy of the frequency band), feature classification (e.g., extracting the power of a specific frequency point as a classification feature) and time-frequency analysis (e.g., combining a sliding window to achieve dynamic power spectrum tracking).
[0185] In step S130, for the specified frequency band EEG signal, based on the specified audio representation form and the specified audio generation method, the specified frequency band EEG signal is segmented into multiple EEG segments according to the EEG characteristics or preset rhythm parameters, and multiple audio representations corresponding to the multiple EEG segments are generated. The feature parameters of the audio representation are determined according to the EEG characteristics of the corresponding EEG segments. The audio representation form includes: note form and / or waveform form, wherein the audio representation corresponding to the note form is a note, and the audio representation corresponding to the waveform form is a waveform. The audio generation method includes: direct mapping method and / or indirect control method using control parameters.
[0186] In step S140, brainwave audio representation data is generated based on the audio representations corresponding to the one or more specified frequency band EEG signals.
[0187] In step S150, brainwave audio is obtained based on the brainwave audio representation data.
[0188] This disclosure, based on steps S110-S120 above, obtains EEG features corresponding to a specified frequency band EEG signal, and then, based on steps S130-S150 above, generates corresponding brainwave audio based on these EEG features. For the brainwave audio generation process, this disclosure implements two real-time EEG-audio generation paths: a symbolic generation path and a waveform generation path. For the symbolic generation path, the final generated audio representation is in the form of musical notes. The characteristic parameters of the audio representation corresponding to the musical note form include, but are not limited to, note duration, pitch, and intensity. For the waveform generation path, the final generated audio representation is in the form of waveforms. The characteristic parameters of the audio representation corresponding to the waveform form include, but are not limited to, the fundamental frequency and harmonic frequencies, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequencies. This disclosure also implements two different audio parameter generation methods: one is control-based generation, which uses the results of an EEG-related discriminant model to control the selection of audio generation parameters, referred to as the "indirect control method." This indirectly controls note parameters (pitch, rhythm, etc.) through a discriminant model (such as classifying attention states), achieving goal-oriented neural feedback and thus improving the accuracy of neural modulation. The second is mapping-based generation, which directly maps EEG features to audio features through a certain mapping relationship, without using other intermediate calculation results to control the mapping parameters, referred to as the "direct mapping method." By directly mapping EEG features, the specific neural electrical activity states of different individuals are intuitively reflected in the audio, meeting the needs of individualized generation and regulation. Furthermore, the parameter generation process of this method is shorter than that of the discriminant model, thus avoiding model computation delays and making it suitable for ultra-real-time response scenarios. The note-form output is more consistent with traditional audio cognition, making it easier for users to understand the relationship between neural states and audio. Waveform synthesis can generate complex soundscapes, suitable for emotion regulation. For the symbol generation path, the audio parameter generation method is divided into control-type generation and mapping-type generation; for the waveform generation path, the parameter generation method is mapping-type generation.
[0189] Therefore, this disclosure provides three audio synthesis paths: a mapping-symbol synthesis path, a control-symbol synthesis path, and a mapping-waveform synthesis path. Specifically, when the specified audio representation is in note form and the specified audio generation method is an indirect control method using control parameters, the corresponding audio synthesis path is a control-symbol synthesis path; when the specified audio representation is in note form and the specified audio generation method is a direct mapping method, the corresponding audio synthesis path is a mapping-symbol synthesis path; and when the specified audio representation is in waveform form and the specified audio generation method is a direct mapping method, the corresponding audio synthesis path is a mapping-waveform synthesis path.
[0190] In practical implementation, the audio representation and audio generation method can be selected based on user-input selection parameters or selection parameters automatically generated by the system within a preset period, thereby determining the corresponding audio synthesis path. The user-input selection parameters can be set according to the application scenario and / or user preferences. For example, for medical rehabilitation (such as ADHD training) scenarios, a control-symbol synthesis path can be selected, using a classification model to control note parameters in a closed loop; for artistic creation scenarios, a mapping-symbol path can be selected, directly mapping EEG to MIDI notes. When the selection parameters are automatically generated by the system, they can be generated based on the current neural state and / or signal quality assessment parameters. For example, in a focused state, a control-symbol generation path can be selected to increase note density; in a relaxed state, a mapping-waveform generation path can be selected to generate soothing audio. Alternatively, the audio representation and audio generation method can be selected based on a specified frequency band within the system, thereby determining the system's default audio synthesis path. For example, EEG signals in the Delta / δ and Theta / θ frequency bands use a mapping-symbol synthesis path, while EEG signals in other frequency bands use a mapping-waveform synthesis path.
[0191] For EEG signals from multiple electrode channels, after obtaining the EEG signals of each specified frequency band in each electrode channel or multiple specified electrode channels based on the above step S120, the corresponding audio representation (note or audio waveform) is generated based on the EEG signals of each specified frequency band. Finally, brainwave audio is generated based on the audio representation corresponding to the EEG signals of each specified frequency band in each electrode channel.
[0192] The following is a detailed description of each audio synthesis path.
[0193] Assume that for an EEG(t) signal with multiple electrode channels, the number of electrode channels is N, and the specified frequency band set is Bands = {δ[1-4Hz), θ[4-8Hz), α[8-12Hz), β[12-30Hz), γ[30-48Hz)}.
[0194] First, the EEG signal of each specified frequency band in each electrode channel is segmented into non-overlapping segments according to a preset duration (e.g., 4 seconds). Then, a note sequence (for mapping-symbol synthesis paths and control-symbol synthesis paths) or audio waveform (for mapping-waveform synthesis paths) is generated for each segment. Each note sequence or audio waveform corresponds to a single specified frequency band EEG signal in a single electrode channel. These note sequences or audio waveforms constitute the EEG audio corresponding to the preset duration. The preset duration is the processing window length of the EEG signal, typically an integer number of segments.
[0195] For mapping-symbol composition paths:
[0196] This process maps EEG features corresponding to EEG signals to audio features, thereby generating corresponding musical notes and scores, which are then used to sample and play the music. To improve real-time performance during EEG-audio mapping, EEG signals from specific frequency bands in each electrode channel can be processed in parallel.
[0197] Figure 4 A flowchart of a mapping-symbol audio synthesis method according to an embodiment of the present disclosure is shown. For each specified frequency band of EEG signal in each electrode channel, an EEG signal of a preset duration is processed as follows: Figure 4 The following steps S410 to S460 are shown:
[0198] For a pre-set duration of EEG signal in a specified frequency band of a specified electrode channel, perform the following steps S410 to S420:
[0199] In step S410, the EEG signal of the preset duration is segmented according to the note duration indicated by the preset rhythm parameters to obtain a set of EEG segments. The set of EEG segments contains multiple EEG segments, and each EEG segment corresponds to a note.
[0200] The rhythm parameters are a set of rules used to control the time structure and rhythmic pattern during the conversion of EEG signals into musical notes, ensuring that the generated audio has a reasonable beat, duration, and rhythm, including but not limited to: BPM (beats per minute), rhythm template, minimum note duration, and number of note segments. For example: BPM = 120, minimum note duration = 16th note (0.125s), rhythm template = 4 / 4 time signature.
[0201] For example: for electrode channel i and a specified frequency band δ, the corresponding EEG signal is EEG_i_δ(t). After segmentation according to preset rhythm parameters, the resulting EEG segment set is notelist = {s}. note1 ,s note2 ,…,s noteN}
[0202] In step S420, for any EEG segment in the set of EEG segments, the pitch and intensity of the note corresponding to any EEG segment are determined according to the EEG frequency domain characteristics corresponding to the any EEG segment, until the pitch and intensity of the note corresponding to each EEG segment in the set of EEG segments are obtained.
[0203] When determining the pitch and intensity of the note corresponding to any given EEG segment based on the frequency domain characteristics of the EEG segment, the following steps S421 to S422 are included:
[0204] In step S421, based on the power spectrum characteristics corresponding to any EEG segment, the peak frequency or weighted average frequency with the largest amplitude in the power spectrum characteristics is selected as the characteristic frequency.
[0205] The power spectrum features corresponding to any EEG segment include: refined power spectrum features obtained through a band-selective spectrum transform algorithm, such as: CZT power spectrum features sp note Its peak frequency is f peak Weighted average frequency f weight It can be calculated using the following formula:
[0206]
[0207] Among them, f k For any given EEG segment, the k-th frequency point (Hz) in the power spectrum is sp note (f k ) represents the power spectral amplitude at the k-th frequency point.
[0208] In step S422, based on a preset mapping method, a specified pitch range is obtained according to the frequency range of the specified frequency band, and the pitch corresponding to any EEG segment is obtained according to the characteristic frequency and the specified pitch range; the sound intensity corresponding to any EEG segment is generated according to the total energy or peak amplitude of the power spectrum feature; the preset mapping method is used to describe the mapping relationship between the frequency range and the pitch range of the specified frequency band.
[0209] When determining the pitch of a note, the pitch range to which it is mapped can first be divided according to the specified frequency band. Taking the mapping to the MIDI pitch range as an example, the MIDI note numbers are 0-127, corresponding to C-1 to G9. Then: the MIDI pitch range corresponding to the EEG signal of the Delta & Theta frequency band is [48,59], the MIDI pitch range corresponding to the EEG signal of the Alpha frequency band is [56,64], the MIDI pitch range corresponding to the EEG signal of the Beta frequency band is [65,88], and the MIDI pitch range corresponding to the EEG signal of the Gamma frequency band is [88,96].
[0210] When determining the preset mapping method, you can choose between linear or logarithmic methods based on the application scenario and specific needs. Taking the linear method as an example, pitch mapping can be performed using the following formula:
[0211]
[0212] Among them, MIDI 音高 The pitch of the note corresponding to any of the EEG segments, MIDI min This represents the minimum value in the MIDI pitch range.max f is the maximum value of the MIDI pitch range. min f is the minimum value of the specified frequency band range. max f is the maximum value of the specified frequency band range, and f is the characteristic frequency.
[0213] In special cases, if the power spectrum amplitude corresponding to any of the EEG segments is lower than the threshold, the forced pitch is 0 (a rest).
[0214] When determining the intensity of a note, the total energy E of the power spectrum characteristic is first determined. total or peak amplitude a peak E total =∑ k sp note (f k ), A peak =max(sp note (f k Then, the sound intensity corresponding to any one of the EEG segments is generated based on the total energy or peak amplitude of the power spectrum feature. When choosing whether to use the total energy or peak amplitude of the power spectrum feature, the following conditions can be met: if the EEG signal is stable (e.g., alpha wave), the total energy is used first; if transient features (e.g., gamma wave burst) need to be captured, the peak amplitude is used.
[0215] When generating the sound intensity corresponding to any EEG segment based on the total energy or peak amplitude of the power spectrum characteristics, a preset sound intensity range is first determined, for example, setting the sound intensity range to a continuous range of normalized values from 0 to 1; then, a preset mapping method is determined, for example:
[0216] When the total energy based on the power spectrum characteristics generates a sound intensity S corresponding to any of the EEG segments, energy When this happens, the following formula can be used for calculation:
[0217]
[0218] When the sound intensity S corresponding to any EEG segment is generated based on the peak amplitude of the power spectrum characteristics... peak When this happens, the following formula can be used for calculation:
[0219]
[0220] Among them, E min and E max These are the minimum and maximum energy thresholds, respectively, calibrated using historical data or experiments. min and A max These are the minimum and maximum calibration thresholds for the peak amplitude, respectively.
[0221] Once the pitch and intensity of the notes corresponding to each EEG segment in the set of EEG segments are determined, the notes corresponding to each EEG segment in the set of EEG segments are determined.
[0222] Repeat steps S410 to S420 to obtain the notes corresponding to each EEG segment in the set of EEG segments of the same preset duration in other specified frequency band EEG signals of each specified electrode channel.
[0223] In step S430, the notes corresponding to each EEG segment in the set of EEG segments of a preset duration in the specified frequency band of the EEG signal of each electrode channel are written into the musical score corresponding to the specified frequency band.
[0224] The purpose of step S430 is to organize all the notes in a single frequency band of each electrode channel into a structured musical score in chronological order. The score for each frequency band consists of a series of notes, which includes information such as the start time, pitch, intensity, duration, electrode channel number, and frequency band identifier for each note.
[0225] In one specific example, the synthesized score is stored in MIDI (Music Instrument Digital Interfacing) format.
[0226] Repeat steps S410 to S430 to obtain the musical scores corresponding to other specified frequency bands of each specified electrode channel.
[0227] In step S440, the musical scores corresponding to each specified frequency band in each electrode channel are written into a complete musical score set corresponding to the preset duration, and the complete musical score set is used as the brainwave audio representation data.
[0228] Among them, brainwave audio representation data is an intermediate data format that converts the time-domain / frequency-domain characteristics of electroencephalogram (EEG) signals into audible audio.
[0229] The purpose of step S440 is to merge the scores of all electrode channels and frequency bands to generate a complete score set with time synchronization. This complete score set contains the scores corresponding to all specified frequency bands across all electrode channels.
[0230] In step S450, the timbre is configured for the musical score corresponding to each specified frequency band in the brainwave audio representation data according to the preset frequency band-timbre mapping table or configuration parameter-timbre mapping table.
[0231] The frequency band-timbre mapping table is used to describe the correspondence between frequency bands and timbres, and the configuration parameter-timbre mapping table is used to describe the correspondence between specified timbre configuration parameters and timbres. The specified timbre configuration parameters include: non-brain-derived control parameters and / or non-brain-derived event occurrence frequencies.
[0232] The configured timbres include MIDI instrument timbres. As mentioned above, these can be configured according to a timbre mapping table. This disclosure provides two types of timbre mapping tables for mapped-symbol synthesis paths: a frequency band-timbre mapping table and a configuration parameter-timbre mapping table. The frequency band-timbre mapping table defines the correspondence between frequency bands and timbres (instruments), allowing different timbres to be configured for different specified frequency bands. For example, a bass timbre can be configured for a score corresponding to the δ frequency band, and a violin timbre for a score corresponding to the β frequency band. The configuration parameter-timbre mapping table defines the correspondence between specified timbre configuration parameters and timbres (instruments), allowing the timbres used in each frequency band to be obtained based on different timbre configuration parameters. In specific applications, one type of timbre mapping table can be selected to configure the timbres according to the needs of the current scenario.
[0233] Specifically, the non-brain-derived control parameters are obtained as follows:
[0234] When the EEG signal of the preset length in the designated electrode channel contains non-brain-derived signals and denoised EEG signals, the non-brain-derived control parameter is obtained by the ratio of the preset fractional part of the absolute values of the non-brain-derived signals and the denoised EEG signals in the EEG signal of the preset length in the designated electrode channel.
[0235] The designated electrode channel can be all electrode channels, or any one or several electrode channels among all electrode channels. Thus, when obtaining the non-brain-derived signal control parameter based on the ratio of the absolute values of the non-brain-derived signal and the denoised EEG signal in the EEG signal of the preset length from the designated electrode channel, the non-brain-derived signal control parameter can be the average of the preset ratios of the absolute values of the non-brain-derived signal and the denoised EEG signal in the EEG signal of the preset length from all electrode channels, or the maximum value among all preset ratios, or a preset ratio using a specific electrode channel of interest. The preset ratio can be the 90th percentile ratio, and can be set according to requirements. The non-brain-derived signal may include, but is not limited to, electrooculogram (EOG) signals.
[0236] The calculation process of the non-brain-derived control parameters is described in detail below:
[0237] For an EEG signal of a preset length from a specified electrode channel, if it contains non-brain-derived signals and denoised EEG signals, the pre-trained brain-derived signal and non-brain-derived signal separation network model described above is first used to preprocess the EEG signal of the preset length from the specified electrode channel to obtain the denoised EEG signal. brain Then, based on the preset length of the EEG signal and the denoised EEG signal from the designated electrode channel... brain The difference yields the non-brain-derived signal EOG. Taking the ratio of the 90th percentiles of the electrode channels of interest as an example, the non-brain-derived control parameter AR can be calculated using the following formula:
[0238] AR = EOG 90 / EEG brain90 ;
[0239] AR represents non-brain-derived signal EOG and denoised EEG signal. brain The ratio of the 90th percentiles after taking the absolute values of each signal is used. The 90th percentile is defined as the value at the 90th position after arranging the signal data in ascending order (i.e., the threshold of the top 10% of high amplitude values). EOG 90 EEG represents the 90th percentile after taking the absolute value of a non-brain-derived signal (EOG). brain90 This represents the denoised EEG signal. brain The 90th percentile after taking the absolute value.
[0240] The non-brain-derived event frequency is obtained as follows: a non-brain-derived event threshold is obtained by taking the absolute value of the denoised EEG signal and dividing it by a preset fractional number; positions in the non-brain-derived signal that are higher than the non-brain-derived event threshold are counted to obtain the number of non-brain-derived events occurring in the non-brain-derived signal; and the non-brain-derived event frequency is obtained based on the number of non-brain-derived events and the preset duration.
[0241] The non-brain-derived control parameter was also defined as the ratio of the 90th percentile of the electrode channels of particular interest.
[0242] Example: First, the non-brain-derived event threshold Th can be calculated using the following formula:
[0243] Th=a*EEG braih90 ;
[0244] Where 'a' is a preset constant value, such as a = 3.
[0245] Then, count the positions in the non-encephalogenic signal EOG that are higher than the non-encephalogenic event threshold. The non-encephalogenic parts exceeding this threshold will be counted and divided into individual events, obtaining the number N of non-encephalogenic events occurring in the non-encephalogenic signal.
[0246] Finally, obtain the non-encephalogenic event occurrence frequency D based on the ratio of the number N of non-encephalogenic events to the preset duration T. For example: In a signal with a preset duration of 4 s, there are 4 non-encephalogenic events exceeding the threshold, so D = 4 / 4 = 1 (times / s).
[0247] After obtaining the non-encephalogenic control parameter AR and the non-encephalogenic event occurrence frequency D, one or both of them can be used in combination to determine different frequency band timbre combinations, forming a corresponding configuration parameter - timbre mapping table.
[0248] Taking the example of using only the AR parameter to determine the timbre used for each frequency band (each note belonging to this frequency band uses this timbre):
[0249] First, preset the parameters c1, c2 (c2 > c1) and rangeConfig used for the mapping. Here, c1 is the relaxation threshold, c2 is the excitation threshold, and rangeConfig is the transition parameter (which can be preset to 0.25). Among them, c1 and c2 can be manually adjusted parameters, and their preset values can be c = 1 and c2 = 2.
[0250] Then, determine the timbre used for each frequency band corresponding to the AR parameter based on the following mapping rules:
[0251] Calculate the transition zone width midrange using the following formula:
[0252] midrange = c1 - c2;
[0253] When AR < c1 - rangeConfig * midrange, use timbre combination 1.
[0254] When c1 - rangeConfig * midrange ≤ AR < c1 + rangeConfig * midrange, use timbre combination 2.
[0255] When c1 + rangeConfig * midrange ≤ AR < c2 - rangeConfig * midrange, use timbre combination 3.
[0256] When c2 - rangeConfig * midrange ≤ AR < c2 + rangeConfig * midrange, use timbre combination 4.
[0257] When c2+rangeConfig*midrange≤AR, use tone combination 5.
[0258] In a specific example, the timbre combinations are shown in Table 1 below (this is just an example and can be adjusted according to actual needs):
[0259] Table 1. Correspondence between timbre combinations of each frequency band
[0260]
[0261] As shown in Table 1, timbre combinations 1 to 5 correspond to five emotions ranging from soothing to intense, and are correlated with the intensity of non-brain-derived signal changes under different emotional states. A similar correspondence can be achieved using combinations of D or AR with D. The specific correspondence method can be determined according to specific needs.
[0262] In step S460, the brainwave audio is generated by using a MIDI synthesizer or an audio synthesis library based on the musical score corresponding to each specified frequency band in the brainwave audio representation data.
[0263] The purpose of this step S460 is to render the sheet music into a playable audio file so that it can be played back by an audio player to provide feedback to the user.
[0264] For control-symbol synthesis paths:
[0265] This approach first uses EEG signals and combines them with a discriminative model to generate control signals. For example, based on the emotion discrimination score, it determines the control signals for generating audio: Arosal (abbreviated as aro) and Valence (abbreviated as val). Then, it uses the control signals and EEG features corresponding to the EEG signals to generate notes corresponding to the EEG signals, and uses instrument sampling to play based on the notes.
[0266] Figure 5 A flowchart of a controlled-symbol audio synthesis method according to an embodiment of the present disclosure is shown. For each specified frequency band of EEG signal in each electrode channel, an EEG signal of a preset duration is processed as follows: Figure 5 The following steps S510 to S580 are shown:
[0267] For a pre-set duration of EEG signal in a specified frequency band of a specified electrode channel, perform the following steps S510 to S550:
[0268] In step S510, the EEG signal of a preset duration in the specified frequency band is used as input, and a pre-trained user state discrimination model is used to generate control parameters, which include: arousal level and valence.
[0269] In a specific example, the user state discrimination model includes a multilayer perceptron emotion discrimination model. When generating control parameters based on this emotion discrimination model, the control parameters are: arousal (aro) and valence (val). Arousal (aro) represents the intensity of physiological or emotional activation, describing a continuous state from "calm" to "excitement," and can be categorized as high arousal, moderate arousal, and low arousal. Valence (val) represents the positive or negative tendency of emotion, describing the emotional polarity from "negative" to "positive," and can be categorized as positive, neutral, and negative. Arousal (aro) and valence (val) are the core dimensions of the emotion quantification model, mapped from EEG features to audio parameters to achieve a closed loop of "physiological signal → emotional expression → audio generation." They can be calculated using the following formula:
[0270] aro = aro base (MotionInst,MotionLong)+meanConf×C0+(1-consistency)×
[0271] variation×C1;
[0272]
[0273] val = val base (MotionInst,MotionLong,meanConf)+meanConf×C2;
[0274] Among them, aro base This is a function of the real-time emotion discrimination result `MotionInst` and the long-term emotion discrimination result `MotionLong`. In a specific example, `MotionLong` is the classification result of the user state discrimination model when a 10-second EEG signal is input, and `MotionInst` is the result of the last 2-second EEG segment in the 10-second EEG signal input. If the preset duration is 4 seconds, then `MotionInst` is the result of the last 2-second EEG segment in the 4-second EEG signal input, and `MotionLong` is the classification result of the 4-second EEG signal combined with the previous 6-second EEG signal input. `meanConf`, `consistency`, and `variation` are the statistical results output by the emotion discrimination model, representing the average confidence, consistency, and average rate of change of the real-time emotion discrimination results over a preset number of times within a preset time period, respectively. base Let C0, C1, and C2 be functions of the real-time sentiment assessment result, the long-term sentiment assessment result, and the average confidence score, where C0, C1, and C2 are constants.
[0275] When specifically executing steps S520 to S550, the probability of note appearance, note duration, pitch, intensity, and timbre generation methods can be set according to the needs of the stimulation task. This flexibility serves objective monitoring or active intervention, meeting diverse needs such as medical treatment and entertainment. For example, if the current stimulation task (e.g., neurofeedback training task or state monitoring task) requires precise mapping of EEG characteristics to audio parameters and accurate representation of the user's neural activity state, for EEG signals expressing high arousal and active states (aro=8, val=8), fast-paced (short notes), dense, loud, high-pitched, and major-key bright brainwave audio is generated; while for EEG signals expressing low arousal and passive states (aro=2, val=3), slow-paced (long notes), sparse, soft, low-pitched, and minor-key melancholic melody is generated; if the current stimulation task ( For example, in clinical intervention tasks, it is necessary to reverse the regulation of the neural activity of the user being monitored. When guiding the user's neural state to adjust in the target direction (such as inhibiting over-excitation or activating the depressive state) through reverse audio feedback, for EEG signals expressing over-excitation (aro=9, val=6), a slow rhythm (long note), sparse, soft, low tone, and gentle melody in minor key is generated to induce relaxation. For EEG signals expressing the depressive state (aro=1, val=2), a fast rhythm (short note), dense, loud, high tone, and bright melody in major key is generated to stimulate emotions.
[0276] In step S520, the duration of a single note in the note sequence corresponding to the EEG signal of the preset duration is calculated based on the arousal level and / or the valence.
[0277] In a specific example, the note duration is obtained by calculating the rhythmic feature tempo, where:
[0278] tempo:note dur (t)=F tempo (aro,val);
[0279] Among them, note dur (t) represents the duration of the note at time t, which determines the rhythmic characteristics of the audio and defines the note. dur After (t), the length of all notes in the segment generated at time t can be note. dur (t), F tempo It is a function that takes `aro` and `val` as parameters. For example:
[0280] note dur (t) = [1 - (aro + val) × 0.4] norm ;
[0281] in,[] normAn operation is defined, for example, with 2 seconds as the duration of a whole note, [a] norm This means changing 'a' to a multiple of the nearest 0.25 (eighth note).
[0282] In step S530, the EEG signal of the preset duration is divided according to the duration of the note to obtain a set of EEG segments. The set of EEG segments contains multiple EEG segments, and each EEG segment corresponds to a note.
[0283] In one specific example, each note has the same duration. Thus, when the EEG signal of the preset duration is segmented according to the note duration, a series of EEG segments of equal duration will be obtained.
[0284] In step S540, the probability of occurrence of the note corresponding to each EEG segment in the EEG segment set is calculated based on the arousal level and / or the valence. When the occurrence probability meets the occurrence condition, the note corresponding to the corresponding EEG segment appears; otherwise, the note corresponding to the corresponding EEG segment does not appear. The note sequence corresponding to the EEG signal of the preset duration is determined based on the calculation result.
[0285] In a specific example, the probability of a note's occurrence is obtained by calculating the melody feature rhythm, where:
[0286] rhythm: p(note=1)=F rhythm (aro,val);
[0287] This parameter determines the complexity of the generated audio. p(note=1) represents the probability of a note appearing within a fixed-length audio segment. The higher the probability, the higher the frequency of the note and the more complex the melody.
[0288] For example, the i-th note of length Δt at time t. i Whether it appears will be determined by the following formula:
[0289]
[0290] Where, roughness is a preset constant or a function of aro and val, for example:
[0291] roughness = (aro + val) / 2;
[0292] activate is a calculation result containing a random number, for example:
[0293] activate = random × aro;
[0294] Where random is a random number uniformly distributed in [0,1].
[0295] In step S550, the intensity of the notes in the note sequence is calculated based on the arousal level, the valence, and the power spectrum energy characteristics of the EEG segments corresponding to the notes in the note sequence; the tonality parameters corresponding to the notes in the note sequence are calculated based on the arousal level and / or the valence; the pitch of the notes in the note sequence is calculated based on the tonality parameters and the power spectrum frequency characteristics of the EEG segments corresponding to the notes in the note sequence; the timbre of the notes in the note sequence is configured based on the arousal level and / or the valence; or, the timbre of all notes in the note sequence is configured based on a preset configuration parameter-timbre mapping table; the configuration parameter-timbre mapping table is used to describe the correspondence between specified timbre configuration parameters and timbres.
[0296] The specified timbre configuration parameters include: non-brain-derived control parameters and / or non-brain-derived event occurrence frequency. The non-brain-derived control parameters are obtained as follows: when the EEG signal of the specified electrode channel of the preset length contains non-brain-derived signals and denoised EEG signals, the ratio of the absolute values of the non-brain-derived signals and the denoised EEG signals in the EEG signal of the specified electrode channel of the preset length to the preset fractional part is used as the non-brain-derived control parameters.
[0297] The non-brain-derived event frequency is obtained as follows: a non-brain-derived event threshold is obtained by taking the absolute value of the denoised EEG signal and dividing it by a preset fractional number; positions in the non-brain-derived signal that are higher than the non-brain-derived event threshold are counted to obtain the number of non-brain-derived events occurring in the non-brain-derived signal; and the non-brain-derived event frequency is obtained based on the number of non-brain-derived events and the preset duration.
[0298] The intensity of a note can be calculated using the following formula:
[0299] loudness:note vel (t)=F loudness (aro,val)*P eeg (t);
[0300] Among them, note vel (t) represents the intensity of the note at time t, F loudness P is a function that takes aro and val as parameters. eeg (t) is the energy characteristic of the EEG power spectrum of the EEG segment extracted at time t.
[0301] The pitch of a note can be calculated using the following formula:
[0302] pitch:note = F pitch (f eeg(t))+mode(i-1);
[0303] This parameter determines the pitch of the note, i.e., the note's number in MIDI (Musical Instrument Digital Interface) representation; F pitch Therefore, f eeg (t) is a function with parameters, representing the mapping of EEG frequency features to musical note pitch features, f eeg (t) represents the frequency characteristics of the electroencephalogram (EEG) at time t, for example: energy-weighted average frequency f. weight mode represents the key, and i represents the note's position in the note sequence. Its calculation method is as follows:
[0304] mode=F mode (aro,val);
[0305] For example:
[0306]
[0307] The generation of the first note starts from mode[0], and the MIDI number of the i-th note differs from that of the previous note by mode[i-1].
[0308] When configuring the timbre of notes in the note sequence based on the arousal and / or valence, a mapping relationship between arousal and / or valence and timbre can be preset. Then, the corresponding timbre to be configured can be obtained based on the arousal and / or valence and the mapping relationship. The timbre can be set based on the characteristics of a specific instrument. For example, the timbre of notes in the note sequence can be configured using the mapping relationship table between arousal and valence and timbre shown in Table 2 below.
[0309] Table 2. Mapping Relationship Between Arousal and Valence and Timbre
[0310] Emotion Quadrant wake-up rate valence Instrument Selection High arousal 7-10 7-10 electric guitar High arousal neutral 7-10 4-7 small High arousal negative 7-10 0-4 Electric bass distortion Awakening Positive 4-7 7-10 piano Neutral Awakening 4-7 4-7 harp Awakening the negative 4-7 0-4 Big Drum Low arousal positive 0-4 7-10 Celesto Low arousal neutral 0-4 4-7 cello Low arousal negative 0-4 0-4 bass flute
[0311] Based on Table 2 above, when the arousal level is 8 and the valence is 4, the corresponding timbre configuration is trumpet.
[0312] When configuring the timbre of all notes in the note sequence according to the preset configuration parameters - timbre mapping table, the above description and combination with Table 1 can be used, and will not be repeated here.
[0313] Once the pitch, intensity, and timbre of the notes in the note sequence are determined, the notes corresponding to each EEG segment in the EEG segment set are determined.
[0314] Repeat steps S510 to S550 to obtain the notes corresponding to each EEG segment in the set of EEG segments of preset duration in other specified frequency band EEG signals of each specified electrode channel.
[0315] In step S560, a musical score corresponding to the specified frequency band is generated based on the note sequence of the specified frequency band in each electrode channel, the intensity, pitch, and timbre of the notes in the note sequence.
[0316] Repeat steps S510 to S560 to obtain the musical scores corresponding to other specified frequency bands of each specified electrode channel.
[0317] In step S570, the musical scores corresponding to each specified frequency band in each electrode channel are written into a complete musical score set corresponding to the preset duration, and the complete musical score set is used as the brainwave audio representation data.
[0318] In step S580, the brainwave audio is generated by a MIDI synthesizer or an audio synthesis library based on the musical score corresponding to each specified frequency band in the brainwave audio representation data.
[0319] In both the control-symbol synthesis path and the mapping-symbol synthesis path, this disclosure calculates the pitch of a note corresponding to an EEG segment based on the power spectrum frequency characteristics of the corresponding EEG segment, and calculates the intensity of a note based on the power spectrum energy characteristics of the corresponding EEG segment. On the one hand, using the power spectrum frequency characteristics of the EEG signal, rather than the time-domain amplitude characteristics of the EEG waveform, to determine the pitch of a note has the following beneficial technical effects: Since the frequency components of the EEG signal (e.g., alpha waves 8-13Hz, beta waves 14-30Hz) are less affected by environmental noise and electrode contact impedance, while the amplitude is easily fluctuated by changes in scalp impedance, motion artifacts, etc., frequency analysis can effectively filter out interference from non-target frequency bands, improving the stability of pitch recognition; furthermore, specific frequency bands are directly related to cognitive states (e.g., gamma waves and auditory processing), and frequency-pitch mapping can more naturally match the user's musical intentions, while amplitude only reflects the intensity of neural activity and has no clear physiological correlation with pitch. Furthermore, while frequency characteristics can be precisely mapped to pitch through frequency bands (e.g., 1Hz steps), amplitude is limited by the signal's dynamic range, making it difficult to distinguish subtle pitch differences. On the other hand, using power spectrum energy characteristics (rather than average power) in EEG signals to determine note intensity can achieve more accurate and interference-resistant intensity mapping by differentiating energy distributions across different frequency bands. Specifically, this manifests in: ① frequency band specificity that allows intensity modulation to align with physiological intentions (e.g., enhanced beta waves corresponding to forte); ② expanded dynamic range supporting subtle dynamic changes (e.g., crescendo, vibrato); and ③ computational efficiency by reducing latency through selective band analysis. Compared to the single scalar output of average power, power spectrum energy characteristics significantly enhance expressiveness and robustness in multimodal music interaction.
[0320] Furthermore, this disclosure provides three different timbre mapping schemes for determining the timbre of notes corresponding to EEG segments: configuring the timbre of different frequency band note sequences based on a preset frequency band-timbre mapping table, a preset configuration parameter-timbre mapping table, and the arousal level and / or the valence. In specific applications, the corresponding timbre mapping scheme can be flexibly selected according to different application scenarios or needs.
[0321] For mapping-type waveform synthesis paths:
[0322] This path maps EEG features (such as frequency domain features and time domain envelope) corresponding to EEG signals to audio features, and uses EEG features (such as time domain envelope) to modulate the audio signal to obtain an audio waveform with a special EEG timbre.
[0323] Figure 6 A flowchart of a mapping-waveform audio synthesis method according to an embodiment of the present disclosure is shown. For each specified frequency band of EEG signal in each electrode channel, an EEG signal of a preset duration is processed as follows: Figure 6The following steps S610 to S650 are shown:
[0324] For a pre-set duration of EEG signal in a specified frequency band of a specified electrode channel, perform the following steps S610 to S630:
[0325] In step S610, the EEG signal of the preset duration is segmented based on the temporal envelope corresponding to the EEG signal of the preset duration to obtain a set of EEG segments.
[0326] The segmentation process includes the following steps S611 to S613:
[0327] In step S611, the effective amplitude threshold is determined based on the mean of the time domain envelope.
[0328] The effective amplitude threshold is used to filter low-amplitude noise. Setting the effective amplitude threshold based on the envelope mean can avoid the sensitivity of a fixed threshold to changes in signal amplitude.
[0329] When determining the effective amplitude threshold, it can be based on the product of a preset scaling factor and the mean of the time domain envelope. In a specific example, the scaling factor can be a value between 1.2 and 1.5.
[0330] In step S612, the set of valleys of the time domain envelope is determined according to the effective amplitude threshold, wherein the valley amplitudes in the set are not greater than the effective amplitude threshold.
[0331] Specifically, a sliding window (e.g., 100ms) is used to detect local minima of the temporal envelope Envelope(t), preserving values with amplitudes ≤ th. amp The valleys, after removing spurious valleys caused by high-frequency noise, are represented by the valley set V as follows:
[0332] V = {v0, v1, ... v} n |v i ≤th amp};
[0333] Among them, th amp The effective amplitude threshold is denoted as .
[0334] In step S613, the EEG signal of the preset duration is segmented according to the preset effective time threshold and the set of troughs of the time domain envelope to obtain the set of EEG segments. The set of EEG segments contains one or more EEG segments, and the time value of each EEG segment is greater than the effective time threshold.
[0335] Each EEG segment corresponds to an audio waveform; the effective time threshold is used to ensure the minimum duration of the note, which can prevent short-term noise from being misjudged as a note.
[0336] The set of troughs in the temporal envelope determines each possible EEG segment (seg). i , of which seg i The left boundary is v i The right boundary is v i+1 Obtain the set of EEG segments Seg = {seg i |v i+1 -v i >th time}, th time The effective time threshold is denoted as .
[0337] Figure 7 This diagram illustrates the segmentation results of an EEG segment in a mapping-waveform audio synthesis method according to an embodiment of the present disclosure. The blue portion represents the EEG signal, the red portion represents the temporal envelope of the EEG signal, and the black portion represents the EEG segmentation result.
[0338] In step S620, for any EEG segment in the set of EEG segments, the characteristic parameters of the waveform corresponding to any EEG segment are determined according to the frequency domain characteristics of the EEG segment, until the characteristic parameters of the waveform corresponding to each EEG segment in the set of EEG segments are obtained.
[0339] The step of determining the characteristic parameters of the waveform corresponding to any EEG segment based on the EEG frequency domain characteristics of any EEG segment includes: based on the power spectrum characteristics of any EEG segment, determining the fundamental frequency and harmonic frequency of the waveform corresponding to any EEG segment, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency, based on the power spectrum characteristics of the power spectrum characteristics and the number of peak frequencies and amplitudes ranked by amplitude of the top preset number of peak frequencies.
[0340] In a specific example, the harmonic frequencies include the second harmonic frequency and the third harmonic frequency. Then: obtain the top 3 peak values with the largest amplitude in the power spectrum characteristics corresponding to any EEG segment, and record the peak frequencies {f0, f1, f2} and their corresponding amplitudes {A0, A1, A2}. The largest peak frequency f0 and its corresponding amplitude A0 can be used as the fundamental frequency and its intensity, f1 and its corresponding amplitude A1 can be used as the second harmonic frequency and its intensity, and f2 and its corresponding amplitude A2 can be used as the third harmonic frequency and its intensity.
[0341] In step S630, for any EEG segment in the set of EEG segments, the audio waveform corresponding to any EEG segment is determined according to the characteristic parameters and time-domain envelope of the waveform corresponding to the any EEG segment, until the audio waveforms corresponding to each EEG segment in the set of EEG segments are obtained.
[0342] Determining the audio waveform corresponding to any EEG segment based on the characteristic parameters and time-domain envelope of the waveform corresponding to any EEG segment includes the following steps S631-S632:
[0343] In step S631, a sine wave of the corresponding frequency is generated and mixed according to the fundamental frequency and harmonic frequency of the waveform corresponding to any EEG segment, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency, to obtain a mixed sine wave.
[0344] Specifically, the fundamental frequency sine wave S0(t) is generated using the following formula:
[0345] S0(t) = A0·sin(2πf0t);
[0346] The second harmonic sine wave S1(t) is generated using the following formula:
[0347] S1(t) = A1·sin(2πf1t);
[0348] The third harmonic sine wave S2(t) is generated using the following formula:
[0349] S2(t) = A2·sin(2πf2t);
[0350] Generate a mixed sine wave s using the following formula. mix (t):
[0351] s mix (t)=S0(t)+S1(t)+S2(t);
[0352] In step S632, the mixed sine wave is modulated using the time-domain envelope corresponding to any one of the EEG segments to obtain the audio waveform corresponding to the any one of the EEG segments.
[0353] Specifically, the audio waveform a corresponding to any of the EEG segments can be obtained using the following formula. note :
[0354] a note (t)=s mix (t)*e nite (t);
[0355] Among them, e note (t) represents the temporal envelope corresponding to any of the EEG segments.
[0356] Repeat steps S610 to S630 to obtain the audio waveforms corresponding to each EEG segment in the set of EEG segments of preset duration in other specified frequency band EEG signals of each specified electrode channel.
[0357] In step S640, the audio waveforms corresponding to each EEG segment in the set of EEG segments of each specified frequency band in the specified electrode channel are used as the audio waveforms corresponding to the EEG signals of the specified electrode channel with the preset duration.
[0358] In step S650, the audio waveforms corresponding to the EEG signals of each electrode channel for the preset duration are used as the brainwave audio representation data; the brainwave audio representation data are directly used as brainwave audio.
[0359] In the mapping-waveform synthesis path, the temporal envelope is used to determine the rhythm of the audio, and the spectral characteristics of the EEG segments are used to determine the fundamental tone, harmonics, and intensity of each component of the audio, and a sine wave of the corresponding frequency is synthesized. The temporal envelope is then applied to the generated harmonic sine wave to generate audio with "EEG color".
[0360] The real-time brainwave audio synthesis feedback loop implemented in this disclosure realizes the path for converting real-time EEG signals into personalized audio stimulation signals. It supports three different audio synthesis paths, can use a variety of instrument timbres or brainwave envelope timbres, and supports multi-band, multi-channel audio synthesis and playback. It has significant application value in personalized brainwave music therapy, real-time brainwave audio feedback stimulation, and other fields.
[0361] Figure 8 A flowchart illustrating another proprioceptive brainwave audio-visual perception synchronous feedback method according to an embodiment of this disclosure is shown. Figure 8 As shown, after performing step S150, the following step S160 is also included:
[0362] In step S160, the brainwave audio is synchronously fed back to the user being collected, so as to dynamically regulate the nervous system activity of the user being collected through auditory perception.
[0363] By synchronously playing the generated brainwave audio to the user, real-time audio stimulation can be achieved. Simultaneously, the synchronously acquired EEG signals following stimulation will continue the audio synthesis process described above, controlling the synthesis of subsequent audio stimulation signals, thus achieving a closed loop from EEG acquisition to audio stimulation. In this closed loop, the default frame length for brainwave audio generation is 4 seconds, and the brainwave audio feedback delay is less than 5 seconds. Furthermore, users can interact with the front-end visual interface to determine the frequency band used for generating brainwave audio, thereby enabling various combinations of audio frequencies corresponding to different brainwave bands.
[0364] Furthermore, with the rapid development of Virtual Reality (VR) technology, users' demand for personalized and interactive content in immersive experiences is increasing. In VR application scenarios, the real-time EEG acquisition and brainwave audio generation scheme provided in this disclosure can be used to provide users with personalized background audio based on individual EEG characteristics, thereby significantly improving the immersion and engagement in VR interactive scenarios.
[0365] For this specific application scenario, a portable EEG acquisition device is used to collect real-time EEG signals while the VR device is being worn. These signals are then processed simultaneously to generate personalized brainwave audio, which is provided to the user as background audio. This background audio fully utilizes the user's EEG characteristics to generate brainwave audio with individual user-specific features. Furthermore, it can analyze the user's emotional state in real time based on changes in the user's EEG signals (determined by changes in the corresponding EEG features). The feature parameters of the generated brainwave audio are adjusted according to the user's emotional state and the scenario, achieving a closed-loop adjustment of the personalized background audio. For example, when the user's EEG characteristics are identified as matching emotions such as anxiety or tension, the generated feature parameters are adjusted to produce calming brainwave audio; when the scenario requires the user to concentrate, brainwave audio that promotes focus is generated.
[0366] Through the implementation of this embodiment, users can experience a highly personalized audio environment in VR scenarios. This audio not only blends perfectly with the virtual scene but also intelligently adjusts according to the user's real-time emotional state. This significantly enhances user immersion and satisfaction and shows broad application prospects in multiple fields such as mental health therapy, rehabilitation training, and entertainment experiences. In the future, with further technological development, this system is expected to achieve breakthrough applications in more fields, bringing users a richer and more personalized virtual reality experience.
[0367] According to embodiments of this disclosure, the step of synchronously feeding back the brainwave audio to the user being collected includes:
[0368] Based on the combination relationship between the spatial positions of electrodes in the one or more electrode channels and the one or more specified frequency bands, brainwave audio corresponding to multiple channels is obtained based on the brainwave audio.
[0369] In a specific example, for two or more vocal channels, spatial regions can be defined based on electrode coordinates in one or more electrode channels (e.g., Fp1 corresponds to the left frontal lobe, Cz corresponds to the central parietal lobe, and O1 corresponds to the left occipital lobe). The scalp can be divided into regions such as the frontal lobe, central lobe, occipital lobe, and temporal lobe. The brainwave audio of each region corresponds to a specific vocal channel group. For example, the brainwave audio of the Theta frequency band in the frontal region (Fp1, Fp2) can be mapped to the left anterior / right anterior vocal channel, the brainwave audio of the Beta frequency band in the central region (C3, C4) can be mapped to the central vocal channel, and the brainwave audio of the Alpha frequency band in the occipital lobe region (O1, O2) can be mapped to the left posterior / right posterior vocal channel.
[0370] In another specific example, for the left and right channels, the left and right brain regions and the central region can be defined according to the electrode coordinates in one or more electrode channels. For example, the left brain region corresponds to F3 (left frontal lobe), C3 (left central lobe), T3 (left temporal lobe), P3 (left parietal lobe), and O1 (left occipital lobe), the right brain region corresponds to F4 (right frontal lobe), C4 (right central lobe), T4 (right temporal lobe), P4 (right parietal lobe), and O2 (right occipital lobe), and the central region corresponds to Cz (central parietal region). Then, the brainwave audio of the corresponding frequency band of the left brain region is mapped to the left channel, and the brainwave audio of the corresponding frequency band of the right brain region is mapped to the right channel. The brainwave audio of the central region is balanced and distributed to the left and right channels according to a preset ratio (e.g., 50:50).
[0371] Next, the brainwave audio from the multiple channels is played through an audio device, for example, the left front channel is played through the left front speaker, and the right rear channel is played through the right rear speaker, so as to synchronously feed the brainwave audio back to the user being collected.
[0372] This disclosure generates and feeds back multiple channels of brainwave audio to the user, enabling the user to perceive the spatial distribution of brain activity through hearing (e.g., when the frontal Theta is enhanced, the volume of the left anterior channel becomes prominent), thus enhancing spatial perception. Through the multi-dimensional feedback of multi-channel brainwave audio, it provides richer information than traditional single-channel audio, thereby effectively assisting neurofeedback training.
[0373] In addition to generating and feeding back multiple channels of brainwave audio to the user to form auditory feedback, this disclosure also provides a scheme to form visual feedback by dynamically displaying the location of sound images.
[0374] Based on the principle of binaural hearing, sound sources played from different locations reach both ears at different times and with varying intensities. The nervous system uses these differences to determine the direction of the sound, giving it a directional quality. Using a playback device with independent dual channels, directional sound—i.e., stereo—can be simulated by adjusting the intensity and phase of the left and right channel audio. Alternatively, a speaker array with more than two channels can be used to synchronously play brainwave audio from multiple channels and spatial locations, creating a richer spatial surround sound effect. Directional sound effectively attracts the listener's attention, and changes in sound image position and intensity can be used to regulate the listener's attentional state.
[0375] In this disclosure, the acoustic image position of real-time brainwave audio is calculated, and a dynamic acoustic image display interface is implemented, which enables dynamic visualization of the changes in the acoustic image position of the brainwave audio while it is being played.
[0376] According to embodiments of this disclosure, the method further includes:
[0377] Obtain the acoustic image position parameters of the brainwave audio corresponding to multiple channels, and dynamically display the acoustic image position of the brainwave audio on the visualization interface based on the acoustic image position parameters.
[0378] According to embodiments of this disclosure, obtaining the acoustic image location parameters of brainwave audio corresponding to multiple channels includes:
[0379] Obtain the brainwave audio corresponding to each specified frequency band in the brainwave audio of the multiple channels, and calculate the intensity of the brainwave audio corresponding to each specified frequency band in each channel within a specified time period; obtain the acoustic image orientation indication parameter corresponding to the brainwave audio of the specified frequency band based on the intensity difference of the brainwave audio corresponding to the specified frequency band in different channels; and use the acoustic image orientation indication parameter and intensity of the brainwave audio corresponding to each specified frequency band in the multiple channels as the acoustic image position parameter.
[0380] According to embodiments of this disclosure, the method further includes:
[0381] The sound image location of the brainwave audio is dynamically and synchronously fed back to the user being collected on a visual interface, so as to dynamically regulate the neural activity of the user being collected through visual perception.
[0382] In a specific example, taking left and right stereo channels as an example, for brainwave audio signals a∈A corresponding to different frequency bands, the waveforms of a are monitored and independent sound image positions are calculated. The audio analysis tools in Tone.js can be used to monitor the brainwave audio waveforms played in the left and right channels for the specified frequency bands in real time, calculate the intensity of the brainwave audio corresponding to the specified frequency band in the left and right channels within a specified duration, obtain the intensity difference of the brainwave audio corresponding to the specified frequency band in the left and right channels, and then derive the sound image orientation indicator parameters based on this intensity difference.
[0383] Specifically, the acoustic orientation indicator parameter c can be calculated using the following formula:
[0384]
[0385] Among them, v right This represents the intensity of brainwave audio in the right channel of a specified frequency band, v left This indicates the intensity of the brainwave audio in the left channel for a specified frequency band, v right -v left It has poor strength.
[0386] In a specific example, c∈[-1,1], c=-1 indicates that the sound image is on the far left, c=1 indicates that the sound image is on the far right, and c=0 indicates that the sound image is in front (middle position).
[0387] Figure 9 A schematic diagram showing the sound intensity location of brainwave audio in a specified frequency band according to an embodiment of the present disclosure is provided. Figure 9 As shown, when dynamically displaying the sound intensity position in real time, the front-end visual dynamic feedback interface uses multiple small balls at different locations to display the sound image position. The small balls are arranged in multiple rows, with each row representing the brainwave audio corresponding to a frequency band. The size and color of the small balls at different positions are determined based on the real-time calculated sound image position parameters and intensity parameters, which are used to represent the sound image position. The sound image orientation indicator parameter can be normalized, and the normalized sound image orientation indicator parameter is used to represent the offset of the sound image position relative to the center point. The intensity is used to represent the size of the small ball, thereby realizing the real-time calculation and dynamic visualization of the brainwave audio sound image position.
[0388] The refresh rate of the front-end feedback animation can be set to [5Hz, 10Hz]. The audio-visual display and audio playback feedback occur simultaneously with a latency of less than 200ms. It supports selecting different frequency bands for visualizing the location of the audio-visual image. There are four selectable frequency bands: δ & θ, α, β, and γ, which correspond to the specified frequency bands used in the generation of brainwave audio.
[0389] This embodiment provides users with an interactive feedback interface for intuitive visual observation of changes in brainwave audio-visual images. This is beneficial for testers to observe changes in brainwaves at different frequency bands and the movement of brainwave audio-visual images in the user being tested. The dynamic visual feedback of the image position also has the potential to be used in the user's attention regulation scenarios. It has the potential to be combined with brainwave audio for stimulation regulation, promoting psychological relaxation and cognitive regulation. It can be extended to multiple fields such as medicine, entertainment, and education.
[0390] In this disclosure, besides extracting EEG features corresponding to one or more specified frequency band EEG signals to directly generate corresponding EEG audio without considering the inter-electrode dependencies of the EEG acquisition device, other relevant features, such as brain network features and source space features, are extracted in real time, taking into account the relationships between multiple EEG electrode leads. These features are then visualized and fed back to doctors or researchers, helping them to assess the patient's current state in real time and adjust treatment plans flexibly and individually. Addressing the difficulties in real-time feature calculation and the single feedback pathway encountered in traditional EEG acquisition-feedback processes, this disclosure improves the methods for extracting brain network features and source space features, proposing a real-time automatic EEG processing and feedback workflow that achieves a low-latency real-time link for EEG signal acquisition, analysis, and feature visualization feedback.
[0391] According to embodiments of this disclosure, the method further includes:
[0392] Brain network features corresponding to the EEG signals of the multiple electrode channels are extracted. The brain network features include: node-node functional connectivity features and / or edge-edge interaction features. The node-node functional connectivity features are used to describe the statistical dependency between two node signals, and the edge-edge interaction features are used to describe the temporal dependency between two edges. In this case, a node corresponds to a single electrode channel, a node signal corresponds to the EEG signal of a single electrode channel, and an edge corresponds to the connectivity between two nodes.
[0393] Specifically, node-to-node functional connectivity features corresponding to the EEG signals of the multiple electrode channels are extracted by calculating the weighted phase lag index (WPLI) between every two node signals. WPLI is used to quantify the phase lag relationship between two signals. By weighting the sign of the imaginary part of the cross spectrum, the influence of zero-lag spurious correlation is reduced. Compared with PLI (Weighted Phase Lag Index), which only calculates the expectation of the imaginary part sign and ignores amplitude information, WPLI can more accurately quantify the statistical significance of phase lag, while reducing the influence of noise or zero-lag spurious correlation and improving the sensitivity to the true phase lag. In addition, the EEG signal segments required for WPLI calculation are relatively short, which can be used in real-time functional connectivity calculation scenarios.
[0394] According to embodiments of this disclosure, calculating the weighted phase lag coefficient between any two node signals includes: performing a Hilbert transform on a first node signal and a second node signal to obtain a first node analytic signal and a second node analytic signal corresponding to the first node signal and the second node signal, respectively; obtaining the instantaneous cross power spectrum of the first node signal and the second node signal based on the first node analytic signal and the second node analytic signal; averaging the imaginary absolute values of the instantaneous cross power spectrum at each time point within a preset time window to obtain the mean of the imaginary absolute values; calculating the weight of the imaginary part of the instantaneous cross power spectrum at each time point within the preset time window, and obtaining a weighted complex covariance based on the instantaneous cross power spectrum at each time point within the preset time window and the weight of the imaginary part corresponding to the instantaneous cross power spectrum; and calculating the weighted phase lag coefficient of the first node signal and the second node signal using the imaginary absolute value of the weighted complex covariance and the mean of the imaginary absolute values.
[0395] The first node parse signal and the second node parse signal, corresponding to the first node signal and the second node signal, are obtained by the following formulas:
[0396] z I (t)=x i (t)+j·H(x I (t));
[0397] z j (t)=x j (t)+j·H(x j (t));
[0398] The instantaneous cross power spectrum is obtained using the following formula:
[0399]
[0400] The weight of the imaginary part of the instantaneous cross-power spectrum at each time point within the preset time window is calculated using the following formula:
[0401]
[0402] c = median(|Im(P ij (t))|);
[0403] The weighted complex covariance is obtained using the following formula:
[0404]
[0405] The weighted phase lag coefficient is calculated using the following formula:
[0406]
[0407] Where, x i (t) represents the signal at the first node, x j (t) represents the signal at the second node, z i (t) is the analytic signal of the first node, z j (t) is the analytic signal of the second node, j is the imaginary unit, and H() is the Hilbert transform operator; P ij (t) represents the instantaneous cross-power spectrum, * represents complex conjugate, and median() performs median operation. The weighted complex covariance is T, where T is the number of time points within the preset time window.
[0408] Traditional methods for calculating weighted phase lag coefficients, while reducing the impact of zero-lag noise, can still be affected by high-frequency or non-stationary noise, exhibiting sensitivity to outliers (such as sudden noise). In this disclosure, a dynamic weighting mechanism is introduced when calculating the weighted phase lag coefficient between any two node signals. Firstly, weights are calculated using the absolute value of the imaginary part, assigning higher weights to time points with high signal-to-noise ratios (SNR) and suppressing the influence of low SNR points, making phase lag estimation more robust in EEG noise environments (such as EEG artifacts and EMG interference). Secondly, weighting using complex covariance preserves the joint information of phase difference amplitude and direction, improving the resolution of phase synchronization features. Furthermore, by suppressing noisy connections through weighting, the correlation between small-world properties of brain networks (such as clustering coefficients and global efficiency) and behavior / disease is significantly enhanced, improving the reliability of graph theory analysis results.
[0409] In extracting edge-to-edge interaction features corresponding to the EEG signals of the multiple electrode channels, this disclosure employs instantaneous phase coherence as a measure of node-to-node connectivity (edges) to ensure sufficient edge time series length within each window for measuring correlation. Then, the Pearson correlation coefficient between the two edge time series is calculated. The Pearson correlation coefficient reflects the orthogonality of signals between electrode channels; a larger value indicates a weaker connection and more independent information between the two electrode channels. Furthermore, the EEG signal segments required for calculating the Pearson correlation coefficient are relatively short, making it suitable for real-time functional connectivity calculation scenarios.
[0410] According to embodiments of this disclosure, edge-to-edge interaction features corresponding to the EEG signals of the multiple electrode channels are extracted by calculating the Pearson correlation coefficient between every two edges in the candidate edge set. The candidate edge set is constructed by calculating the average phase synchronization index of all edges, sorting the average phase synchronization indices of all edges from largest to smallest, and retaining the first preset number of edges. The average phase synchronization index of the edges is used to indicate the phase synchronization strength of the edges. When the average phase synchronization index of an edge is 0 or approximately 0, it indicates that the phase difference between the two nodes is randomly distributed and there is no synchronization. When the average phase synchronization index of an edge is 1 or approximately 1, it indicates that the phase difference between the two nodes is constant and they are completely synchronized. By pre-screening all edges based on the average phase synchronization index, only highly synchronized functional connection edges are retained. Then, the instantaneous linear correlation of highly synchronized edges within a local window is analyzed within a sliding window. By calculating the sparse Pearson correlation coefficient between highly synchronized edges to capture dynamic interactions and detect rapidly changing connections, fully connected computation is avoided, thus significantly reducing the computational load. Therefore, by combining pre-screening and sparse Pearson computation, computational efficiency can be significantly improved while maintaining accuracy.
[0411] According to an embodiment of this disclosure, when calculating the Pearson correlation coefficient between any two edges in the candidate edge set, the method includes: calculating the instantaneous phase difference between the first edge and the second edge at each time point within a preset time window, wherein the instantaneous phase difference is the difference between the instantaneous phases of one node and another node in the edge; and calculating the sparse Pearson correlation coefficient between the first edge and the second edge using a sliding window based on the instantaneous phase difference between the first edge and the second edge at each time point within the preset time window.
[0412] The average phase synchronization index of the edge is calculated using the following formula:
[0413]
[0414] The instantaneous phase of a node is calculated using the following formula:
[0415] z(t) = x(t) + j·H(x(t));
[0416]
[0417] The instantaneous phase difference of the first side and the instantaneous phase difference of the second side can be obtained using the following formulas:
[0418]
[0419]
[0420] The Pearson correlation coefficient between the first side and the second side is calculated using the following formula:
[0421]
[0422] Where x(t) is the node signal, z(t) is the node analytic signal, j is the imaginary unit, and H() is the Hilbert transform operator; This represents the instantaneous phase of the node at time point t, and arg() represents taking the phase angle. This represents the instantaneous phase of a node in the first edge at time point t. This represents the instantaneous phase of the other node in the first edge at time point t. This represents the instantaneous phase of a node in the second edge at time point t. This represents the instantaneous phase of the other node in the second edge at time point t. This represents the instantaneous phase difference of the first edge at time point t. This represents the instantaneous phase difference of the second side at time point t, where T is the number of time points within the preset time window, and W is the length of the sliding window. <T,μ ij Let μ be the average instantaneous phase difference of the first side within the sliding window. uv The average instantaneous phase difference of the second side within the sliding window.
[0423] According to embodiments of this disclosure, the method further includes:
[0424] Electrodes corresponding to the multiple electrode channels are dynamically drawn on the individualized 3D head model of the visualization interface. The electrodes are distributed on the scalp surface of the individualized 3D head model. The connection relationship and connection strength between the electrodes and / or between the electrode connections are dynamically displayed in real time according to the brain network features, the electrode display parameters input by the user, and the feature display parameters are used to indicate the display of one or more electrodes corresponding to the multiple electrode channels. The feature display parameters are used to indicate the display of node-node functional connection features and / or edge-edge interaction features in the brain network features.
[0425] Specifically, lines can be used to represent the connections between different electrodes, with line thickness and color representing connection strength. Electrodes can also be selected for display; unselected electrodes will not have their connections shown. The refresh rate and latency of the functional connection feedback are the same as the parameters for displaying the EEG waveform.
[0426] EEG source localization refers to the technique of calculating the brain's internal neural electrical activity using scalp EEG signals and head electromagnetic models. This technique maps signals from the scalp's electrode space to the cerebral cortex (source space), thus visually reflecting the distribution of neural electrical activity within the brain. In this disclosure, individual head conductivity models were calculated using individual T1-weighted MRI data for different individuals. These individual head conductivity models were then used as parameters to achieve real-time EEG source localization using the Minimum Norm Estimation (MNE) method, thereby realizing real-time mapping of EEG features from the electrode space to the source space.
[0427] Specifically, according to embodiments of this disclosure, the method further includes: extracting source space features corresponding to the EEG signals of the plurality of electrode channels, including: obtaining a pre-calculated lead field pseudo-inverse matrix, and obtaining the source space features based on the pre-calculated lead field pseudo-inverse matrix and the EEG signals of the plurality of electrode channels; wherein, the lead field pseudo-inverse matrix is determined based on the lead field matrix and a regularization parameter optimized offline, the lead field matrix is used to describe the electrical signal transmission relationship from the source space to the electrode space, the lead field matrix is calculated using the T1-weighted magnetic resonance imaging data of the user being collected as input, and using a pre-constructed individualized three-dimensional head model corresponding to the user being collected; the regularization parameter optimized offline is the regularization parameter corresponding to the maximum curvature point of the relationship curve of the residual norm and the solution norm.
[0428] The source space features are calculated using the following formula:
[0429]
[0430] Alternatively, the source space features can be calculated using the following formula:
[0431]
[0432] L + =L T (LL T +λI) -1 ;
[0433] R = L T (LL T +λI)-1 L;
[0434] in, Let L be the source space feature at time t, x(t) be the EEG signal of the multiple electrode channels corresponding to time t, and L be the source space feature at time t. + Let L be the pseudo-inverse matrix of the lead field, and L be the lead field matrix, where L∈R C×D C represents the number of scalp electrodes, and D represents the source space dimension. T This is the transpose operation of a matrix, where I is the identity matrix, λ is the regularization parameter after offline optimization, and diag() represents the extraction of diagonal elements of the matrix.
[0435] In this disclosure, when extracting source spatial features corresponding to the EEG signals of the multiple electrode channels, on the one hand, the complexity of source spatial feature calculation is reduced and computational efficiency is improved by pre-calculating the lead field pseudo-inverse matrix, ensuring the real-time performance of source spatial feature extraction; on the other hand, when calculating the lead field pseudo-inverse matrix, the lead field matrix used is calculated based on the individualized three-dimensional head model corresponding to the user being sampled, thereby improving the source localization accuracy; in addition, when calculating the lead field pseudo-inverse matrix, a regularization parameter optimized offline is used. The offline optimization of the regularization parameter is determined by plotting the relationship curve between the residual norm and the solution norm, and selecting the regularization parameter corresponding to the point of maximum curvature of the relationship curve. The residual norm is used to measure the difference between the electrode signal predicted by the estimated source activity through the forward model and the actual observed signal. The smaller the value, the closer the solution of the inverse problem is to the observed data. The solution norm is used to quantify the overall amplitude (energy) of the source activity estimation. The larger the value, the more unstable or sparse the solution is. By balancing these two parameters, overfitting noise or excessively smoothed solutions can be avoided, thereby improving the reliability of source localization. Specifically, the point of maximum curvature corresponds to the optimal trade-off between residuals and solution complexity, ensuring that source localization is both stable and accurate.
[0436] According to embodiments of this disclosure, the method further includes:
[0437] Based on the model display parameters input by the user, the source space activation map of the corresponding part of the individualized three-dimensional head model is dynamically drawn on the visualization interface. The individualized three-dimensional head model includes: the surface part of the brain, the surface part of the skull, and the surface part of the scalp. Based on the source space features, the source space electrical activity intensity is dynamically displayed in real time in the source space activation map of the corresponding part of the displayed individualized three-dimensional head model.
[0438] In a specific example, a web-based interface dynamically displays the source space electrical activity distribution of a five-channel EEG signal in real-time. A boundary element model (BEM) of the individual's head is calculated based on individual MRI data. The web-based interface displays the individual's head geometry model, including brain surface modeling, skull surface modeling, and scalp surface modeling, according to the user's selection. The user can choose whether to display each part of the geometry model on the web-based interface, adjust its transparency, and adjust the viewing angle by dragging the model. The individual's source space is distributed within the skull, and the intensity of source space electrical activity is dynamically displayed based on the real-time calculated source tracking results (i.e., source space features). The refresh rate and latency of the dynamic feedback of the source tracking results are the same as those of the EEG waveform display.
[0439] According to embodiments of this disclosure, the method further includes: acquiring dynamic adjustment parameters; wherein the dynamic adjustment parameters are generated based on stimulus task requirements, or based on brain network features corresponding to the EEG signals of the plurality of electrode channels; and adjusting feature parameters of the audio performance based on the dynamic adjustment parameters.
[0440] This disclosure provides two implementation modes for obtaining dynamic control parameters:
[0441] Mode 1: Manual Intervention Mode. During neuromodulation, experts observe the waveforms of the user's EEG signals, dynamically displayed brain network features, or source space features through a visual interface. They then combine this with prior knowledge to create a stimulation task tailored to the current EEG signal. Dynamic modulation parameters are generated based on the task's requirements, and these parameters are manually input to adjust the audio characteristics and alter the current brainwave audio. For example, if an expert observes insufficient energy in the user's prefrontal cortex theta waves (4-8Hz) and wishes to enhance attention by increasing low-frequency note feedback, the theta band gain can be adjusted from 1.0 to 1.5 to change the note intensity, lowering the pitch (making it more subdued) and increasing the volume. The note density can be adjusted from 5 to 8 to affect the note frequency, increasing the number of notes per unit time and creating a more dense rhythm. These adjustments alter the brainwave audio, thereby achieving a neuromodulation effect that enhances attention.
[0442] Mode 2: Automatic Analysis Mode. Dynamic control parameters can be generated based on the brain network characteristics (e.g., node-to-node functional connectivity features) corresponding to the EEG signals of the multiple electrode channels. That is, for an EEG signal corresponding to a time period T, when generating the corresponding EEG audio based on the EEG frequency domain features and / or EEG time domain features of the EEG signal in a specified frequency band, dynamic control parameters are generated based on the functional connectivity features between the electrode channel to which the EEG signal belongs and other electrode channels, thereby triggering the modulation of cross-channel audio feature parameters.
[0443] In a specific example, for an EEG signal corresponding to a time period T, the EEG signal is represented as EEG(t), t∈T. First, the node-to-node functional connectivity features between the electrode channel i containing the EEG signal and other electrode channels among the plurality of electrode channels are obtained. Then, the electrode channel j with the strongest connectivity relationship is selected, that is, the largest FC among the obtained node-to-node functional connectivity features is selected. ij To avoid noise interference, it is determined whether the largest value meets a preset condition (e.g., FC). ij If the value is greater than a preset threshold (th), then the node-to-node functional connection feature FC with the largest value is selected. ij The dynamic control parameter is used as the basis for adjusting the characteristic parameters of the audio performance; if the condition is not met, the modulation of the cross-channel audio characteristic parameters is not triggered.
[0444] When adjusting the characteristic parameters of the audio performance according to the dynamic control parameters, the characteristic parameters of the audio performance that can be adjusted include, but are not limited to, envelope and pitch.
[0445] The envelope of the audio representation generated based on the specified frequency band of the EEG signal can be adjusted using the following formula according to the dynamic control parameters:
[0446]
[0447] The pitch of the audio representation generated based on a specified frequency band of EEG signals can be adjusted using the following formula, according to the dynamic control parameters:
[0448]
[0449] Wherein, envelope is the envelope of the audio representation generated from the specified frequency band of the EEG signal based on electrode channel i after adjustment. i The envelope is the original envelope of the audio representation generated from the specified frequency band of the EEG signal based on electrode channel i. j FC is the envelope of the audio representation generated from the specified frequency band of the EEG signal based on electrode channel j. ij The node-to-node functional connectivity features between electrode channel i and electrode channel j are defined as follows: pitch is the pitch of the audio representation generated from the EEG signal of a specified frequency band based on electrode channel i. i The pitch is the original pitch of the audio representation generated from the specified frequency band of the EEG signal based on electrode channel i. jC1 represents the pitch of the audio representation generated from the EEG signal of the specified frequency band based on electrode channel j, and C2 represents the cross-channel envelope influence coefficient and the cross-channel pitch influence coefficient.
[0450] In practical applications, you can choose to use either Mode 1 or Mode 2 according to the application scenario and task requirements, or you can choose to use a combination of Mode 1 and Mode 2.
[0451] For Mode 1:
[0452] This disclosure dynamically displays the waveforms, brain network characteristics, and source space features of EEG signals on a visual interface, allowing experts to observe and understand the current neural activity of the user in real time. This enables the development of targeted stimulation programs, and the dynamic adjustment of audio parameters through an interactive interface. This allows for targeted intervention against specific neural oscillation defects, achieving precise neural feedback reinforcement. For example, increasing the theta band gain enhances low-frequency note feedback, thereby promoting prefrontal theta wave synchronization and improving attention; decreasing the alpha band volume reduces over-excitation, thus alleviating anxiety. Furthermore, strategies can be flexibly adjusted based on the user's real-time responses (such as brainwave changes and behavioral performance), avoiding the limitations of a "one-size-fits-all" algorithm. For example, increasing rhythmic complexity for children with ADHD maintains interest, while simplifying the melody for anxious adults induces relaxation. Additionally, by manually identifying abnormal rhythms (such as epileptiform discharges) and instantly modifying audio parameters (such as inserting silence segments or low-frequency pulses), abnormal neural activity can be interrupted, reducing the risk of seizures and thus correcting abnormal EEG patterns.
[0453] For mode 2:
[0454] This disclosure automatically generates dynamic regulation parameters based on brain network characteristics, enabling millisecond-level responses to EEG changes and forming closed-loop regulation, thus improving the efficiency of real-time closed-loop feedback. For example, when high connectivity between the motor cortex and auditory cortex is detected, rhythm is automatically enhanced to strengthen the training effect of motor imagery. Furthermore, by driving audio modulation through node-to-node functional connections, functional integration of distal brain regions (such as the default mode network and executive control network) is promoted, thereby more accurately capturing functional associations between brain regions and improving the neural relevance of synthesized audio. For example, when there is high connectivity between the parietal and prefrontal lobes, the pitches of the two areas are fused to generate complex melodies, enhancing working memory.
[0455] In summary, this disclosure improves the neural modulation precision and artistic expressiveness of brainwave audio systems from the perspectives of human intelligence and computational intelligence through the two modes described above.
[0456] Figure 10This diagram illustrates the structure of a proprioceptive brainwave audio-visual perception synchronous feedback device according to an embodiment of the present disclosure. The device 1000 is mounted on a host computer, which is connected to an EEG acquisition device via an EEG data transmission interface. The device 1000 includes: an EEG signal receiving module, an EEG feature extraction module, and an EEG-audio generation module. The EEG signal receiving module is configured to receive EEG signals from the EEG acquisition device online via the EEG data transmission interface. The EEG signals are acquired by the EEG acquisition device through one or more electrode channels, where each electrode channel corresponds to one or more electrodes. The EEG feature extraction module is configured to acquire one or more specified frequency band EEG signals from the EEG signals, and extract and specify EEG features based on the one or more specified frequency band EEG signals. The EEG features corresponding to the specified frequency band EEG signals include EEG time-domain features and / or EEG frequency-domain features; the EEG-audio generation module is configured to generate an audio representation corresponding to the specified audio representation based on the specified audio representation form and a specified audio generation method, according to the EEG features, for the specified frequency band EEG signals, wherein the feature parameters of the audio representation are determined based on the EEG features; the audio representation form includes: note form and / or waveform form, wherein the audio representation corresponding to the note form is a note, and the audio representation corresponding to the waveform form is a waveform, and the audio generation method includes: direct mapping method and / or indirect control method using control parameters; generating brainwave audio representation data based on the audio representation corresponding to the one or more specified frequency band EEG signals; and obtaining brainwave audio based on the brainwave audio representation data.
[0457] Figure 11 The diagram illustrates the structure of another proprioceptive brainwave audio auditory perception synchronous feedback device according to an embodiment of the present disclosure. The EEG signal includes a denoised EEG signal. The device 1100 further includes an EEG preprocessing module, which is connected to the EEG signal receiving module and the EEG feature extraction module. The EEG preprocessing module is configured to: preprocess the EEG signals of the one or more electrode channels using a pre-trained brain-derived signal and non-brain-derived signal separation network model before acquiring one or more specified frequency band EEG signals from the EEG signal through the EEG feature extraction module, to obtain a denoised EEG signal, and transmit the denoised EEG signal to the EEG feature extraction module.
[0458] Figure 12The diagram illustrates the structure of another proprioceptive brainwave audio auditory perception synchronous feedback device according to an embodiment of the present disclosure. The device 1200 further includes a feedback module, which includes an audio feedback module. The feedback module is configured to synchronously feed back the brainwave audio to the user being collected through the audio feedback module, so as to dynamically regulate the nervous system activity of the user being collected through auditory perception.
[0459] According to an embodiment of this disclosure, the device further includes: a dynamic control module connected to the EEG feature extraction module, configured to: acquire dynamic control parameters; wherein the dynamic control parameters are generated based on stimulus task requirements, or based on brain network features corresponding to the EEG signals of the plurality of electrode channels; and adjust the feature parameters of the audio performance according to the dynamic control parameters.
[0460] According to an embodiment of this disclosure, the EEG preprocessing module is connected to the feedback module, and the feedback module further includes an EEG waveform feedback module; the feedback module is also configured to dynamically display the waveform of the denoised EEG signal in real time on a visualization interface through the EEG waveform feedback module.
[0461] According to embodiments of this disclosure, the EEG feature extraction module is further configured as follows:
[0462] Brain network features corresponding to the EEG signals of the multiple electrode channels are extracted. These features include node-to-node functional connectivity features and edge-to-edge interaction features. The node-to-node functional connectivity features describe the statistical dependency between two node signals, and the edge-to-edge interaction features describe the temporal dependency between two edges. Each node corresponds to a single electrode channel, each node signal corresponds to the EEG signal of a single electrode channel, and each edge corresponds to the connectivity between two nodes. The node-to-node functional connectivity features corresponding to the EEG signals of the multiple electrode channels are extracted by calculating the weighted phase lag coefficient between each pair of node signals. Calculating the weighted phase lag coefficient between any two node signals includes: performing a Hilbert transform on the first node signal and the second node signal to obtain a first node analytic signal and a second node analytic signal corresponding to the first node signal and the second node signal, respectively; obtaining the instantaneous cross-power spectrum of the first node signal and the second node signal based on the first node analytic signal and the second node analytic signal; averaging the imaginary absolute values of the instantaneous cross-power spectrum at each time point within a preset time window to obtain the mean of the imaginary absolute values; and calculating the... The weighted complex covariance is obtained based on the weights of the imaginary part of the instantaneous cross-power spectrum at each time point within a preset time window and the weights of the imaginary part corresponding to the instantaneous cross-power spectrum. The weighted phase lag coefficients of the first node signal and the second node signal are calculated using the absolute value of the imaginary part and the mean of the absolute values of the imaginary part of the weighted complex covariance. Edge-to-edge interaction features corresponding to the EEG signals of the multiple electrode channels are extracted by calculating the Pearson correlation coefficient between every two edges in the candidate edge set. The candidate edge set is obtained by calculating the average of all edges. The phase synchronization index is calculated by sorting all edges by their average phase synchronization index from largest to smallest and retaining the first preset number of edges. When calculating the Pearson correlation coefficient between any two edges in the candidate edge set, the process includes: calculating the instantaneous phase difference between the first edge and the second edge at each time point within a preset time window, where the instantaneous phase difference is the difference between the instantaneous phases of one node and another node in the edge; and calculating the sparse Pearson correlation coefficient between the first edge and the second edge using a sliding window based on the instantaneous phase differences of the first edge and the second edge at each time point within the preset time window.
[0463] According to embodiments of this disclosure, the EEG feature extraction module is connected to the feedback module, and the feedback module further includes a brain network functional connectivity feedback module; the feedback module is further configured to: dynamically draw electrodes corresponding to the plurality of electrode channels on an individualized three-dimensional head model of the visualization interface through the brain network functional connectivity feedback module, the electrodes being distributed on the scalp surface of the individualized three-dimensional head model; dynamically display the connection relationship and connection strength between the electrodes and / or between the electrode connections in real time according to the brain network features, the electrode display parameters input by the user, and the feature display parameters; the electrode display parameters are used to indicate the display of one or more electrodes corresponding to the plurality of electrode channels, and the feature display parameters are used to indicate the display of node-to-node functional connectivity features and / or edge-to-edge interaction features in the brain network features.
[0464] According to embodiments of this disclosure, the EEG feature extraction module is further configured as follows:
[0465] Extracting source space features corresponding to the EEG signals of the multiple electrode channels includes: obtaining a pre-calculated lead field pseudo-inverse matrix, and obtaining the source space features based on the pre-calculated lead field pseudo-inverse matrix and the EEG signals of the multiple electrode channels; wherein, the lead field pseudo-inverse matrix is determined based on the lead field matrix and a regularization parameter optimized offline, the lead field matrix is used to describe the electrical signal transmission relationship from the source space to the electrode space, the lead field matrix is calculated using the T1-weighted magnetic resonance imaging data of the user being collected as input, and a pre-constructed individualized three-dimensional head model corresponding to the user being collected; the regularization parameter optimized offline is the regularization parameter corresponding to the point of maximum curvature of the curve corresponding to the relationship curve of the residual norm and the solution norm.
[0466] According to embodiments of this disclosure, the EEG feature extraction module is connected to the feedback module, and the feedback module further includes an EEG source tracing result feedback module; the feedback module is further configured to: dynamically draw the source space activation map of the corresponding part of the individualized three-dimensional head model on the visualization interface according to the model display parameters input by the user, the individualized three-dimensional head model including: brain surface part, skull surface part and scalp surface part; and dynamically display the source space electrical activity intensity in real time in the source space activation map of the corresponding part of the displayed individualized three-dimensional head model according to the source space features.
[0467] According to an embodiment of this disclosure, the feedback module further includes: an acoustic image position feedback module; the feedback module is further configured to: obtain acoustic image position parameters of brainwave audio corresponding to multiple channels through the acoustic image position feedback module, and dynamically display the acoustic image position of the brainwave audio on a visualization interface according to the acoustic image position parameters.
[0468] This disclosure also discloses an electronic device. Figure 13 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown. Figure 13 As shown, the electronic device 1300 includes the apparatus described in the embodiments of this disclosure.
[0469] Those skilled in the art will understand that the device can be installed on one electronic device or on multiple electronic devices.
[0470] Taking the device as an example where it is mounted on two electronic devices, specifically, the electronic devices include: a first electronic device and a second electronic device. In a specific embodiment, such as... Figure 14 As shown, the EEG signal receiving module, EEG preprocessing module, and feedback module of the device are respectively disposed on the first electronic device, while the EEG feature extraction module and EEG-audio generation module of the device are disposed in the second electronic device; in another specific embodiment, as Figure 15 As shown, the EEG signal receiving module and feedback module of the device are respectively located on the first electronic device, while the EEG preprocessing module, EEG feature extraction module and EEG-audio generation module of the device are located in the second electronic device.
[0471] Figure 16 A structural block diagram of another electronic device according to an embodiment of the present disclosure is shown. (As follows) Figure 16 As shown, the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to embodiments of the present disclosure.
[0472] Figure 17 A structural block diagram of a proprioceptive brainwave audio-visual perception synchronous feedback system according to an embodiment of the present disclosure is shown. Figure 17 As shown, the system 1700 includes: an EEG acquisition device and an electronic device as described in the embodiments of this disclosure, wherein the EEG acquisition device and the electronic device are connected via an EEG data transmission interface, wherein:
[0473] The EEG acquisition device is configured to acquire the EEG signals of the user through one or more electrode channels, and transmit the EEG signals of the one or more electrode channels online to the electronic device based on the EEG data transmission interface.
[0474] In one specific embodiment, the proprioceptive brainwave audio auditory perception synchronous feedback system will be implemented based on an edge-cloud collaborative system. Specifically, an edge-edge connection is established with the EEG acquisition device as the edge and the accompanying computer (host computer) as the edge. The edge sends EEG signals to the edge at 30ms transmission intervals, and the edge sends the EEG signals to the cloud in real time. The cloud performs functions such as EEG feature extraction and EEG-audio generation, and returns the audio generation results to the edge in real time. The edge performs functions such as real-time dynamic feedback of various EEG feature calculation results, audio synthesis, and real-time playback through web-based development.
[0475] This disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system described in the above embodiments; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs, which are used by one or more processors to perform the methods described in this disclosure.
[0476] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements any of the methods described in this disclosure.
[0477] According to the technical solution provided in this disclosure, a host computer receives the EEG signals of the user being collected online from the EEG acquisition device, acquires one or more EEG signals of specified frequency bands, extracts EEG features corresponding to the specified frequency band EEG signals, generates audio representations corresponding to the specified audio representations based on specified audio representation formats and specified audio generation methods, and generates brainwave audio representation data based on the audio representation data corresponding to the one or more specified frequency band EEG signals. Brainwave audio is then obtained based on the brainwave audio representation data. When this brainwave audio is synchronously fed back to the user being collected, the user's nervous system activity can be dynamically controlled in real time through auditory perception. The brainwave audio obtained based on this technical solution can accurately match the user's EEG features. Since the EEG features are generated based on online real-time acquired EEG signals, this closed-loop EEG-audio real-time feedback interaction simultaneously satisfies the physiological adaptability of neural regulation and the artistic expressiveness of audio generation, achieving an organic unity between neural regulation and audio generation.
[0478] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A method of body brainwave audio hearing perception synchronization feedback, characterized in that, The method is applied to a host computer, which is connected to the EEG acquisition device via an EEG data transmission interface. The method includes: The EEG data transmission interface receives the EEG signals of the user being collected online from the EEG acquisition device. The EEG signals are acquired by the EEG acquisition device through one or more electrode channels, wherein each electrode channel corresponds to one electrode or a combination of multiple electrodes. One or more specified frequency band EEG signals are acquired from the EEG signals, and EEG features corresponding to the specified frequency band EEG signals are extracted based on the one or more specified frequency band EEG signals. The specified frequency band EEG signals refer to EEG signals located in the specified frequency band, and the EEG features include EEG time domain features and / or EEG frequency domain features. For the specified frequency band EEG signal, based on a specified audio representation and a specified audio generation method, the specified frequency band EEG signal is segmented into multiple EEG segments according to the EEG characteristics or preset rhythm parameters, and multiple audio representations corresponding to the multiple EEG segments are generated. The feature parameters of the audio representation are determined according to the EEG characteristics of the corresponding EEG segment. The audio representation includes: note form and / or waveform form, wherein the audio representation corresponding to the note form is a note, and the audio representation corresponding to the waveform form is a waveform. The audio generation method includes: direct mapping method and / or indirect control method using control parameters. Brainwave audio representation data is generated based on the audio representations corresponding to one or more specified frequency band EEG signals. Brainwave audio is obtained based on the brainwave audio representation data; When the specified audio representation is in the form of musical notes and the specified audio generation method is an indirect control method using control parameters, the characteristic parameters of the audio representation include: note duration, pitch, and intensity; the step of segmenting the specified frequency band EEG signal into multiple EEG segments based on the EEG characteristics and generating multiple audio representations corresponding to the multiple EEG segments respectively includes: The EEG signal of a preset duration in the specified frequency band is used as input, and a pre-trained user state discrimination model is used to generate control parameters, which include: arousal level and valence. The duration of a single note in the note sequence corresponding to the EEG signal of the preset duration is calculated based on the arousal level and / or the valence. The EEG signal of the preset duration is divided according to the duration of the note to obtain a set of EEG segments. The set of EEG segments contains multiple EEG segments, and each EEG segment corresponds to a note. The probability of occurrence of a note corresponding to each EEG segment in the set of EEG segments is calculated based on the arousal level and / or the valence. When the occurrence probability meets the occurrence condition, the note corresponding to the corresponding EEG segment appears; otherwise, the note corresponding to the corresponding EEG segment does not appear. The note sequence corresponding to the EEG signal of the preset duration is determined based on the calculation results. The intensity of the notes in the note sequence is calculated based on the arousal level, the valence, and the power spectrum energy characteristics of the EEG segments corresponding to the notes in the note sequence. Calculate the tonality parameters corresponding to the notes in the note sequence based on the arousal level and / or the valence; The pitch of the notes in the note sequence is calculated based on the tonality parameter and the power spectrum frequency characteristics of the EEG segments corresponding to the notes in the note sequence; Configure the timbre of the notes in the note sequence according to the arousal level and / or the valence; or, According to the preset configuration parameter-timbre mapping table, configure the timbre of all notes in the note sequence; the configuration parameter-timbre mapping table is used to describe the correspondence between the specified timbre configuration parameters and the timbre, and the specified timbre configuration parameters include: non-brain-derived control parameters and / or non-brain-derived event occurrence frequency.
2. The method according to claim 1, characterized in that, The method further includes: The brainwave audio is synchronously fed back to the user being collected, so as to dynamically regulate the nervous system activity of the user being collected through auditory perception.
3. The method according to claim 2, characterized in that, The step of synchronously feeding back the brainwave audio to the user whose brainwaves were collected includes: Based on the combination relationship between the spatial positions of electrodes in the one or more electrode channels and the one or more specified frequency bands, brainwave audio corresponding to multiple channels is obtained based on the brainwave audio. The brainwave audio from the multiple channels is played through an audio device so that the brainwave audio is synchronously fed back to the user whose brainwaves were collected.
4. The method according to claim 1, characterized in that, The EEG signal includes a denoised EEG signal, and before acquiring one or more specified frequency band EEG signals from the EEG signal, the method further includes: The EEG signals of one or more electrode channels are preprocessed using a pre-trained brain-derived signal and non-brain-derived signal separation network model to obtain denoised EEG signals. The brain-derived signal and non-brain-derived signal separation network model includes: a denoising module and a skip connection module. The denoising module includes: a decomposition module, a channel spatiotemporal attention processing module, and a reconstruction module. The preprocessing of the EEG signals from one or more electrode channels using a pre-trained brain-derived signal and non-brain-derived signal separation network model to obtain denoised EEG signals includes: The denoising module removes non-brain-derived signals from the EEG signals of one or more electrode channels, comprising: decomposing the EEG signals of one or more electrode channels into multidimensional embedding vectors corresponding to multiple signal channels, the multiple signal channels including brain-derived signal channels and non-brain-derived signal channels; performing channel attention processing and spatiotemporal attention processing on the multidimensional embedding vectors respectively by the channel spatiotemporal attention processing module to generate channel attention weights and temporal attention weights respectively; fusing the generated channel attention weights and temporal attention weights with the multidimensional embedding vectors corresponding to the multiple signal channels to obtain processed multidimensional embedding vectors; and generating reconstructed EEG signals by the reconstruction module based on the processed multidimensional embedding vectors. The processed multidimensional embedding vector is obtained through the skip connection module. Based on the processed multidimensional embedding vector and the EEG signals of the one or more electrode channels, dynamic fusion parameters are generated collaboratively using a gated linear unit (GLU) and a noise perception module (NAM). The EEG signals of the one or more electrode channels are fused with the reconstructed EEG signals according to the dynamic fusion parameters to obtain denoised EEG signals. The dynamic fusion parameters are used to control the fusion ratio of the EEG signals of the one or more electrode channels with the reconstructed EEG signals.
5. The method according to claim 4, characterized in that, The method further includes: The waveform of the denoised EEG signal is displayed dynamically in real time on a visual interface.
6. The method according to claim 1, characterized in that, in, The frequency domain features of the EEG include: EEG power spectrum features, and the time domain features of the EEG include: time domain envelope; When extracting EEG power spectrum features corresponding to the EEG signals of the specified frequency bands based on one or more specified frequency bands, the refined EEG power spectrum features corresponding to the EEG signals of the specified frequency bands are extracted using a frequency band-selective spectrum transformation algorithm. This includes: calculating a rotation factor and a starting point based on the start and end frequencies of the specified frequency band and the target resolution; performing the frequency band-selective spectrum transformation algorithm on the EEG signals of the specified frequency bands based on the rotation factor and the starting point to obtain the spectrum corresponding to the EEG signals of the specified frequency bands; and calculating the refined EEG power spectrum features corresponding to the EEG signals of the specified frequency bands based on the spectrum. The frequency band-selective spectrum transformation algorithm includes any one of the following transformation algorithms: Chirp-Z transform, Zoom-FFT, Goertzel algorithm, and Continuous Wavelet Transform (CWT).
7. The method according to claim 1, characterized in that, When the specified audio representation is in the form of musical notes and the specified audio generation method is direct mapping, the characteristic parameters of the audio representation include: note duration, pitch, and intensity; the step of segmenting the specified frequency band EEG signal into multiple EEG segments according to preset rhythm parameters and generating multiple audio representations corresponding to the multiple EEG segments respectively includes: For a predetermined duration of EEG signal in the specified frequency band of the specified electrode channel: The EEG signal of the preset duration is divided according to the note duration indicated by the preset rhythm parameters to obtain a set of EEG segments. The set of EEG segments contains multiple EEG segments, and each EEG segment corresponds to a note. For any EEG segment in the set of EEG segments, the pitch and intensity of the note corresponding to the any EEG segment are determined based on the frequency domain characteristics of the EEG segment, until the pitch and intensity of the note corresponding to each EEG segment in the set of EEG segments are obtained. This includes: selecting the peak frequency or weighted average frequency with the largest amplitude in the power spectrum characteristics of the power spectrum characteristics as the feature frequency; obtaining the corresponding specified pitch range based on the frequency range of the specified frequency band based on a preset mapping method; obtaining the pitch corresponding to the any EEG segment based on the feature frequency and the specified pitch range; generating the intensity corresponding to the any EEG segment based on the total energy or peak amplitude of the power spectrum characteristics; the preset mapping method is used to describe the mapping relationship between the frequency range of the specified frequency band and the pitch range.
8. The method according to claim 7, characterized in that, The step of generating brainwave audio representation data based on the audio representations corresponding to the one or more specified frequency band EEG signals includes: The notes corresponding to each EEG segment in the set of EEG segments of a preset duration in the specified frequency band of each electrode channel are written into the musical score corresponding to the specified frequency band. The musical score corresponding to each specified frequency band in each electrode channel is written into a complete musical score set corresponding to the preset duration, and the complete musical score set is used as the brainwave audio representation data. The step of obtaining brainwave audio based on the brainwave audio representation data includes: The timbre is configured for the musical score corresponding to each specified frequency band in the brainwave audio representation data according to a preset frequency band-timbre mapping table or a configuration parameter-timbre mapping table; the frequency band-timbre mapping table is used to describe the correspondence between frequency bands and timbres, and the configuration parameter-timbre mapping table is used to describe the correspondence between specified timbre configuration parameters and timbres; the specified timbre configuration parameters include: non-brain-derived control parameters and / or non-brain-derived event occurrence frequencies; The non-brain-derived control parameter is obtained as follows: when the EEG signal of the preset length in the specified electrode channel contains both non-brain-derived signals and denoised EEG signals, the non-brain-derived control parameter is obtained by taking the ratio of the absolute values of the non-brain-derived signals and the denoised EEG signals in the preset length of the EEG signal in the specified electrode channel to the preset fractional part. The non-brain-derived event frequency is obtained as follows: a non-brain-derived event threshold is obtained by taking the absolute value of the denoised EEG signal and dividing it by a preset fractional number; positions in the non-brain-derived signal that are higher than the non-brain-derived event threshold are counted to obtain the number of non-brain-derived events in the non-brain-derived signal; and the non-brain-derived event frequency is obtained based on the number of non-brain-derived events and the preset duration. The brainwave audio is generated using a MIDI synthesizer or audio synthesis library based on the musical score corresponding to each specified frequency band in the brainwave audio representation data.
9. The method according to claim 1, characterized in that, in, The non-brain-derived control parameter is obtained as follows: when the EEG signal of the preset length in the specified electrode channel contains both non-brain-derived signals and denoised EEG signals, the ratio of the preset fractional part of the absolute values of the non-brain-derived signals and the denoised EEG signals in the preset length of the EEG signal in the specified electrode channel is taken as the non-brain-derived control parameter. The non-brain-derived event frequency is obtained as follows: a non-brain-derived event threshold is obtained by taking the absolute value of the denoised EEG signal and dividing it by a preset fractional number; positions in the non-brain-derived signal that are higher than the non-brain-derived event threshold are counted to obtain the number of non-brain-derived events occurring in the non-brain-derived signal; and the non-brain-derived event frequency is obtained based on the number of non-brain-derived events and the preset duration.
10. The method according to claim 1, characterized in that, When the specified audio representation is in the form of musical notes and the specified audio generation method is an indirect control method using control parameters, generating brainwave audio representation data based on the audio representations corresponding to the one or more specified frequency band EEG signals includes: Based on the note sequence of the specified frequency band in each electrode channel, and the intensity, pitch, and timbre of the notes in the note sequence, a musical score corresponding to the specified frequency band is generated; The musical score corresponding to each specified frequency band in each electrode channel is written into a complete musical score set corresponding to the preset duration, and the complete musical score set is used as the brainwave audio representation data. The step of obtaining brainwave audio based on the brainwave audio representation data includes: The brainwave audio is generated using a MIDI synthesizer or audio synthesis library based on the musical score corresponding to each specified frequency band in the brainwave audio representation data.
11. The method according to claim 1, characterized in that, When the specified audio representation is a waveform and the specified audio generation method is a direct mapping method, the characteristic parameters of the audio representation include: fundamental frequency and harmonic frequency, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency; the step of segmenting the specified frequency band EEG signal into multiple EEG segments according to the EEG characteristics and generating multiple audio representations corresponding to the multiple EEG segments respectively includes: For a predetermined duration of EEG signal in the specified frequency band of the specified electrode channel: Based on the temporal envelope corresponding to the EEG signal of the preset duration, the EEG signal of the preset duration is segmented to obtain a set of EEG segments, including: determining an effective amplitude threshold based on the mean of the temporal envelope; determining a set of troughs in the temporal envelope based on the effective amplitude threshold, wherein the amplitude of the troughs in the set of troughs is not greater than the effective amplitude threshold; segmenting the EEG signal of the preset duration based on a preset effective time threshold and the set of troughs in the temporal envelope to obtain the set of EEG segments, wherein the set of EEG segments contains one or more EEG segments, and the time value of each EEG segment is greater than the effective time threshold; For any EEG segment in the set of EEG segments, the characteristic parameters of the waveform corresponding to any EEG segment are determined based on the frequency domain characteristics of the EEG segment, until the characteristic parameters of the waveform corresponding to each EEG segment in the set of EEG segments are obtained. This includes: based on the power spectrum characteristics of any EEG segment, determining the fundamental frequency and harmonic frequency of the waveform corresponding to any EEG segment, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency, based on the peak frequencies and amplitudes of the top preset number of peak frequencies ranked by amplitude in the power spectrum characteristics.
12. The method according to claim 11, characterized in that, The step of generating brainwave audio representation data based on the audio representations corresponding to the one or more specified frequency band EEG signals includes: For any EEG segment in the set of EEG segments, the audio waveform corresponding to any EEG segment is determined based on the characteristic parameters and time-domain envelope of the waveform corresponding to any EEG segment, until the audio waveforms corresponding to each EEG segment in the set of EEG segments are obtained. This includes: generating sine waves of corresponding frequencies based on the fundamental frequency and harmonic frequency of the waveform corresponding to any EEG segment, as well as the intensity of the fundamental frequency and the intensity of the harmonic frequency, and mixing them to obtain a mixed sine wave; applying the time-domain envelope corresponding to any EEG segment to the mixed sine wave for modulation to obtain the audio waveform corresponding to any EEG segment. The audio waveforms corresponding to each EEG segment in the set of EEG segments in each specified frequency band of the specified electrode channel are used as the audio waveforms corresponding to the EEG signals of the specified electrode channel with the preset duration. The audio waveforms corresponding to the EEG signals of each electrode channel for the preset duration are used as the EEG audio representation data. The step of obtaining brainwave audio based on the brainwave audio representation data includes: The brainwave audio representation data is directly used as brainwave audio.
13. The method according to claim 1, characterized in that, The method further includes: Brain network features corresponding to the EEG signals of the multiple electrode channels are extracted. The brain network features include: node-node functional connectivity features and / or edge-edge interaction features. The node-node functional connectivity features are used to describe the statistical dependency between two node signals, and the edge-edge interaction features are used to describe the temporal dependency between two edges. Wherein, a node corresponds to a single electrode channel, a node signal corresponds to the EEG signal of a single electrode channel, and an edge corresponds to the connectivity between two nodes. The method involves extracting node-to-node functional connectivity features corresponding to the EEG signals of the multiple electrode channels by calculating the weighted phase lag coefficient between every two node signals. The calculation of the weighted phase lag coefficient between any two node signals includes: performing a Hilbert transform on the first node signal and the second node signal to obtain a first node analytic signal and a second node analytic signal corresponding to the first node signal and the second node signal, respectively; obtaining the instantaneous cross-power spectrum of the first node signal and the second node signal based on the first node analytic signal and the second node analytic signal; averaging the imaginary absolute values of the instantaneous cross-power spectrum at each time point within a preset time window to obtain the mean of the imaginary absolute values; calculating the weight of the imaginary part of the instantaneous cross-power spectrum at each time point within the preset time window; obtaining the weighted complex covariance based on the instantaneous cross-power spectrum at each time point within the preset time window and the weight of the imaginary part corresponding to the instantaneous cross-power spectrum; and calculating the weighted phase lag coefficient between the first node signal and the second node signal using the imaginary absolute value of the weighted complex covariance and the mean of the imaginary absolute values. Edge-to-edge interaction features corresponding to the EEG signals of the multiple electrode channels are extracted by calculating the Pearson correlation coefficient between every two edges in the candidate edge set. The candidate edge set is constructed by calculating the average phase synchronization index of all edges, sorting the average phase synchronization index of all edges from largest to smallest, and retaining the first preset number of edges. When calculating the Pearson correlation coefficient between any two edges in the candidate edge set, the process includes: calculating the instantaneous phase difference of the first edge and the second edge at each time point within a preset time window, wherein the instantaneous phase difference is the difference between the instantaneous phases of one node and another node in the edge; and calculating the sparse Pearson correlation coefficient between the first edge and the second edge using a sliding window based on the instantaneous phase difference of the first edge and the instantaneous phase difference of the second edge at each time point within the preset time window.
14. The method according to claim 13, characterized in that, in, The first node parse signal and the second node parse signal, corresponding to the first node signal and the second node signal, are obtained respectively by the following formulas: ; ; The instantaneous cross power spectrum is obtained using the following formula: ; The weight of the imaginary part of the instantaneous cross-power spectrum at each time point within the preset time window is calculated using the following formula: ; ; The weighted complex covariance is obtained using the following formula: ; The weighted phase lag coefficient is calculated using the following formula: ; in, This is the first node signal. This is the signal from the second node. Analyze the signal for the first node. The second node's analyzed signal, where j is the imaginary unit. For Hilbert transform operators; Represents the instantaneous cross power spectrum. Represents complex conjugation, median For median operations, The weighted complex covariance is T, where T is the number of time points within the preset time window.
15. The method according to claim 13, characterized in that, in, The average phase synchronization index of the edge is calculated using the following formula: ; The instantaneous phase of a node is calculated using the following formula: ; ; The instantaneous phase difference of the first side and the instantaneous phase difference of the second side can be obtained using the following formulas: ; ; The Pearson correlation coefficient between the first side and the second side is calculated using the following formula: ; ; ; in, For node signals, For the node-analyzed signal, j is the imaginary unit. For Hilbert transform operators; This represents the instantaneous phase of the node at time point t. This represents taking the phase angle. This represents the instantaneous phase of a node in the first edge at time point t. This represents the instantaneous phase of the other node in the first edge at time point t. This represents the instantaneous phase of a node in the second edge at time point t. This represents the instantaneous phase of the other node in the second edge at time point t. This represents the instantaneous phase difference of the first edge at time point t. This represents the instantaneous phase difference of the second side at time point t, where T is the number of time points within the preset time window. The length of the sliding window. The average instantaneous phase difference of the first side within the sliding window. The average instantaneous phase difference of the second side within the sliding window.
16. The method according to claim 13, characterized in that, The method further includes: Electrodes corresponding to the multiple electrode channels are dynamically drawn on the individualized 3D head model of the visualization interface. The electrodes are distributed on the scalp surface of the individualized 3D head model. The connection relationship and connection strength between the electrodes and / or between the electrode connections are dynamically displayed in real time according to the brain network features, the electrode display parameters input by the user, and the feature display parameters are used to indicate the display of one or more electrodes corresponding to the multiple electrode channels. The feature display parameters are used to indicate the display of node-node functional connection features and / or edge-edge interaction features in the brain network features.
17. The method according to claim 1, characterized in that, The method further includes: Extracting source space features corresponding to the EEG signals of the multiple electrode channels includes: obtaining a pre-calculated lead field pseudo-inverse matrix, and obtaining the source space features based on the pre-calculated lead field pseudo-inverse matrix and the EEG signals of the multiple electrode channels; wherein, the lead field pseudo-inverse matrix is determined based on the lead field matrix and a regularization parameter optimized offline, the lead field matrix is used to describe the electrical signal transmission relationship from the source space to the electrode space, the lead field matrix is calculated using the T1-weighted magnetic resonance imaging data of the user being collected as input, and a pre-constructed individualized three-dimensional head model corresponding to the user being collected; the regularization parameter optimized offline is the regularization parameter corresponding to the point of maximum curvature of the curve corresponding to the relationship curve of the residual norm and the solution norm.
18. The method according to claim 17, characterized in that, in, The source space features are calculated using the following formula: ; Alternatively, the source space features can be calculated using the following formula: ; ; L; in, Let be the source space features at time t. The EEG signals of the multiple electrode channels corresponding to time t. The pseudo-inverse matrix of the lead field is... Let be the lead field matrix. , Where D is the number of scalp electrodes, and D is the dimension of the source space. This is the transpose operation of a matrix, where I is the identity matrix. The regularization parameters are those optimized offline. This indicates the extraction of diagonal elements from a matrix.
19. The method according to claim 17, characterized in that, The method further includes: Based on the model display parameters input by the user, the source space activation map of the corresponding part of the individualized three-dimensional head model is dynamically drawn on the visualization interface. The individualized three-dimensional head model includes: the surface part of the brain, the surface part of the skull, and the surface part of the scalp. Based on the source space features, the source space electrical activity intensity is dynamically displayed in real time in the source space activation map of the corresponding part of the displayed individualized three-dimensional head model.
20. The method according to claim 1, characterized in that, The method further includes: Obtain the acoustic image position parameters of the brainwave audio corresponding to multiple channels, and dynamically display the acoustic image position of the brainwave audio on the visualization interface based on the acoustic image position parameters.
21. The method according to claim 20, characterized in that, The acquisition of the acoustic image location parameters of brainwave audio corresponding to multiple channels includes: The brainwave audio corresponding to each specified frequency band in the brainwave audio of the multiple channels is obtained, and the intensity of the brainwave audio corresponding to each specified frequency band in each channel within a specified time period is calculated; based on the intensity difference of the brainwave audio corresponding to the same specified frequency band in different channels, the acoustic image orientation indication parameter corresponding to the brainwave audio of the specified frequency band is obtained. The acoustic image orientation indicator parameters and intensity corresponding to the brainwave audio of each specified frequency band in the multiple channels are used as the acoustic image position parameters.
22. The method according to claim 20, characterized in that, The method further includes: The sound image location of the brainwave audio is dynamically and synchronously fed back to the user being collected on a visual interface, so as to dynamically regulate the neural activity of the user being collected through visual perception.
23. The method according to claim 1, characterized in that, The method further includes: Acquire dynamic control parameters; wherein the dynamic control parameters are generated based on the stimulus task requirements, or based on the brain network features corresponding to the EEG signals of the multiple electrode channels; The characteristic parameters of the audio performance are adjusted according to the dynamic control parameters.
24. A proprioceptive brainwave audio-visual perception synchronous feedback device, characterized in that, The device is installed on a host computer, which is connected to the EEG acquisition device via an EEG data transmission interface. The device includes: an EEG signal receiving module, an EEG feature extraction module, and an EEG-audio generation module. The EEG signal receiving module is configured to receive EEG signals of the user being collected online from the EEG acquisition device through the EEG data transmission interface. The EEG signals are acquired by the EEG acquisition device through one or more electrode channels, wherein each electrode channel corresponds to one electrode or a combination of multiple electrodes. The EEG feature extraction module is configured to acquire one or more EEG signals of a specified frequency band from the EEG signals, and extract EEG features corresponding to the EEG signals of the specified frequency bands based on the one or more EEG signals of the specified frequency bands. The EEG signals of the specified frequency bands refer to EEG signals located in the specified frequency bands. The EEG features include EEG time-domain features and / or EEG frequency-domain features. The EEG-audio generation module is configured to, for the specified frequency band EEG signal, based on a specified audio representation format and a specified audio generation method, segment the specified frequency band EEG signal into multiple EEG segments according to the EEG characteristics or preset rhythm parameters, and generate multiple audio representations corresponding to the multiple EEG segments respectively. The feature parameters of the audio representation are determined based on the EEG characteristics of the corresponding EEG segment. The audio representation format includes: note form and / or waveform form, wherein the audio representation corresponding to the note form is a note, and the audio representation corresponding to the waveform form is a waveform. The audio generation method includes: direct mapping method and / or indirect control method using control parameters. The module generates brainwave audio representation data based on the audio representations corresponding to one or more specified frequency band EEG signals; and obtains brainwave audio based on the brainwave audio representation data. When the specified audio representation is in the form of musical notes and the specified audio generation method is an indirect control method using control parameters, the characteristic parameters of the audio representation include: note duration, pitch, and intensity; the step of segmenting the specified frequency band EEG signal into multiple EEG segments based on the EEG characteristics and generating multiple audio representations corresponding to the multiple EEG segments respectively includes: The EEG signal of a preset duration in the specified frequency band is used as input, and a pre-trained user state discrimination model is used to generate control parameters, which include: arousal level and valence. The duration of a single note in the note sequence corresponding to the EEG signal of the preset duration is calculated based on the arousal level and / or the valence. The EEG signal of the preset duration is divided according to the duration of the note to obtain a set of EEG segments. The set of EEG segments contains multiple EEG segments, and each EEG segment corresponds to a note. The probability of occurrence of a note corresponding to each EEG segment in the set of EEG segments is calculated based on the arousal level and / or the valence. When the occurrence probability meets the occurrence condition, the note corresponding to the corresponding EEG segment appears; otherwise, the note corresponding to the corresponding EEG segment does not appear. The note sequence corresponding to the EEG signal of the preset duration is determined based on the calculation results. The intensity of the notes in the note sequence is calculated based on the arousal level, the valence, and the power spectrum energy characteristics of the EEG segments corresponding to the notes in the note sequence. Calculate the tonality parameters corresponding to the notes in the note sequence based on the arousal level and / or the valence; The pitch of the notes in the note sequence is calculated based on the tonality parameter and the power spectrum frequency characteristics of the EEG segments corresponding to the notes in the note sequence; Configure the timbre of the notes in the note sequence according to the arousal level and / or the valence; or, According to the preset configuration parameter-timbre mapping table, configure the timbre of all notes in the note sequence; the configuration parameter-timbre mapping table is used to describe the correspondence between the specified timbre configuration parameters and the timbre, and the specified timbre configuration parameters include: non-brain-derived control parameters and / or non-brain-derived event occurrence frequency.
25. The apparatus according to claim 24, characterized in that, The device further includes a feedback module, which includes an audio feedback module. The feedback module is configured to synchronously feed back the brainwave audio to the user being collected through the audio feedback module, so as to dynamically regulate the nervous system activity of the user being collected through auditory perception.
26. The apparatus according to claim 25, characterized in that, The EEG signal includes a denoised EEG signal, and the device further includes an EEG preprocessing module, which is connected to the EEG signal receiving module and the EEG feature extraction module respectively. The EEG preprocessing module is configured to: before acquiring one or more specified frequency band EEG signals from the EEG signals through the EEG feature extraction module, preprocess the EEG signals from the one or more electrode channels using a pre-trained brain-derived signal and non-brain-derived signal separation network model to obtain denoised EEG signals, and transmit the denoised EEG signals to the EEG feature extraction module.
27. The apparatus according to claim 26, characterized in that, The EEG preprocessing module is connected to the feedback module, and the feedback module further includes an EEG waveform feedback module; The feedback module is also configured to dynamically display the waveform of the denoised EEG signal in real time on a visualization interface through the EEG waveform feedback module.
28. The apparatus according to claim 25, characterized in that, The EEG feature extraction module is further configured to: Brain network features corresponding to the EEG signals of the multiple electrode channels are extracted. The brain network features include: node-node functional connectivity features and edge-edge interaction features. The node-node functional connectivity features are used to describe the statistical dependency between two node signals, and the edge-edge interaction features are used to describe the temporal dependency between two edges. In this case, a node corresponds to a single electrode channel, a node signal corresponds to the EEG signal of a single electrode channel, and an edge corresponds to the connectivity between two nodes. The method involves extracting node-to-node functional connectivity features corresponding to the EEG signals of the multiple electrode channels by calculating the weighted phase lag coefficient between every two node signals. The calculation of the weighted phase lag coefficient between any two node signals includes: performing a Hilbert transform on the first node signal and the second node signal to obtain a first node analytic signal and a second node analytic signal corresponding to the first node signal and the second node signal, respectively; obtaining the instantaneous cross-power spectrum of the first node signal and the second node signal based on the first node analytic signal and the second node analytic signal; averaging the imaginary absolute values of the instantaneous cross-power spectrum at each time point within a preset time window to obtain the mean of the imaginary absolute values; calculating the weight of the imaginary part of the instantaneous cross-power spectrum at each time point within the preset time window; obtaining the weighted complex covariance based on the instantaneous cross-power spectrum at each time point within the preset time window and the weight of the imaginary part corresponding to the instantaneous cross-power spectrum; and calculating the weighted phase lag coefficient between the first node signal and the second node signal using the imaginary absolute value of the weighted complex covariance and the mean of the imaginary absolute values. Edge-to-edge interaction features corresponding to the EEG signals of the multiple electrode channels are extracted by calculating the Pearson correlation coefficient between every two edges in the candidate edge set. The candidate edge set is constructed by calculating the average phase synchronization index of all edges, sorting the average phase synchronization index of all edges from largest to smallest, and retaining the first preset number of edges. When calculating the Pearson correlation coefficient between any two edges in the candidate edge set, the process includes: calculating the instantaneous phase difference of the first edge and the second edge at each time point within a preset time window, wherein the instantaneous phase difference is the difference between the instantaneous phases of one node and another node in the edge; and calculating the sparse Pearson correlation coefficient between the first edge and the second edge using a sliding window based on the instantaneous phase difference of the first edge and the instantaneous phase difference of the second edge at each time point within the preset time window.
29. The apparatus according to claim 28, characterized in that, The EEG feature extraction module is connected to the feedback module, and the feedback module further includes a brain network function connection feedback module. The feedback module is further configured to dynamically draw electrodes corresponding to the multiple electrode channels on the individualized 3D head model of the visualization interface through the brain network functional connection feedback module, the electrodes being distributed on the scalp surface of the individualized 3D head model; and to dynamically display the connection relationship and connection strength between the electrodes and / or between the electrode connections in real time according to the brain network features, the electrode display parameters input by the user, and the feature display parameters; the electrode display parameters are used to indicate the display of one or more electrodes corresponding to the multiple electrode channels, and the feature display parameters are used to indicate the display of node-node functional connection features and / or edge-edge interaction features in the brain network features.
30. The apparatus according to claim 25, characterized in that, The EEG feature extraction module is further configured to: Extracting source space features corresponding to the EEG signals of the multiple electrode channels includes: obtaining a pre-calculated lead field pseudo-inverse matrix, and obtaining the source space features based on the pre-calculated lead field pseudo-inverse matrix and the EEG signals of the multiple electrode channels; wherein, the lead field pseudo-inverse matrix is determined based on the lead field matrix and a regularization parameter optimized offline, the lead field matrix is used to describe the electrical signal transmission relationship from the source space to the electrode space, the lead field matrix is calculated using the T1-weighted magnetic resonance imaging data of the user being collected as input, and a pre-constructed individualized three-dimensional head model corresponding to the user being collected; the regularization parameter optimized offline is the regularization parameter corresponding to the point of maximum curvature of the curve corresponding to the relationship curve of the residual norm and the solution norm.
31. The apparatus according to claim 30, characterized in that, The EEG feature extraction module is connected to the feedback module, and the feedback module further includes an EEG source tracing result feedback module; The feedback module is further configured to dynamically draw the source space activation map of the corresponding part of the individualized three-dimensional head model on the visualization interface according to the model display parameters input by the user through the EEG source tracing result feedback module. The individualized three-dimensional head model includes: the surface part of the brain, the surface part of the skull, and the surface part of the scalp. The source space electrical activity intensity is dynamically displayed in real time in the source space activation map of the corresponding part of the displayed individualized three-dimensional head model according to the source space characteristics.
32. The apparatus according to claim 25, characterized in that, The feedback module further includes: an acoustic position feedback module; The feedback module is further configured to obtain the sound image position parameters of the brainwave audio corresponding to multiple channels through the sound image position feedback module, and dynamically display the sound image position of the brainwave audio on the visualization interface according to the sound image position parameters.
33. An electronic device, characterized in that, The apparatus includes any one of claims 24 to 32, or includes a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method of any one of claims 1 to 23.
34. A proprioceptive brainwave audio-visual perception synchronous feedback system, characterized in that, The system includes: an electroencephalogram (EEG) acquisition device and the electronic device as described in claim 33, wherein the EEG acquisition device and the electronic device are connected via an EEG data transmission interface, wherein: The EEG acquisition device is configured to acquire the EEG signals of the user through one or more electrode channels, and transmit the EEG signals of the one or more electrode channels online to the electronic device based on the EEG data transmission interface.
35. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, they implement the method of any one of claims 1 to 23.
36. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 23.
Citation Information
Patent Citations
Method and device for converting electroencephalogram into audio, electronic equipment and storage medium
CN115317001A