Method and system for generating natural sound based on personalized EEG signals
By extracting and fusion of feature vectors of scene spatial data and user EEG data, combining self-attention mechanism and comparative learning to generate Mel spectral features of natural sound, the HiFi-GAN model is finally generated, which solves the problem that the existing technology cannot dynamically adjust natural sound in real time and realizes personalized adjustment of user creativity.
Patent Information
- Application Number
- CN202510253398.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The prior art is unable to dynamically adjust natural sounds in real time based on environmental changes and user physiological data, resulting in limitations in improving personal creativity.
By obtaining multi-scene data and natural sound data, the feature vectors are extracted using BERT, LaBraM, and HTS-AT pre-trained models, and the Mel spectral features of natural sound are generated using self-attention mechanism and comparison learning. Finally, the corresponding natural sound is generated through the HiFi-GAN model.
It realizes the generation of natural sounds based on the user's EEG data and scene data collected in real time, and dynamically adjusts natural sounds through the creativity adjustment mechanism, improving the personalized adjustment effect of user creativity.
Smart Images

Figure CN119767214B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to a method and system for generating natural sounds based on personalized electroencephalogram signals. Background Art
[0002] With the development of science and technology, people pay more and more attention to the role of creativity in personal and social development. Studies have shown that the sensory experience brought by the sound in the natural environment can significantly enhance the creativity of individuals, and the level of individual creativity is positively correlated with the amplitude of EEG alpha waves.
[0003] However, the traditional method of acquiring natural sounds through streaming media only involves audio acquisition but not personalized audio generation. It cannot be adjusted dynamically in real time according to environmental changes and user physiological data, and therefore has certain limitations in improving personal creativity. At the same time, existing audio generation technologies are mostly concentrated in the field of music generation. They usually use preset music libraries, parameter mapping, symbol generation, and audio generation to generate music to adapt to different scenarios. These methods lack adaptability and interactivity, and the generated audio is mostly music-based, with few natural sounds. In addition, existing application scenarios for using data to generate audio are widely used in specific fields such as music therapy, psychological healing, and sleep assistance, but few focus on the application scenario of creativity regulation. Summary of the invention
[0004] Based on this, it is necessary to provide a method and system for personalized generation of natural sounds based on EEG signals to address the above technical problems.
[0005] A method for generating natural sounds based on personalized brain electrical signals, the method comprising:
[0006] Acquire multi-scenario data and natural sound data; multi-scenario data includes scene space data and user EEG data.
[0007] Use the BERT pre-trained model to extract the feature vector of the text corresponding to the scene space data;
[0008] The LaBraM pre-trained model is used to extract the feature vector of the user's EEG data;
[0009] The feature vector of the text corresponding to the scene space data and the feature vector of the user's EEG data are fused and processed using the self-attention mechanism to obtain the feature vector of the multivariate scene data.
[0010] The HTS-AT pre-trained model is used to extract feature vectors of natural sound data.
[0011] The feature vectors of multivariate scene data and the feature vectors of natural sound data are compared and learned to obtain a high-dimensional feature representation of the multivariate scene data.
[0012] According to the high-dimensional feature representation of random Gaussian noise vector and multivariate scene data, the conditional controlled diffusion model is used to generate the Mel spectrum features of natural sound.
[0013] According to the Mel-spectrogram features of natural sounds, the HiFi-GAN pre-trained model is used to generate the corresponding natural sounds.
[0014] Calculate the difference in the alpha wave amplitude values of the user's EEG data in adjacent time windows, determine the change trend of the alpha wave, and through the creativity adjustment mechanism, maintain the natural sound of the previous time window or generate a new natural sound.
[0015] A system for generating natural sounds based on personalized brain electrical signals, the system comprising: an information acquisition module, a data processing module, a personalized natural sound generation module and a creativity adjustment module based on natural sounds.
[0016] The information collection module is used to collect user EEG data and on-site scene space data using sensor devices, mobile devices, and audio devices; and obtain natural sound data from open source natural sound datasets.
[0017] The data processing module is used to extract the feature vector of the text corresponding to the scene space data by using the BERT pre-training model; extract the feature vector of the user's EEG data by using the LaBraM pre-training model; fuse the feature vector of the text corresponding to the scene space data with the feature vector of the user's EEG data and process them by using the self-attention mechanism to obtain the feature vector of the multivariate scene data; extract the feature vector of the natural sound data by using the HTS-AT pre-training model; perform comparative learning on the feature vectors of the multivariate scene data and the feature vectors of the natural sound data to obtain the high-dimensional feature representation of the multivariate scene data.
[0018] The personalized natural sound generation module is used to generate the Mel spectrum features of natural sounds based on the high-dimensional feature representation of random Gaussian noise vectors and multivariate scene data, and use the conditional controlled diffusion model for processing; based on the Mel spectrum features of natural sounds, the HiFi-GAN pre-trained model is used to generate the corresponding natural sounds.
[0019] The creativity adjustment module based on natural sounds is used to calculate the difference in the alpha wave amplitude values of the user's EEG data in adjacent time windows, determine the change trend of the alpha wave, and through the creativity adjustment mechanism, maintain the natural sound of the previous time window or generate a new natural sound.
[0020] The above-mentioned method and system for generating natural sounds based on personalized EEG signals, the method generates natural sounds based on the user's EEG data and scene data collected in real time, calculates the alpha wave amplitude value related to creativity, and judges the change trend of the alpha wave based on the alpha wave amplitude value of adjacent time windows, and dynamically generates natural sounds that are highly adapted to the personalized user EEG data and scene space data through the creativity adjustment mechanism, thereby personalizing the user's creativity. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram of a process for generating natural sounds based on personalized EEG signals in one embodiment;
[0022] Figure 2 A flowchart of a method for generating natural sounds based on personalized EEG signals in one embodiment;
[0023] Figure 3 A block diagram of a system for generating natural sounds based on EEG signals in an individualized manner in one embodiment. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0025] In one embodiment, Figure 1 , Figure 2 As shown, a method for generating natural sounds based on personalized EEG signals is provided, the method comprising the following steps:
[0026] Step 100: Acquire multivariate scene data and natural sound data; the multivariate scene data includes scene space data and user EEG data.
[0027] Specifically, all devices in the scene where the current user is located are obtained, including sensor devices, mobile devices, and audio devices. Personalized user EEG data and scene space data are obtained, and the user EEG data and scene space data are integrated to obtain multi-dimensional scene data. Specifically, it includes: through the sensor devices in the acquired devices, according to the preset collection frequency and time window, personalized user EEG data and scene space data are synchronously obtained.
[0028] Get natural sound data from the open source natural sound dataset.
[0029] Scene spatial data includes: lighting, weather conditions, time and region; among them, lighting elements include natural light (such as sunlight and moonlight), artificial light (such as lamplight and firelight), etc., which affect the visual effect of the scene; weather conditions such as temperature, such as sunny days, rainy days, snowy days, foggy days, etc., can change the atmosphere and visual effect of the scene; time elements can be different times of the day (such as morning, noon, and evening), seasonal changes (such as spring, summer, autumn, and winter), and festivals (such as the Spring Festival, Mid-Autumn Festival, and Double Ninth Festival); region includes the judgment of region type, such as plateau, plain, mountain, coast, lakeside, etc., as well as local altitude, etc.
[0030] The user's EEG data includes: channels related to creativity and frequency bands related to creativity. Taking the 32-bit electrode EEG device as an example, the user's EEG data is collected, and the EEG data of the frontal lobe and right parietal lobe are used. The data of these points are related to the user's creativity. Alpha waves are frequency bands related to creativity: the individual's creativity level is positively correlated with the amplitude of the alpha wave; the frequency range of the alpha wave is 8-13HZ (the amplitude range of the normal person's EEG signal is 10µV-200µV, and the frequency is 0.2Hz-90Hz).
[0031] Step 102: Use the BERT pre-trained model to extract the feature vector of the text corresponding to the scene space data.
[0032] Step 104: Use the LaBraM pre-trained model to extract the feature vector of the user's EEG data.
[0033] Step 106: The feature vector of the text corresponding to the scene space data and the feature vector of the user's EEG data are fused and processed using a self-attention mechanism to obtain a feature vector of the multivariate scene data.
[0034] Step 108: Use the HTS-AT pre-trained model to extract feature vectors of natural sound data.
[0035] Specifically, the comparison scene natural sound model (CSNSP model) includes: BERT pre-training model, LaBraM pre-training model, HTS-AT pre-training model and self-attention mechanism.
[0036] Step 110: Perform comparative learning on the feature vectors of the multivariate scene data and the feature vectors of the natural sound data to obtain a high-dimensional feature representation of the multivariate scene data.
[0037] Step 112: Based on the high-dimensional feature representation of the random Gaussian noise vector and the multivariate scene data, a conditional controlled diffusion model is used for processing to generate a Mel-spectrogram feature of the natural sound.
[0038] Specifically, the Mel spectrum features of natural sound are mapped to low-dimensional latent feature vectors through the encoder in the VAE pre-trained model, and the forward diffusion process and reverse diffusion process of the conditional controlled diffusion model are performed in sequence. In the forward diffusion process, Gaussian noise of T time steps is added to the low-dimensional latent vector; in the reverse diffusion process, the feature vector of the multivariate scene data is used as a control condition to predict the Gaussian noise added to the low-dimensional latent vector in the forward diffusion process, and finally a low-dimensional latent feature vector is generated after the noise is removed and reconstructed, and the reconstructed low-dimensional latent feature vector is mapped to the Mel spectrum features of the reconstructed natural sound data through the decoder in the VAE pre-trained model.
[0039] The training process of the conditional controlled diffusion model includes: obtaining the Mel spectrum features of natural sound data through short-time Fourier transform, Mel frequency scale and Mel filter group, and obtaining the high-dimensional feature representation of multivariate scene data through the CSNSP model. The Mel spectrum features of natural sound data and the feature vector of multivariate scene data are input into the conditional controlled diffusion model. Preferably, the Mel spectrum features of natural sound are mapped to low-dimensional latent feature vectors through the encoder in the VAE pre-trained model, and the forward diffusion process and the reverse diffusion process of the conditional controlled diffusion model are performed in sequence. In the forward diffusion process, Gaussian noise of T time steps is added to the low-dimensional latent vector; in the reverse diffusion process, the high-dimensional feature representation of the multivariate scene data is used as a control condition to predict the Gaussian noise added to the low-dimensional latent vector in the forward diffusion process, and finally a low-dimensional latent feature vector reconstructed after noise removal is generated, and the reconstructed low-dimensional latent feature vector is mapped to the Mel spectrum features of the reconstructed natural sound data through the decoder in the VAE pre-trained model. Back propagation of errors is performed, and the parameters of the conditional controlled diffusion model are updated. Repeat the steps until the training is completed to obtain a trained conditional controlled diffusion model.
[0040] Step 114: Generate the corresponding natural sound using the HiFi-GAN pre-trained model according to the Mel-spectrogram features of the natural sound.
[0041] Specifically, after the natural sound is generated, the natural sound will affect the user's EEG data in the current time window, that is, the natural sound generated in the previous time window affects the EEG data in the current time window.
[0042] Step 116: Calculate the difference in the alpha wave amplitude values of the user's EEG data in adjacent time windows, determine the change trend of the alpha wave, and maintain the natural sound of the previous time window or generate a new natural sound through the creativity adjustment mechanism.
[0043] In the above-mentioned method for generating natural sounds based on personalized EEG signals, the method generates natural sounds based on the real-time collected user EEG data and scene space data, calculates the alpha wave amplitude value related to creativity, and judges the change trend of the alpha wave based on the alpha wave amplitude value of adjacent time windows, and dynamically generates natural sounds that are highly adapted to the personalized user EEG data and scene space data through the creativity adjustment mechanism, thereby personalizing the user's creativity.
[0044] In one embodiment, a CSNSP model is formed by combining a BERT pre-trained model, a LaBraM pre-trained model, a self-attention mechanism and an HTS-AT pre-trained model; the trained BERT pre-trained model, the LaBraM pre-trained model, the HTS-AT pre-trained model and the self-attention mechanism are obtained by training the CSNSP model; the training process of the CSNSP model includes: obtaining a multivariate scene data-natural sound data set; the multivariate scene data-natural sound data set is constructed according to a predefined correspondence between multivariate scene data and natural sound data; wherein the multivariate scene data includes: scene space data and user EEG data; the text corresponding to N scene space data is encoded through the BERT pre-trained model to obtain a feature vector of the text corresponding to the N scene space data; the scene in a time window is encoded The user computer data corresponding to the scene space data is encoded by the LaBraM pre-trained model to obtain the feature vector of the user's EEG data; the feature vector of the user's EEG data and the feature vectors of N scene space data are merged to obtain a fused feature vector; the fused feature vector is processed by the self-attention mechanism to obtain the feature vector of the multi-scene data; the corresponding natural sound data is encoded by the HTS-AT pre-trained model to obtain the feature vector of the natural sound data; the feature vector of the natural sound data and the feature vector of the multi-scene data are mapped to the latent space to obtain the natural sound data feature normalized vector and the multi-scene data feature normalized vector; according to the natural sound data feature normalized vector and the multi-scene data feature normalized vector, the CSNSP model is trained by contrastive learning to obtain a trained CSNSP model.
[0045] Specifically, the construction process of the multi-scenario data-natural sound dataset includes:
[0046] The natural sound data is processed by using short-time Fourier transform, Mel frequency scale and Mel filter bank to obtain the Mel spectrum features of the natural sound data;
[0047] Use the device in the scene where the user is located to collect scene space data and user EEG data. Through data cleaning and feature extraction, remove noise and invalid data, and determine the scene space data and user EEG data ;
[0048] The user's EEG data and scene space data are integrated to obtain multivariate scene data; then, the correspondence between the Mel-spectrum features of the multivariate scene data and the natural sound data is determined to construct a multivariate scene data-natural sound data set.
[0049] Preprocess the user's EEG data. The specific method is to filter to remove high-frequency noise and low-frequency drift in the EEG data, and then detect and remove artifacts in the EEG data through independent component analysis. Filter out the channels related to the user's creativity in the EEG data, and calculate the frequency domain data corresponding to the EEG data. The specific method is to perform short-time Fourier transform on the preprocessed user's EEG data to obtain its corresponding frequency domain data. Calculate the amplitude value of the alpha wave of a specific channel in the user's EEG data. The specific method is to filter out the alpha wave from the frequency domain data of the EEG data according to the frequency band of the alpha wave, and calculate the amplitude value of the alpha wave.
[0050] The specific steps of training the CSNSP model using the multi-scenario data-natural sound dataset include:
[0051] Step S1: Build the CSNSP model. The CSNSP model is a three-encoder architecture, which is used as the scene space data encoder. BERT pre-trained model as an EEG data encoder LaBraM pre-trained model and encoder for natural sound data HTS-AT pre-trained model, where Extract the feature vector of the text corresponding to the scene space data, Extract the feature vector of the user's EEG data, Extract feature vectors of natural sound data.
[0052] Given The text data corresponding to the scene space data (such as time data, weather data) is After encoding, indivual Dimensional vector , written as:
[0053] ;
[0054] in, Indicates scene space data.
[0055] Given a user's EEG data within a time window, After encoding, we get a Dimensional vector , written as:
[0056] ;
[0057] The feature vector of EEG data and The feature vector of the scene data Merge to get the feature vector of fused data , where the vector The shape is .
[0058] Using the Query Matrix in the Self-Attention Mechanism The feature vector Mapping to vector , written as:
[0059] ;
[0060] Using the Key matrix in the self-attention mechanism The feature vector Mapping to vector , written as:
[0061] ;
[0062] Use the Value matrix in the self-attention mechanism The feature vector Mapping to vector , written as:
[0063] ;
[0064] For each position in the feature vector , calculate the attention score , the attention score is used to measure the position and location The correlation between them is recorded as:
[0065] ;
[0066] in, yes No. vectors, yes No. vectors, , divided by This is to prevent the dot product result from being too large, causing the gradient of the softmax function to disappear.
[0067] Use the softmax function to calculate the attention score Normalize and get the attention weight , written as:
[0068] ;
[0069] By attention weight Perform weighted summation on the Value vector to get the output of the self-attention mechanism , written as:
[0070] ;
[0071] in, It is A Value vector.
[0072] Select vector The vector at the first position in is used as the feature vector of the multivariate scene data .
[0073] Given a natural sound data ,go through After encoding, we get a Dimensional vector , written as:
[0074] ;
[0075] Using the Multivariate Scene Transformation Matrix Vector Mapping to the latent space yields dimensional normalized vector , written as:
[0076] ;
[0077] Using the Natural Sound Conversion Matrix Vector Mapping to the latent space yields dimensional normalized vector , written as:
[0078] .
[0079] Step S2: training the contrast scene natural sound model. During the training process, the contrast scene natural sound model uses contrast learning to make the distance between matching multivariate scene data and natural sound data pairs in the vector space closer, while the distance between mismatching multivariate scene data and natural sound data pairs is farther.
[0080] Step S3: After the training is completed, the natural sound pre-training model of the comparison scene is obtained .
[0081] In one embodiment, the loss function used in the process of training the CSNSP model by contrastive learning includes: a loss function for retrieving natural sound data using multi-scenario data and a loss function for retrieving multi-scenario data using natural sound data.
[0082] Among them: The loss function expression of using multi-scene data to retrieve natural sound data is:
[0083] ;
[0084] in, Represents the similarity of the positive sample (i.e. Multi-scenario data and similarity between natural sound data), Indicates The sum of the similarities between multi-scenario data and all natural sound data, Represents the number of multivariate scene data.
[0085] The loss function expression for retrieving multi-scene data using natural sound data is:
[0086] ;
[0087] in, Represents the similarity of the positive sample (i.e. Natural sound data and similarity between multi-scenario data), Indicates The sum of the similarities between natural sound data and all multi-scenario data, Indicates the amount of natural sound data.
[0088] In one embodiment, the conditional controlled diffusion model in step 112 includes a VAE pre-trained model and a UNet+Attention model.
[0089] Specifically, step 1: determine the natural sound data . From multi-scene data - natural sound dataset Select natural sound data .
[0090] Step 2: Determine multi-scenario data . From multi-scene data - natural sound dataset Selecting multi-scenario data .
[0091] Step 3: Determine the natural sound data Mel spectrum features . Use short-time Fourier transform, Mel frequency scale and Mel filter bank to analyze natural sound data Processing to obtain the Mel spectrum features of natural sound .
[0092] Step 4: Build a conditional controlled diffusion model, including the VAE pre-training model and the UNet+Attention model.
[0093] Step 5: Train the conditional controlled diffusion model. Specifically, first use the encoder in the VAE pre-trained model Mel-spectrogram features of natural sound data Mapping to a low-dimensional latent vector , and in the forward diffusion process of the conditional controlled diffusion model, the latent vector Add to time steps of Gaussian noise.
[0094] The forward diffusion process is defined as a Markov chain process, denoted as:
[0095] ;
[0096] ;
[0097] in, , , is each time step The noise diffusion coefficient, represents the low-dimensional latent vector at the 0th time step, Indicates The low-dimensional latent vector after noise addition at each time step, is standard normal Gaussian noise.
[0098] After the forward diffusion process is completed, the reverse diffusion process is carried out. Here, the reverse diffusion process of the DDIM model is taken as an example.
[0099] The purpose of the reverse diffusion process is to Denoising is performed step by step, using the feature vectors of multivariate scene data To control the conditions, use the UNet+Attention model to predict the noise data , through which The model obtains high-dimensional feature representation of multi-scenario data .
[0100] The entire training process of the conditional controlled diffusion model uses the predicted noise Fitting at each time step Added Noise , the loss function of the conditional controlled diffusion model , written as:
[0101] ;
[0102] The final sampling formula for the reverse diffusion process is written as:
[0103] ;
[0104] in, ,variance It is a hyperparameter that can be adjusted manually.
[0105] In one embodiment, step 106 includes: fusing the feature vector of the text corresponding to the scene space data and the feature vector of the user's EEG data to obtain a fused feature vector; inputting the fused feature vector into the self-attention mechanism to generate a self-attention mechanism output; wherein the expression of the self-attention mechanism output is:
[0106] ;
[0107] ;
[0108] ;
[0109] ;
[0110] ;
[0111] ;
[0112] in, is the output of the self-attention mechanism, is the attention weight, represents the Query vector corresponding to the fused feature vector, Represents the Key vector corresponding to the fused feature vector, Represents the Value vector corresponding to the fused feature vector, yes No. vectors, yes No. vectors, yes No. vectors, is the Query matrix in the self-attention mechanism, is the Key matrix in the self-attention mechanism, is the Value matrix in the self-attention mechanism, is the dimension of the Query vector, is the dimension of the Key vector, is the dimension of the Value vector, is the amount of scene space data, is the fusion feature vector, Score for attention.
[0113] The vector at the first position in the output of the self-attention mechanism is used as the feature vector of the multivariate scene data.
[0114] In one embodiment, step 116 includes: calculating the average difference in alpha wave amplitude values of user EEG data in adjacent time windows; analyzing the change trend of the alpha wave amplitude in adjacent windows based on the average difference in the alpha wave amplitude values; when the change trend of the alpha wave amplitude in adjacent windows is an upward trend or remains unchanged, the natural sound in the next time window is consistent with the natural sound in the current time window; when the change trend of the alpha wave amplitude in adjacent windows is a downward trend, generating a new natural sound based on the scene space data and user EEG data at the current moment.
[0115] In one embodiment, the average difference of alpha wave amplitude values of user EEG data in adjacent time windows is calculated, including: if the current time window is the first time window, directly taking the amplitude value of the alpha wave of the EEG data in the time window as the average difference of the alpha wave amplitude values in adjacent time windows; if the current time window is not the first time window, calculating the difference between the alpha wave amplitude values of different channels in the previous time window and the current time window; adding the differences between the alpha wave amplitude values of different channels, and then dividing by the number of channels to obtain the average difference of the alpha wave amplitude values in adjacent time windows.
[0116] In one embodiment, the steps for calculating the alpha wave amplitude value of the user's EEG data include: collecting the user's EEG data within the current time window; filtering and performing independent component analysis on the user's EEG data within the current time window, detecting and removing artifacts in the user's EEG data, and then selecting channels related to the user's creativity from the processed user's EEG data; pre-emphasizing, framing, and windowing the user's EEG data, and then performing short-time Fourier transform to obtain frequency domain data corresponding to each channel of the EEG data; determining the frequency domain data corresponding to the alpha wave from the frequency domain data according to the frequency band of the alpha wave; calculating the amplitude values of different frequencies in the frequency domain data corresponding to the alpha wave of each channel after transforming the frequency domain data using the Euler formula; calculating the average of the alpha wave amplitude values of each channel as the amplitude value of the alpha wave of the current channel; and calculating the average of the alpha wave amplitude values of all channels as the amplitude value of the alpha wave of the current time window.
[0117] In a specific embodiment, the process of adjusting creativity based on natural sounds includes:
[0118] Step 1: Collect user EEG data within a time window , where the sampling frequency of EEG data is .
[0119] Step 2: User EEG data Filtering, using high-pass filter and low-pass filter to remove user EEG data The high-frequency noise and low-frequency drift in the user’s EEG data are then detected and removed through independent component analysis. Artifacts in the image, such as eye movement artifacts and electrocardiogram artifacts.
[0120] Step 3: Select user EEG data Channels related to user creativity.
[0121] Step 4: User EEG data After pre-emphasis, framing, and windowing, the short-time Fourier transform Get the frequency domain data corresponding to each channel of EEG data .
[0122] ;
[0123] in, Indicates Frame No. frequency components, Indicates the maximum value of frequency, Indicates Frame No. The signal value of the time domain sampling point, Represents the complex twiddle factors in Fourier transform.
[0124] Step 5: From frequency domain data According to the frequency band of alpha waves Determine the frequency domain data corresponding to the alpha wave .
[0125] Step 6: Calculate frequency domain data Different frequencies within Amplitude value .
[0126] First, the short-time Fourier transform After Euler's formula transformation, we get , written as:
[0127] ;
[0128] Then, according to Calculate frequency domain data Amplitude value , written as:
[0129] ;
[0130] in, express The real part of express The imaginary part of .
[0131] Step 7: Calculate frequency domain data Different frequencies within The average of the amplitude values is taken as the amplitude value of the alpha wave .
[0132] Step 8: Calculate the average difference of the alpha wave amplitude values in adjacent time windows.
[0133] First, the differences between the alpha wave amplitude values of different channels in adjacent time windows are calculated in turn; then, the differences between the alpha wave amplitude values of different channels are added and divided by the number of channels to obtain the average difference between the alpha wave amplitude values of adjacent time windows.
[0134] If the current time window is the first time window, the amplitude value of the alpha wave of the EEG data in the time window is directly used as the average difference of the alpha wave amplitude values of adjacent time windows.
[0135] Step nine: Analyze the changing trend of the alpha wave amplitude values in adjacent time windows based on the average difference analysis of the alpha wave amplitude values.
[0136] Step 10: Determine whether to generate a new natural sound based on the changing trend of the amplitude values of the EEG data in adjacent time windows.
[0137] If the change trend of the amplitude values of the EEG data in adjacent time windows shows an upward trend or remains unchanged, the natural sound in the next time window will be consistent with the natural sound in the current time window; if the change trend of the amplitude values of the EEG data in adjacent time windows shows a downward trend, a new natural sound will be generated based on the multivariate scene data at the current moment.
[0138] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0139] In one embodiment, Figure 3 As shown, a system for generating natural sounds based on personalized brain wave signals is provided, the system comprising: an information acquisition module, a data processing module, a personalized natural sound generation module and a creativity adjustment module based on natural sounds.
[0140] The information collection module is used to collect user EEG data and on-site scene space data using sensor devices, mobile devices, and audio devices; and obtain natural sound data from open source natural sound datasets.
[0141] Specifically, through the sensor devices in the acquired devices, personalized EEG data and scene space data are synchronously obtained according to the preset 256Hz acquisition frequency and 4s time window. Specifically as follows: EEG data acquisition: obtain EEG data monitoring equipment, the equipment uses a high-precision 32-bit electrode EEG machine, and continuously samples the EEG signals of the auxiliary object, and the sampling frequency is set to 256Hz. And the collected EEG signals are processed periodically. Scene space data acquisition: obtain meteorological monitoring equipment, obtain the time, light intensity, temperature, and humidity information of the location; obtain the GPS system, and obtain the geographical location information of the device, including longitude, latitude, and altitude. The specific method for obtaining natural sounds is to obtain natural sound data from open source natural sound datasets. The specific steps are as follows: comprehensively screen the existing open source natural sound datasets, and download the required natural sound data from the selected datasets.
[0142] Sensor equipment includes: EEG monitoring equipment, meteorological monitoring equipment and positioning system.
[0143] Among them: EEG monitoring equipment can be but is not limited to 32-electrode EEG, which is based on the mature electroencephalogram (EEG) technology principle, with 32 electrodes at specific positions on the scalp to achieve high-resolution collection of brain bioelectric activity. These electrodes are precisely placed according to the international standard 10-20 system.
[0144] Meteorological monitoring equipment includes: light monitoring sensors, temperature sensors, and humidity sensors.
[0145] Mobile devices include, but are not limited to, smartphones, smartwatches, tablets, and laptops.
[0146] Audio devices include, but are not limited to, mobile device speakers, external headphones, microphones, and speakers.
[0147] The data processing module is used to extract the feature vector of the text corresponding to the scene space data by using the BERT pre-training model; extract the feature vector of the user's EEG data by using the LaBraM pre-training model; fuse the feature vector of the text corresponding to the scene space data with the feature vector of the user's EEG data, and then process them using the self-attention mechanism to obtain the feature vector of the multi-scene data; extract the feature vector of the natural sound data by using the HTS-AT pre-training model; perform comparative learning on the feature vectors of the multi-scene data and the feature vectors of the natural sound data to obtain the high-dimensional feature representation of the multi-scene data.
[0148] The personalized natural sound generation module is used to generate the Mel spectrum features of natural sounds based on the high-dimensional feature representation of random Gaussian noise vectors and multivariate scene data, and use the conditional controlled diffusion model for processing; based on the Mel spectrum features of natural sounds, the HiFi-GAN pre-trained model is used to generate the corresponding natural sounds.
[0149] Specifically, a control interface installed in a mobile device is provided to ensure that users can control the playback and pause of sounds according to their actual needs.
[0150] Match the natural sound output device to ensure that the generated natural sound can be output to the user through the most appropriate device in the scene.
[0151] The creativity adjustment module based on natural sounds is used to calculate the difference in the alpha wave amplitude values of the user's EEG data in adjacent time windows, determine the change trend of the alpha wave, and through the creativity adjustment mechanism, maintain the natural sound of the previous time window or generate a new natural sound.
[0152] In one of the embodiments, the creativity adjustment module based on natural sounds is also used to calculate the average difference in alpha wave amplitude values of user EEG data in adjacent time windows; analyze the changing trend of the alpha wave amplitude in adjacent windows based on the average difference in the alpha wave amplitude values; when the changing trend of the alpha wave amplitude in adjacent windows is an upward trend or remains unchanged, the natural sound in the next time window is consistent with the natural sound in the current time window; when the changing trend of the alpha wave amplitude in adjacent windows is a downward trend, a new natural sound is generated based on the scene space data and user EEG data at the current moment.
[0153] In one of the embodiments, the BERT pre-trained model, LaBraM pre-trained model, self-attention mechanism and HTS-AT pre-trained model in the data processing module are combined into a CSNSP model; the trained BERT pre-trained model, LaBraM pre-trained model, self-attention mechanism and HTS-AT pre-trained model are obtained by training the CSNSP model; the training process of the CSNSP model includes: obtaining a multi-scenario data-natural sound data set; the multi-scenario data-natural sound data set is constructed according to a pre-defined correspondence between multi-scenario data and natural sound data; wherein the multi-scenario data includes: scene space data and user EEG data; N scene space data are encoded through the BERT pre-trained model to obtain feature vectors of text corresponding to the N scene space data; the scene in a time window is encoded into a scene space data set; the scene space data set ... The user computer data corresponding to the spatial data is encoded by the LaBraM pre-trained model to obtain the feature vector of the user's EEG data; the feature vector of the user's EEG data and the feature vector of the text corresponding to the N scene spatial data are merged to obtain a fused feature vector; the fused feature vector is processed by the self-attention mechanism to obtain the feature vector of the multi-scene data; the corresponding natural sound data is encoded by the HTS-AT pre-trained model to obtain the feature vector of the natural sound data; the feature vector of the natural sound data and the feature vector of the multi-scene data are mapped to the latent space to obtain the natural sound data feature normalized vector and the multi-scene data feature normalized vector; according to the natural sound data feature normalized vector and the multi-scene data feature normalized vector, the CSNSP model is trained by contrastive learning to obtain a trained CSNSP model.
[0154] In one of the embodiments, the loss function used in the data processing module in the process of training the CSNSP model by contrastive learning includes: a loss function for retrieving natural sound data using multivariate scene data and a loss function for retrieving multivariate scene data using natural sound data; wherein: the loss function for retrieving natural sound data using multivariate scene data is as shown in the above-mentioned loss function expression for retrieving natural sound data using multivariate scene data; the loss function for retrieving multivariate scene data using natural sound data is as shown in the above-mentioned loss function expression for retrieving multivariate scene data using natural sound data.
[0155] In one embodiment, the conditional controlled diffusion model in the personalized natural sound generation module includes a VAE pre-trained model and a UNet+Attention model.
[0156] In one of the embodiments, the data processing module is also used to fuse the feature vector of the text data and the feature vector of the user's EEG data to obtain a fused feature vector; input the fused feature vector into the self-attention mechanism to generate a self-attention mechanism output as shown in the expression of the self-attention mechanism output above; and use the vector at the first position in the self-attention mechanism output as the feature vector of the multivariate scene data.
[0157] In one of the embodiments, the creativity adjustment module based on natural sounds is also used to calculate the difference in alpha wave amplitude values of user EEG data in adjacent time windows, determine the changing trend of the alpha wave, and maintain the natural sound of the previous time window or generate a new natural sound through the creativity adjustment mechanism, including: calculating the average difference in alpha wave amplitude values of user EEG data in adjacent time windows; analyzing the changing trend of the alpha wave amplitudes in adjacent windows based on the average difference in the alpha wave amplitudes; when the changing trend of the alpha wave amplitudes in adjacent windows is an upward trend or remains unchanged, the natural sound in the next time window is consistent with the natural sound in the current time window; when the changing trend of the alpha wave amplitudes in adjacent windows is a downward trend, a new natural sound is generated based on the scene space data and user EEG data at the current moment.
[0158] In one embodiment, the average difference of alpha wave amplitude values of user EEG data in adjacent time windows is calculated, including: if the current time window is the first time window, directly taking the amplitude value of the alpha wave of the EEG data in the time window as the average difference of the alpha wave amplitude values in adjacent time windows; if the current time window is not the first time window, calculating the difference between the alpha wave amplitude values of different channels in the previous time window and the current time window; adding the differences between the alpha wave amplitude values of different channels, and then dividing by the number of channels to obtain the average difference of the alpha wave amplitude values in adjacent time windows.
[0159] In one of the embodiments, the steps for calculating the alpha wave amplitude value of the user's EEG data in the creativity adjustment module based on natural sounds include: collecting the user's EEG data within the current time window; filtering and performing independent component analysis on the user's EEG data within the current time window, detecting and removing artifacts in the user's EEG data, and then selecting channels related to the user's creativity from the processed user EEG data; pre-emphasizing, framing, and windowing the user's EEG data, and then performing short-time Fourier transform to obtain frequency domain data corresponding to each channel of the EEG data; determining the frequency domain data corresponding to the alpha wave from the frequency domain data according to the frequency band of the alpha wave; calculating the amplitude values of different frequencies in the frequency domain data corresponding to the alpha wave of each channel after transforming the frequency domain data through the Euler formula; calculating the average value of the alpha wave amplitude value of each channel as the amplitude value of the alpha wave of the current channel; calculating the average value of the alpha wave amplitude values of all channels as the amplitude value of the alpha wave of the current time window.
[0160] For the specific limitations of the system for generating natural sounds based on personalized EEG signals, please refer to the limitations of the method for generating natural sounds based on personalized EEG signals above, which will not be repeated here. The various modules in the above-mentioned system for generating natural sounds based on personalized EEG signals can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0161] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0162] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for generating natural sounds based on personalized EEG signals, characterized in that: The method comprises: Acquire multi-scenario data and natural sound data; the multi-scenario data includes scene space data and user EEG data; A BERT pre-trained model is used to extract a feature vector of the text corresponding to the scene space data; The LaBraM pre-trained model is used to extract the feature vector of the user's EEG data; The feature vector of the text corresponding to the scene space data and the feature vector of the user's EEG data are fused and processed using a self-attention mechanism to obtain a feature vector of the multivariate scene data; The HTS-AT pre-trained model is used to extract feature vectors of natural sound data; Performing comparative learning on the feature vectors of the multivariate scene data and the feature vectors of the natural sound data to obtain a high-dimensional feature representation of the multivariate scene data; Based on the high-dimensional feature representation of random Gaussian noise vectors and multivariate scene data, the conditional controlled diffusion model is used to generate the Mel spectrum features of natural sound; According to the Mel spectrum features of natural sounds, the HiFi-GAN pre-trained model is used to generate the corresponding natural sounds; Calculate the difference in the alpha wave amplitude values of the user's EEG data in adjacent time windows, determine the change trend of the alpha wave, and maintain the natural sound of the previous time window or generate a new natural sound through the creativity adjustment mechanism; The steps for calculating the alpha wave amplitude value of the user's EEG data include: Collect the user's EEG data within the current time window; Filtering and performing independent component analysis on the user's EEG data within the current time window, detecting and removing artifacts in the user's EEG data, and then selecting channels related to the user's creativity from the processed user's EEG data; After pre-emphasis, framing and windowing of the user's EEG data, the frequency domain data corresponding to each channel of the EEG data is obtained through short-time Fourier transform; Determine frequency domain data corresponding to the alpha wave from the frequency domain data according to the frequency band of the alpha wave; After the frequency domain data is transformed by the Euler formula, the amplitude values of different frequencies in the frequency domain data corresponding to each channel alpha wave are calculated; the average value of the amplitude value of each channel alpha wave is calculated as the amplitude value of the current channel alpha wave; the average value of the amplitude value of all channels alpha waves is calculated as the amplitude value of the alpha wave in the current time window.
2. The method for generating natural sounds based on personalized EEG signals according to claim 1, characterized in that: The CSNSP model is composed of the BERT pre-training model, the LaBraM pre-training model, the self-attention mechanism, and the HTS-AT pre-training model; The training process of the CSNSP model includes: Acquire a multivariate scene data-natural sound data set; the multivariate scene data-natural sound data set is constructed according to a predefined correspondence between multivariate scene data and natural sound data; wherein the multivariate scene data includes: scene space data and user EEG data; Encode the text corresponding to the N scene space data through the BERT pre-trained model to obtain feature vectors of the N scene space data; The user's computer data corresponding to the scene space data in a time window is encoded through the LaBraM pre-trained model to obtain the feature vector of the user's EEG data; Merge the feature vector of the user's EEG data and the feature vector of the text corresponding to the N scene space data to obtain a fused feature vector; The fused feature vector is processed by a self-attention mechanism to obtain a feature vector of the multivariate scene data; the corresponding natural sound data is encoded by an HTS-AT pre-trained model to obtain a feature vector of the natural sound data; Mapping the feature vector of the natural sound data and the feature vector of the multivariate scene data to a latent space to obtain a normalized feature vector of the natural sound data and a normalized feature vector of the multivariate scene data; According to the normalized feature vectors of natural sound data and the normalized feature vectors of multivariate scene data, the CSNSP model is trained by contrastive learning to obtain a trained CSNSP model.
3. The method for generating natural sounds based on personalized EEG signals according to claim 2, characterized in that: The loss functions used in the process of training the CSNSP model by contrastive learning include: a loss function for retrieving natural sound data using multi-scenario data and a loss function for retrieving multi-scenario data using natural sound data; Among them: The loss function of retrieving natural sound data using multi-scene data is: in, Represents the similarity of positive samples, Indicates The sum of the similarities between multi-scenario data and all natural sound data, Indicates the number of multivariate scene data; The loss function for retrieving multi-scene data using natural sound data is: in, Represents the similarity of positive samples, Indicates The sum of the similarities between natural sound data and all multi-scenario data, Indicates the amount of natural sound data.
4. The method for generating natural sounds based on personalized EEG signals according to claim 1, characterized in that: The conditional controlled diffusion model includes a VAE pre-training model and a UNet+Attention model.
5. The method for generating natural sounds based on personalized EEG signals according to claim 1, characterized in that: The feature vector of the text corresponding to the scene space data and the feature vector of the user's EEG data are fused and processed using a self-attention mechanism to obtain a feature vector of the multivariate scene data, including: The feature vector of the text corresponding to the scene space data and the feature vector of the user's EEG data are fused to obtain a fused feature vector; the fused feature vector is input into the self-attention mechanism to generate a feature vector of the multivariate scene data: in, is the output of the self-attention mechanism, is the attention weight, represents the Query vector corresponding to the fused feature vector, Represents the Key vector corresponding to the fused feature vector, Represents the Value vector corresponding to the fused feature vector, yes No. vectors, yes No. vectors, yes No. vectors, is the Query matrix in the self-attention mechanism, is the Key matrix in the self-attention mechanism, is the Value matrix in the self-attention mechanism, is the dimension of the Query vector, is the dimension of the Key vector, is the dimension of the Value vector, is the amount of scene space data, is the fusion feature vector, Score for attention; The vector at the first position in the output of the self-attention mechanism is used as the feature vector of the multivariate scene data.
6. The method for generating natural sounds based on personalized EEG signals according to claim 1, characterized in that: Calculate the difference in the alpha wave amplitude values of the user's EEG data in adjacent time windows, determine the change trend of the alpha wave, and maintain the natural sound of the previous time window or generate new natural sounds through the creativity adjustment mechanism, including: singing Calculate the average difference of the alpha wave amplitude values of the user's EEG data in adjacent time windows; The changing trend of the alpha wave amplitudes in adjacent windows was analyzed based on the average difference of the alpha wave amplitudes; When the change trend of the alpha wave amplitude in adjacent windows is increasing or remains unchanged, the natural sound in the next time window is consistent with the natural sound in the current time window; When the change trend of the alpha wave amplitude of the adjacent windows shows a downward trend, a new natural sound is generated based on the scene space data and the user's EEG data at the current moment.
7. The method for generating natural sounds based on personalized EEG signals according to claim 6, characterized in that: Calculate the average difference of the alpha wave amplitude values of the user's EEG data in adjacent time windows, including: If the current time window is the first time window, the amplitude value of the alpha wave of the EEG data in the time window is directly used as the average difference of the alpha wave amplitude values of adjacent time windows; If the current time window is not the first time window, the difference between the alpha wave amplitude values of different channels in the previous time window and the current time window is calculated; the differences between the alpha wave amplitude values of different channels are added and then divided by the number of channels to obtain the average difference of the alpha wave amplitude values of adjacent time windows.
8. A system for generating natural sounds based on personalized EEG signals, characterized in that: The system includes: an information collection module, a data processing module, a personalized natural sound generation module, and a creativity adjustment module based on natural sounds; The information collection module is used to collect user EEG data and on-site scene space data using sensor devices, mobile devices and audio devices; and obtain natural sound data from an open source natural sound data set; The data processing module is used to extract the feature vector of the text corresponding to the scene space data by using the BERT pre-training model; extract the feature vector of the user's EEG data by using the LaBraM pre-training model; fuse the feature vector of the text corresponding to the scene space data with the feature vector of the user's EEG data and process them by using the self-attention mechanism to obtain the feature vector of the multivariate scene data; extract the feature vector of the natural sound data by using the HTS-AT pre-training model; perform comparative learning on the feature vector of the multivariate scene data and the feature vector of the natural sound data to obtain the high-dimensional feature representation of the multivariate scene data; The personalized natural sound generation module is used to generate Mel spectrum features of natural sound based on the high-dimensional feature representation of random Gaussian noise vector and multivariate scene data by using a conditional controlled diffusion model; and generate corresponding natural sound based on the Mel spectrum features of natural sound by using a HiFi-GAN pre-trained model; The creativity adjustment module based on natural sound is used to calculate the difference in alpha wave amplitude values of user EEG data in adjacent time windows, determine the change trend of alpha waves, and maintain the natural sound of the previous time window or generate new natural sounds through the creativity adjustment mechanism; the calculation steps of the alpha wave amplitude value of user EEG data include: collecting user EEG data in the current time window; filtering and independent component analysis of the user EEG data in the current time window, detecting and removing artifacts in the user EEG data, and then selecting the user creativity-related sound from the processed user EEG data; channel; after pre-emphasis, framing and windowing processing of the user's EEG data, the frequency domain data corresponding to each channel of the EEG data is obtained through short-time Fourier transform; the frequency domain data corresponding to the alpha wave is determined from the frequency domain data according to the frequency band of the alpha wave; after the frequency domain data is transformed through Euler formula, the amplitude values of different frequencies in the frequency domain data corresponding to the alpha wave of each channel are calculated; the average value of the amplitude value of the alpha wave of each channel is calculated as the amplitude value of the alpha wave of the current channel; the average value of the amplitude values of the alpha waves of all channels is calculated as the amplitude value of the alpha wave of the current time window.
9. The system for generating natural sounds based on personalized EEG signals according to claim 8, characterized in that: The creativity adjustment module based on natural sounds is also used to calculate the average difference in alpha wave amplitude values of user EEG data in adjacent time windows; analyze the change trend of the alpha wave amplitude in adjacent windows according to the average difference in the alpha wave amplitude values; when the change trend of the alpha wave amplitude in adjacent windows is an upward trend or remains unchanged, the natural sound in the next time window is consistent with the natural sound in the current time window; when the change trend of the alpha wave amplitude in adjacent windows is a downward trend, a new natural sound is generated according to the scene space data and user EEG data at the current moment.
Citation Information
Patent Citations
Multifunctional brain wave intelligent device and multifunctional brain wave earphone
CN112969117A
Neural regulation and control system based on alpha wave closed-loop flour noise stimulation
CN119327038A