Intelligent music regulation and control system and method based on electroencephalogram signal multi-modal feature recognition
By performing anti-interference acquisition and adaptive filtering of EEG signals, combined with multi-dimensional feature extraction and cross-modal information fusion, a personalized music regulation system based on EEG signals is realized, solving the problems of inaccurate brain simulation, poor adaptability and insufficient safety in the existing technology, and improving the accuracy and safety of the regulatory system.
Patent Information
- Application Number
- CN202510953829.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing music regulation system based on biological models fails to fully reflect the complex connection structure and dynamic characteristics of the real brain, the ability to integrate cross-modal information is limited, the adaptive regulation ability is lacking, the neural computing model is insufficiently interpreted, and the data security protection mechanism is lacking, resulting in large individual differences in the treatment effect.
By performing anti-interference acquisition and adaptive filtering on EEG signals, multi-dimensional features are extracted and cross-modal information fusion is performed, personalized music regulation is carried out in combination with music-emotional mapping network, a closed-loop feedback mechanism is established and data security hierarchical protection is carried out.
It realizes accurate identification and personalized regulation of user psychological state, improves regulation efficiency and safety, adapts to changes in user needs, provides immediate emotional regulation and long-term emotional management, and enhances trust in use.
Smart Images

Figure CN120437460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological computing technology, and more specifically, to an intelligent music control system and method based on multimodal feature recognition of electroencephalogram (EEG) signals. Background Art
[0002] With the deep integration of neuroscience and computer science, computer systems based on biological models, particularly neural network models, have made significant progress in simulating human cognitive functions. Music, as a non-invasive stimulation method, has a significant regulatory effect on human emotions, cognition, and physiological states, and has been widely used in clinical psychotherapy, cognitive enhancement, mood management, and stress relief. With the development of brain science and artificial intelligence technologies, the application of neural computational models to music control systems has gradually become a research hotspot.
[0003] Traditional music therapy relies primarily on the therapist's professional experience and subjective judgment, lacking a deep understanding and simulation of the human brain's neural mechanisms. For example, in clinical treatment, therapists often select specific musical pieces based on experience, rather than on a real-time assessment of the patient's brain neural network state. This approach, divorced from a foundation in brain science, struggles to precisely regulate the nervous system, leading to significant individual variability in treatment outcomes.
[0004] In recent years, with the development of electroencephalography (EEG) technology and neural network models, researchers have begun to explore the construction of music control systems based on biological models. These systems typically consist of three key modules: an EEG signal processing model based on the principles of the human brain's neural network, used to extract and analyze neural activity characteristics; a neural computational model that simulates the human brain's emotional processing mechanisms, used to identify emotional states; and a music-emotion mapping network based on the response characteristics of the auditory nervous system, used to achieve personalized music control.
[0005] However, existing music control systems based on biological models still have many limitations. First, the simulation of the human brain's neural network is overly simplified, failing to fully reflect the complex connectivity structure and dynamic characteristics of the real brain. Second, the cross-modal information integration capability is limited, making it difficult to simulate the human brain's mechanism for integrating multisensory information. Third, the simulation of neural plasticity is insufficient, lacking the adaptive adjustment capabilities similar to the human brain's long-term memory and learning mechanisms. In addition, the neural computational model is insufficiently interpretable, making it difficult to provide understandable decision-making basis similar to that of human experts. Finally, insufficient consideration is given to the safety of the nervous system, lacking a data security protection mechanism similar to the blood-brain barrier.
[0006] In actual application scenarios, there is an urgent need for an intelligent system that can more accurately simulate the working mechanism of the human brain and achieve precise regulation at the neural level.
[0007] In view of this, the present invention proposes an intelligent music control system and method based on multimodal feature recognition of EEG signals to solve the above problems. Summary of the Invention
[0008] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides an intelligent music control system and method based on multimodal feature recognition of EEG signals, which provides innovative technical support for the fields of emotion management, cognitive enhancement, psychotherapy, etc.
[0009] In a first aspect, the present application provides an intelligent music control method based on multimodal feature recognition of EEG signals, comprising: Step 1: Anti-interference acquisition and adaptive filtering of EEG signals are performed to obtain a high-quality EEG basic data set; Step 2: performing multi-dimensional feature extraction and cross-modal information fusion based on the high-quality EEG basic data set to obtain a multimodal feature vector; Step 3: performing emotional state recognition and cognitive load assessment based on the multimodal feature vector to obtain a representation of the user's current psychological state; Step 4: Determine the target mental state through user input or statistical analysis based on user historical data, and perform music feature screening based on the difference between the target mental state and the user's current mental state representation in combination with a pre-trained music-emotion mapping network to obtain a personalized music control strategy; Step 5: Perform real-time music parameter adjustment and audio stream generation based on the personalized music control strategy to obtain an adaptive music output stream; Step 6: Collecting EEG feedback and evaluating the adjustment effect based on the adaptive music output stream to obtain feedback data and use it to optimize the music-emotion mapping network; Step 7: Perform data security classification and privacy protection processing on the high-quality EEG basic data set, the multimodal feature vector, and the feedback data, and store them in a secure storage database; Step 8: Based on the feedback data in the secure storage database, continuously evaluate the effect of the adaptive music output stream to obtain an effect evaluation report.
[0010] In a second aspect, the present application provides an intelligent music control system based on multimodal feature recognition of EEG signals, comprising: The EEG signal acquisition and processing module is used to perform anti-interference acquisition and adaptive filtering on EEG signals to obtain high-quality EEG basic data sets; Multimodal feature fusion module, which is used to extract multi-dimensional features and fuse cross-modal information based on high-quality EEG basic data sets to obtain multimodal feature vectors; The psychological state recognition module is used to perform emotional state recognition and cognitive load assessment based on multimodal feature vectors to obtain the user's current psychological state representation; A music-emotion mapping module is used to determine a target mental state through user input or statistical analysis based on historical user data. Based on the difference between the target mental state and the user's current mental state representation, music features are screened in combination with a pre-trained music-emotion mapping network to obtain a personalized music control strategy. A personalized control strategy module is used to adjust music parameters and generate audio streams in real time based on the personalized music control strategy to obtain an adaptive music output stream; A music reconstruction output module is used to collect EEG feedback and evaluate the effect of regulation based on the adaptive music output stream, obtain feedback data and use it to optimize the music-emotion mapping network; The data security protection module is used to perform data security classification and privacy protection processing on high-quality EEG basic data sets, multimodal feature vectors and feedback data, and store them in a secure storage database; A long-term tracking and evaluation module, configured to continuously evaluate the effect of the adaptive music output stream based on the feedback data in the secure storage database and obtain an effect evaluation report; The modules are connected via wired and / or wireless means to achieve data transmission between modules.
[0011] The technical effects and advantages of the intelligent music control system and method based on multimodal feature recognition of EEG signals of the present invention are as follows: The present invention improves the accuracy, personalization and safety of music regulation, and completely solves the pain points of traditional music therapy, which has unstable effects and lacks scientific basis. By establishing a precise mapping relationship between EEG signals and emotional states, real-time and accurate grasp of the user's psychological state is achieved, which enables music regulation to shift from empirical to scientific, greatly improving the efficiency of regulation. At the same time, the closed-loop feedback mechanism ensures the continuous optimization of the regulation effect, enabling the system to continuously self-learn and evolve to adapt to the changing needs of different users. This personalized customization capability greatly expands the application scenarios, and can provide targeted support from clinical psychotherapy to daily emotional management, from improving learning efficiency to improving sleep quality. In addition, the complete data security and privacy protection system provides users with reliable protection and enhances their trust in use. The present invention can not only provide users with instant emotional regulation, but also help users establish a healthier emotional management model and improve psychological resilience through long-term data accumulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 Schematic diagram of the intelligent music control method based on multimodal feature recognition of EEG signals of the present invention; Figure 2Schematic diagram of the intelligent music control system based on multimodal feature recognition of EEG signals of the present invention. DETAILED DESCRIPTION
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0014] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0015] This application example provides an intelligent music control system and method based on multimodal feature recognition of EEG signals. The execution subjects of the intelligent music control system and method based on multimodal feature recognition of EEG signals include but are not limited to: brain-computer interface devices, signal processing platforms, music generation systems, cloud server nodes, etc. equipped with the system, which can be regarded as general computing nodes of this application. The signal processing platform includes but is not limited to: at least one of an EEG signal processing system, an emotion recognition system, and a music control generation system.
[0016] See also Figure 1 The present invention provides an intelligent music control method based on multimodal feature recognition of EEG signals, comprising: Step 1: Anti-interference acquisition and adaptive filtering of EEG signals are performed to obtain a high-quality EEG basic data set; Step 2: performing multi-dimensional feature extraction and cross-modal information fusion based on the high-quality EEG basic data set to obtain a multimodal feature vector; Step 3: performing emotional state recognition and cognitive load assessment based on the multimodal feature vector to obtain a representation of the user's current psychological state; Step 4: Determine the target mental state through user input or statistical analysis based on user historical data, and perform music feature screening based on the difference between the target mental state and the user's current mental state representation in combination with a pre-trained music-emotion mapping network to obtain a personalized music control strategy; Step 5: Perform real-time music parameter adjustment and audio stream generation based on the personalized music control strategy to obtain an adaptive music output stream; Step 6: Collecting EEG feedback and evaluating the adjustment effect based on the adaptive music output stream to obtain feedback data and use it to optimize the music-emotion mapping network; Step 7: Perform data security classification and privacy protection processing on the high-quality EEG basic data set, the multimodal feature vector, and the feedback data, and store them in a secure storage database; Step 8: Based on the feedback data in the secure storage database, continuously evaluate the effect of the adaptive music output stream to obtain an effect evaluation report.
[0017] The present invention obtains a high-quality EEG basic data set by performing anti-interference collection and adaptive filtering processing on EEG signals, generates a multimodal feature vector by combining multi-dimensional feature extraction and cross-modal information fusion, performs emotional state recognition and cognitive load assessment based on the multimodal feature vector to obtain a representation of the user's current psychological state, determines the target psychological state through a pre-trained music-emotion mapping network and the user's current psychological state representation, performs music feature screening to generate a personalized music regulation strategy, performs real-time music parameter adjustment and audio stream generation to obtain an adaptive music output stream, utilizes EEG feedback collection and regulation effect evaluation to obtain feedback data and optimize the music-emotion mapping network, establishes a secure storage database through data security classification and privacy protection processing, and finally realizes continuous evaluation of the effect of the adaptive music output stream, generates an effect evaluation report, and provides a scientific basis for personalized music emotion regulation.
[0018] Step 1: Anti-interference acquisition and adaptive filtering of EEG signals are performed to obtain a high-quality EEG basic data set; In this embodiment, the spatial distribution of multi-lead EEG electrodes is optimized to obtain a full-brain coverage acquisition scheme, and active shielding technology is applied to the full-brain coverage acquisition scheme to obtain an anti-interference acquisition circuit; the original EEG data output by the anti-interference acquisition circuit is sampled and quantized in real time to obtain a digitized EEG signal, and adaptive notch filtering is applied to the digitized EEG signal to obtain a power frequency interference-free signal; the power frequency interference-free signal is subjected to wavelet decomposition and reconstruction to obtain a denoised EEG signal; independent component analysis is performed on the denoised EEG signal to obtain an independent neural source signal; artifact detection and removal are performed on the independent neural source signal to obtain a pure EEG signal; time-frequency analysis is performed on the pure EEG signal to obtain enhanced EEG features; data standardization is performed on the enhanced EEG features to obtain standardized EEG data, and the standardized EEG data is stored as a structured data set to obtain a high-quality EEG basic data set.
[0019] In this embodiment, the electrode layout is designed according to the international 10-20 system, covering key brain areas such as the frontal lobe, parietal lobe, temporal lobe, and occipital lobe to ensure that complete EEG activity information is collected; wet electrode technology is used to optimize the electrode-skin contact interface, reduce contact impedance, and improve signal quality; an active shielding electrode system is designed, and a shielding layer is set around the electrode and connected to the drive circuit to form an electric field isolation structure to suppress environmental electromagnetic interference; for mobile scenarios, a fixed headband technology is used to ensure the stability of the electrode position; an analog front-end circuit including preamplification and filtering is constructed to achieve early signal conditioning and noise suppression; a 24-bit A high-precision analog-to-digital converter collects EEG data at a sampling rate of 1000Hz to ensure the capture of high-frequency neural oscillation components; an adaptive notch filtering algorithm is applied to automatically adjust the filtering parameters based on real-time spectrum analysis to accurately remove 50 / 60Hz power frequency interference and its harmonics, while retaining the effective signals in adjacent frequency bands to the greatest extent possible; a Butterworth bandpass filter is used for fundamental frequency band limitation, with a set passband of 0.1-100Hz to remove baseline drift and high-frequency electromyographic interference; zero-phase filtering technology is used in the filtering process to avoid phase distortion introduced by filtering and ensure the accurate retention of the signal's temporal characteristics.
[0020] Perform wavelet decomposition and reconstruction on the power frequency interference-free signal to obtain the denoised EEG signal: In this embodiment, a wavelet basis function suitable for the characteristics of the EEG signal is selected to obtain an optimized wavelet transform operator; the optimized wavelet transform operator is applied to perform multi-level decomposition on the power frequency interference-free signal to obtain a set of wavelet coefficients in different frequency ranges; the wavelet coefficient set is subjected to threshold denoising to obtain denoised wavelet coefficients; the denoised wavelet coefficients are subjected to inverse wavelet transform to obtain a denoised EEG signal.
[0021] In this embodiment, by comparing the characteristics of Daubechies and Symlet wavelet basis functions and evaluating their ability to retain the transient characteristics of EEG signals and computational complexity, the Daubechies wavelet family is selected as the optimized wavelet transform operator; for the selected wavelet family, the number of decomposition levels is optimized to 8 levels based on the signal-to-noise ratio (SNR) and mean square error (MSE) indicators; the optimized wavelet transform operator is applied to perform multi-level decomposition on the power frequency interference-free signal, and the approximate coefficients and detail coefficients corresponding to different frequency ranges are extracted to form a set of wavelet coefficients; the soft threshold method is applied to the wavelet coefficient set for denoising, and the threshold is determined by the Donoho-Johnstone method to generate denoised wavelet coefficients; the denoised wavelet coefficients are inversely transformed to obtain a denoised EEG signal.
[0022] Perform artifact detection and removal on independent neural source signals to obtain pure EEG signals: In this embodiment, neural source signals of training samples containing eye movement, heartbeat and electromyographic artifact features are obtained and labeled to obtain artifact labeled signal samples; features of the artifact labeled signal samples are extracted to obtain an artifact feature template set; based on the artifact feature template set, an artifact recognition classifier is trained using a support vector machine algorithm to obtain an artifact detection model; features of the independent neural source signals are extracted to make them consistent with the feature format of the artifact feature template set to obtain neural source features to be detected; the neural source features to be detected are classified into components using the artifact detection model to obtain artifact component labeling results; based on the artifact component labeling results, the artifact energy ratio of each independent neural source signal is calculated to obtain an artifact impact score; based on the artifact impact score, an artifact energy ratio threshold θ is set, and the artifact energy ratio threshold θ is calculated by the formula θ=θ_bas e×(1+α×S), where θ_base is the basic threshold, α is the adjustment coefficient, and S is the artifact influence score; the artifact components in the independent neural source signal are filtered according to the artifact energy ratio threshold θ to obtain a preliminary purified EEG signal; principal component analysis is applied to the preliminary purified EEG signal to obtain the main neural source signal; signal integrity evaluation is performed on the main neural source signal, and the signal-to-noise ratio score and the spectrum energy distribution balance score are calculated by weighted average to obtain a data quality score, in which the signal-to-noise ratio score accounts for 60% and the spectrum energy distribution balance score accounts for 40%; based on the data quality score, if the score is lower than the preset quality score threshold, the main neural source signal is interpolated and compensated to obtain a pure EEG signal; if the score is not lower than the preset quality score threshold, the main neural source signal is directly used as a pure EEG signal.
[0023] In this embodiment, typical samples of physiological artifacts such as electrooculogram (VEOG / HEOG), electrocardiogram (ECG) and electromyography (EMG) are collected to construct artifact annotation signal samples; time domain features (mean, variance) and frequency domain features (power spectral density) are extracted for each type of artifact to form an artifact feature template set; a support vector machine algorithm is used to train an artifact recognition classifier to learn the feature patterns of different artifact types and to construct an artifact detection model; features that are consistent with the format of the artifact feature template set are extracted from the independent neural source signals obtained by independent component analysis (ICA) separation to obtain the neural source features to be detected; the artifact detection model is used to evaluate the neural source features to be detected component by component, identify suspected artifact components, and generate artifact component labeling results; based on the labeling results, the artifact energy proportion of each independent component is calculated, the formula is S=E_pseudo / E_total, where E_pseudo is the artifact energy and E_total is the total energy, to obtain the artifact impact score S; and a threshold for the artifact energy proportion is set. θ, the base threshold θbase=0.3, the adjustment coefficient α=0.5, and the adaptive threshold is calculated using the formula θ=θ_base×(1+α×S). Based on the artifact energy ratio threshold θ, independent components with artifact energy ratios exceeding the threshold are set to zero to obtain a preliminary purified EEG signal. Principal component analysis is applied to the preliminary purified EEG signal, retaining the top 10 principal components with the largest amount of information to obtain the main neural source signal. The signal-to-noise ratio (SNR) and spectral energy distribution balance of the main neural source signal are calculated. The spectral energy distribution balance score is calculated by the standard deviation of the power ratio of each frequency band (δ, θ, α, β, γ). The smaller the standard deviation, the higher the balance score. The data quality score is calculated by weighted average, ranging from 0 to 100. If the data quality score is lower than the preset quality score threshold of 80, the missing signal is compensated using linear interpolation to obtain a purified EEG signal. If the data quality score is not lower than the preset quality score threshold of 80, the main neural source signal is directly used as the purified EEG signal.
[0024] Perform time-frequency analysis on the pure EEG signal to obtain enhanced EEG features: In this embodiment, a short-time Fourier transform (STFT) is applied to the pure EEG signal, with a window length of 1 second and a step size of 0.5 seconds, to obtain a time-spectrogram. The power features of the δ (0.5-4 Hz), θ (4-8 Hz), α (8-13 Hz), β (13-30 Hz), and γ (>30 Hz) frequency bands are extracted from the time-spectrogram to obtain enhanced EEG features.
[0025] Normalize the enhanced EEG features to obtain standardized EEG data: In this embodiment, the Z-score normalization method is applied to the enhanced EEG features to obtain standardized EEG data; the standardized EEG data is organized into a structured data set according to timestamps and electrode positions, and stored in CSV format to obtain a high-quality EEG basic data set.
[0026] Step 2: Perform multi-dimensional feature extraction and cross-modal information fusion based on high-quality EEG basic data sets to obtain a multimodal feature vector; In this embodiment, a time domain analysis is performed on a high-quality EEG basic data set to extract mean, variance and peak features to obtain a statistical feature set; a frequency domain analysis is performed on a high-quality EEG basic data set to extract power features of the δ, θ, α, β and γ frequency bands to obtain a spectral feature set; a nonlinear analysis is performed on a high-quality EEG basic data set to calculate sample entropy and fractal dimension to obtain a nonlinear feature set; the Pearson correlation coefficient of different brain regions is calculated for the high-quality EEG basic data set to obtain a functional connectivity matrix; graph theory analysis is applied to the functional connectivity matrix to calculate node degree and clustering coefficient to obtain a network feature set; a random forest model is trained based on pre-labeled training data to obtain a feature importance evaluation model; and the feature importance evaluation model is applied. The evaluation model calculates the importance score of each feature in the statistical feature set, the spectral feature set, the nonlinear feature set and the network feature set to obtain an important feature set; applies the principal component analysis algorithm to perform dimensionality reduction processing on the important feature set to obtain a feature vector after dimensionality reduction; collects auxiliary physiological signals, which include heart rate variability and galvanic skin response signals, and extracts features from the auxiliary physiological signals to obtain an auxiliary physiological feature vector; time-aligns the feature vector after dimensionality reduction with the auxiliary physiological feature vector, and synchronizes data through a linear interpolation method to obtain synchronized feature data; applies a weighted fusion method to the synchronized feature data, and the fusion weight is calculated through the Pearson correlation coefficient to obtain a multimodal feature vector.
[0027] Perform time domain analysis on high-quality EEG basic data sets to obtain statistical feature sets: In this embodiment, the mean, variance and peak characteristics of each signal segment are calculated for the high-quality EEG basic data set; the mean is calculated by the formula Calculate, where X_i is the i-th signal sampling point and N is the number of sampling points; the peak value is obtained by finding the maximum and minimum values of each signal segment to generate a statistical feature set.
[0028] Perform frequency domain analysis on high-quality EEG basic data sets to obtain the spectrum feature set: In this embodiment, a fast Fourier transform (FFT) is applied to a high-quality EEG basic data set to convert the time domain signal into the frequency domain and obtain the power spectrum density of the signal. The absolute power of the δ (0.5-4 Hz), θ (4-8 Hz), α (8-13 Hz), β (13-30 Hz), and γ (>30 Hz) frequency bands is extracted from the power spectrum density. The calculation method is: , where P(f) is the power spectral density at frequency f, and band is the corresponding frequency band range; a spectrum feature set is formed.
[0029] Perform nonlinear analysis on high-quality EEG basic data sets to obtain nonlinear feature sets: In this embodiment, the sample entropy and fractal dimension are calculated for a high-quality EEG basic data set; the sample entropy is calculated by the formula Calculation, where m is the embedding dimension (taken as 2), r is the similarity threshold (taken as 0.2×standard deviation), A is the number of matches of a template vector of length m+1, B is the number of matches of a template vector of length m, and N is the signal length; the fractal dimension is calculated using the Higuchi algorithm to generate a nonlinear feature set.
[0030] The Pearson correlation coefficients of different brain regions were calculated for high-quality EEG basic data sets to obtain the functional connectivity matrix: In this embodiment, the Pearson correlation coefficient of the EEG signal time series between different brain regions is calculated to obtain a functional connectivity strength matrix; a fixed threshold is applied to the functional connectivity strength matrix to retain connections with a correlation coefficient greater than 0.6 to obtain a significant connectivity matrix; graph theory analysis is applied to the significant connectivity matrix to calculate the node degree of each brain region, and brain regions with node degrees higher than the average node degree of the whole brain are defined as key brain regions to obtain the distribution of key brain regions; modular analysis is performed on the significant connectivity matrix, and the Louvain algorithm is used to divide the functional modules to obtain a module division result; the average connection strength between the modules in the module division result is calculated to obtain the inter-module interaction strength; the key brain region distribution and the inter-module interaction strength are integrated to obtain a network feature set.
[0031] In this embodiment, the Pearson correlation coefficient of the EEG signal time series between different brain regions is calculated to form a functional connectivity strength matrix; a fixed threshold of 0.6 is applied to the functional connectivity strength matrix to retain significant connections, thereby obtaining a significant connectivity matrix; the significant connectivity matrix is represented as a graph structure, with nodes representing different brain regions and edges representing functional connections between brain regions; the degree of each node is calculated using the formula: , where a_ij is the element of the significant connection matrix, representing the connection strength between node i and node j. The average node degree of the whole brain is calculated as follows: , where n is the total number of nodes. The brain regions with node degree d_i higher than d_avg are defined as key brain regions, and the distribution of key brain regions is obtained. The Louvain algorithm is used to perform modular analysis on the significant connection matrix, divide the functional modules, and obtain the module division results. The average connection strength between each module is calculated using the formula: , where N_inter is the number of inter-module connections and a_ij is the inter-module connection strength, and the inter-module interaction strength is obtained.
[0032] Perform feature importance evaluation and dimensionality reduction: In this embodiment, the statistical feature set, spectral feature set, nonlinear feature set and network feature set are combined into a complete feature set; training samples containing known emotional states and cognitive loads are collected and labeled, and a random forest model is trained based on these labeled samples. The number of trees is set to 100, and the number of features is randomly selected. , where n is the total number of features, and the feature importance evaluation model is obtained; the trained feature importance evaluation model is used to calculate the Gini index reduction value of each feature in the complete feature set, and the formula is , where p_k is the probability of a feature in a random forest node split, and the importance score is obtained; based on the importance score, the features with the top 50% scores are selected to obtain the important feature set; the principal component analysis (PCA) algorithm is applied to the important feature set, and the principal components with a cumulative variance contribution rate of 90% are retained to obtain the feature vector after dimensionality reduction.
[0033] Collect auxiliary physiological signals and perform feature extraction: In this embodiment, the heart rate variability (HRV) signal is collected, and the time domain feature SDNN (standard deviation) and the frequency domain feature LF / HF (low frequency / high frequency power ratio) are extracted to obtain the heart rate variability feature; the galvanic skin response (GSR) signal is collected, and the galvanic skin response amplitude (SCR) and rise time features are extracted to obtain the galvanic skin response feature; the heart rate variability and galvanic skin response features are combined to obtain an auxiliary physiological feature vector.
[0034] Perform time alignment and weighted fusion: In this embodiment, the reduced-dimensional feature vector is time-aligned with the auxiliary physiological feature vector, and data synchronization is achieved through linear interpolation to obtain synchronized feature data. The Pearson correlation coefficient between the features of the synchronized feature data is calculated using the same formula as above to form a cross-modal correlation matrix. Based on the cross-modal correlation matrix, the fusion weight of each feature is calculated, and the weight formula is: , where r_ij is the correlation matrix element, N is the number of features, and the fusion weight is obtained; the weighted fusion method is applied to the synchronized feature data, and the fusion formula is , where F_i is the i-th eigenvector and w_i is the corresponding weight, and the multimodal eigenvector is obtained.
[0035] In this embodiment, emotional state recognition and cognitive load assessment are performed based on multimodal feature vectors to obtain a representation of the user's current mental state, including: based on a pre-trained support vector machine classifier, emotional state classification is performed on the multimodal feature vector to obtain the user's emotional state label, and the emotional state label includes positive, negative and neutral; the power ratio of the θ frequency band and the α frequency band in the spectral features of the multimodal feature vector is extracted to obtain a working memory load index; the power of the β frequency band in the spectral features of the multimodal feature vector is extracted to obtain an attention allocation index; the emotional state label is converted into a numerical representation, wherein positive is assigned a value of 1, neutral is assigned a value of 0, and negative is assigned a value of -1 to obtain an emotional state value; the emotional state value, the working memory load index and the attention allocation index are respectively linearly normalized so that their value ranges are all [0,1] to obtain a representation of the user's current mental state.
[0036] In this embodiment, EEG and auxiliary physiological signal training samples containing multiple emotional states are collected and labeled, and emotional states include three types: positive, negative, and neutral. A support vector machine (SVM) algorithm is used to train an emotional state classifier using labeled samples, and the radial basis function (RBF) is selected as the kernel function. The classifier parameters C and γ are optimized through 5-fold cross-validation to improve the classification accuracy and obtain an emotional state classifier. The multimodal feature vector is input into the emotional state classifier, and the user's emotional state label is output, indicating that the user's current emotional state is positive, negative, or neutral. For the EEG component in the multimodal feature vector, the power ratio of the θ frequency band (4-8Hz) and the α frequency band (8-13Hz) in the spectral feature is extracted to obtain a working memory load index, which reflects the degree of cognitive resource consumption. The power of the β frequency band (13-30Hz) in the spectral feature is extracted from the multimodal feature vector, and the calculation formula is: , where P(f) is the power spectral density at frequency f, and the attention allocation index is obtained to reflect the attention level; the emotional state label is converted into a numerical representation, with positive assigned to 1, neutral assigned to 0, and negative assigned to -1, to obtain the emotional state value; the emotional state value is linearly mapped to the [0,1] interval, and the formula is , where E is the emotional state value; the working memory load index and attention allocation index are normalized to the minimum and maximum values respectively; the normalized emotional state, working memory load index, and attention allocation index are combined into a three-dimensional vector to obtain the representation of the user's current psychological state.
[0037] In this embodiment, a target psychological state is determined through user input or statistical analysis based on user historical data, and music features are screened based on the difference between the target psychological state and the representation of the user's current psychological state in combination with a pre-trained music-emotion mapping network to obtain a personalized music regulation strategy, including: extracting acoustic features of music in a music library, wherein the extracted features include rhythm, pitch, and volume, to obtain a music feature vector; based on pre-labeled music emotion tag data, a support vector regression algorithm is used to train the mapping relationship between music features and emotional states to obtain a music-emotion mapping network; the target psychological state is determined through user input or statistical analysis based on user historical data, and the target psychological state is represented as a vector consistent with the format of the user's current psychological state representation; using the music-emotion mapping network, an initial music feature vector set that matches the user's current psychological state representation is queried; the difference between the user's current psychological state representation and the target psychological state is calculated to obtain a state difference vector; based on the state difference vector, the parameter value of each music feature in the initial music feature vector set is adjusted to obtain an optimized music feature vector set; and the optimized music feature vector set is mapped to specific music parameters to obtain a personalized music regulation strategy.
[0038] In this embodiment, a feature extraction algorithm is applied to the audio files in the music library to extract rhythm features (playback speed BPM), pitch features (spectral centroid) and volume features (loudness) to construct a music feature vector; labeled data of music features and corresponding emotional state labels (positive, negative, neutral) are collected and obtained through music psychology experiments to establish an association database between music features and emotional labels; based on this database, a support vector regression (SVR) algorithm is adopted, and the radial basis function (RBF) is selected as the kernel function. The parameters C and ε are optimized through 5-fold cross validation to train the mapping relationship between music features and emotional states and obtain a music-emotion mapping network; according to the emotional regulation needs clearly specified by the user, a music-emotion mapping network is generated. The target mental state is determined by seeking (such as "helping to relax" corresponds to a neutral emotional state, or "enhancing vitality" corresponds to a positive emotional state) or statistical analysis based on the user's historical data (for example, calculating the average of the user's historical emotional state), and expressed as a target emotional state label (positive, negative or neutral); the target mental state is expressed as a three-dimensional vector consistent with the user's current mental state representation format, including the target emotional state, target working memory load and target attention level; using the music-emotion mapping network, the initial music feature vector set that matches the emotional state in the user's current mental state representation is searched; the difference between the user's current mental state representation and the target mental state is calculated, and the formula is: , where D is the state difference vector, T is the target psychological state vector, and C is the current psychological state vector, to obtain the state difference vector; based on the state difference vector, adjust the parameter value of each music feature in the initial music feature vector set. The adjustment method is F_opt=F_init+α×D, where F_opt is the optimized music feature vector, F_init is the initial music feature vector, and α is the adjustment coefficient (taken as 0.1), to obtain the optimized music feature vector set; map the optimized music feature vector set to specific music parameters (rhythm, pitch, volume) to generate a personalized music control strategy.
[0039] In this embodiment, real-time music parameter adjustment and audio stream generation are performed based on a personalized music control strategy to obtain an adaptive music output stream, including: parameterizing the original music material library to extract adjustable music parameters, including rhythm, pitch and volume, to obtain an adjustable music element library; selecting music materials that meet the optimized music feature vector set from the adjustable music element library according to the personalized music control strategy to obtain an initial music material set; based on the personalized music control strategy, adjusting the music parameters of the initial music material set in real time to obtain an adjusted music stream; applying a smooth transition algorithm to the adjusted music stream to achieve seamless connection between different music materials to obtain a smooth music stream; performing sound quality optimization processing on the smooth music stream, the optimization method including dynamic range compression and equalization processing to obtain a high-quality music stream; transmitting the high-quality music stream to a terminal device for real-time playback to obtain an adaptive music output stream.
[0040] In this embodiment, an original music material library is constructed, which contains music clips of various styles and emotional types. The music in the material library is parameterized and decomposed using digital signal processing technology to extract adjustable music parameters, including rhythm (playback speed BPM), pitch (spectral centroid) and volume (loudness), and a controllable music element library is constructed. According to the optimized music feature vector set in the personalized music control strategy, based on the Euclidean distance matching algorithm, music materials that meet the requirements are retrieved from the controllable music element library. The formula is: , where F_opt_i is the i-th component of the optimized music feature vector, F_lib_i is the i-th component of the music material feature vector in the music element library, d is the distance, and the music material with the smallest distance is selected to obtain the initial music material set; based on the personalized music control strategy, the music parameters of the initial music material set are adjusted in real time, and the adjustment method is P_new=P_init+β×ΔF, where P_new is the adjusted parameter value, P_init is the initial parameter value, β is the adjustment coefficient (taken as 0.05), and ΔF is the difference between the optimized music feature vector and the initial music material feature vector, and the adjusted music stream is obtained; a smooth transition algorithm is applied to the adjusted music stream, and a linear interpolation method is used to achieve seamless connection between different music materials. The interpolation formula is , where t is the target time point, t1 and t2 are adjacent sampling time points, and X(t1) and X(t2) are corresponding sampling values, to obtain a smooth music stream; the smooth music stream is subjected to sound quality optimization processing, and the optimization methods include dynamic range compression (compression ratio set to 4:1) and equalization processing (increasing the mid-frequency band gain by 3dB) to obtain a high-quality music stream; the high-quality music stream is transmitted to the terminal device and played in real time through a low-latency audio transmission protocol (such as RTSP) to obtain an adaptive music output stream.
[0041] In this embodiment, EEG feedback is collected and the adjustment effect is evaluated based on the adaptive music output stream, and feedback data is obtained and used to optimize the music-emotion mapping network, including: during the playback of the adaptive music output stream, the user's EEG signals and auxiliary physiological signals are continuously collected to obtain real-time feedback data; feature extraction and processing are performed on the real-time feedback data to obtain a real-time feedback feature vector that is consistent with the representation format of the user's current psychological state; the real-time feedback feature vector is calculated with the determined target psychological state by Euclidean distance to obtain an effect difference value; based on the effect difference value, the adjustment effect is judged, and if the effect difference value is greater than a preset effect difference threshold, the music parameters in the personalized music control strategy are adjusted. ; Accumulate real-time feedback data and effect difference values to obtain an effect data time series; perform trend analysis on the effect data time series, use a linear regression method to calculate the effect change trend, and obtain an effect evaluation report; Based on the effect evaluation report, analyze the effectiveness of the music parameters in the adaptive music output stream, if the effect difference value of a certain music parameter decreases by more than 10% after adjustment, it is marked as valid, otherwise it is marked as invalid, and a parameter effectiveness classification result is obtained; Based on the parameter effectiveness classification result, update the training data set of the music-emotion mapping network, increase the sample weights corresponding to the valid parameters, reduce the sample weights corresponding to the invalid parameters, and retrain the music-emotion mapping network to obtain an optimized mapping network.
[0042] In this embodiment, the generated adaptive music output stream is played to the user through headphones, while the EEG acquisition device is kept working continuously, the user's EEG signals are collected in real time during the music regulation process, and the heart rate variability and skin electrodermal response signals are recorded synchronously to form real-time feedback data; the real-time feedback data is feature extracted and processed, consistent with the processing method of step 3, to obtain the emotional state value, working memory load index and attention allocation index, and normalized to obtain a real-time feedback feature vector consistent with the user's current psychological state representation format; the real-time feedback feature vector is Euclidean distance calculated with the target psychological state vector determined in step 4, and the formula is , where E is the effect difference value, R_i is the i-th component of the real-time feedback feature vector, and T_i is the i-th component of the target psychological state vector; based on the effect difference value, the adjustment effect is judged, and the preset effect difference threshold is 0.5. If the effect difference value is greater than 0.5, the music parameters in the personalized music control strategy are adjusted. The adjustment method is P_adj=P_curr+γ×E, where P_adj is the adjusted parameter value, P_curr is the current parameter value, and γ is the adjustment coefficient (taken as 0.01); accumulate real-time feedback data and effect difference values, record them according to timestamps, and obtain the effect data time series; perform trend analysis on the effect data time series using linear regression method, The formula is y=a×t+b, where y is the effect difference value, t is time, a is the trend slope, and b is the intercept. The effect change trend is obtained and an effect evaluation report is generated. Based on the effect evaluation report, the effectiveness of the music parameters in the adaptive music output stream is analyzed. If the effect difference value of a certain music parameter decreases by more than 10% after adjustment, it is marked as valid, otherwise it is marked as invalid, and the parameter effectiveness classification result is obtained. Based on the parameter effectiveness classification result, the weight of the music-emotion sample corresponding to the valid parameter in the training data set is increased by 10%, and the weight of the sample corresponding to the invalid parameter is reduced by 10%. Then, the updated weighted training data set is used to retrain the music-emotion mapping network to obtain the optimized mapping network.
[0043] In this embodiment, data security classification and privacy protection processing are performed on the high-quality EEG basic data set, multimodal feature vectors and feedback data, and stored in a secure storage database, including: performing sensitivity assessment on the high-quality EEG basic data set, multimodal feature vectors and feedback data, where the assessment criteria include data type and purpose, to obtain data sensitivity levels, wherein the data sensitivity levels include high sensitivity, medium sensitivity and low sensitivity; applying anonymization processing to the high-sensitivity data to obtain de-identified data by removing user identification information; applying a differential privacy algorithm to the de-identified data and adding Gaussian noise to obtain privacy-protected data; setting access rights for data of different sensitivity levels, where the rights are divided into three levels: read-only, read-write and no permission, to obtain permission control rules; encrypting and storing the privacy-protected data using the AES-256 algorithm to obtain an encrypted database; recording access behavior to the encrypted database to generate an access log, and detecting abnormal access behavior based on the access log, where the abnormal access behavior is judged by an access frequency threshold; integrating the encrypted database, the permission control rules and the access log to form a secure storage database for storing the high-quality EEG basic data set, the multimodal feature vectors and the feedback data.
[0044] In this embodiment, the types and uses of high-quality EEG basic data sets (original EEG signals), multimodal feature vectors (feature extraction results) and feedback data (real-time feedback feature vectors) are analyzed to formulate data sensitivity assessment standards; according to the standards, the original EEG signals are marked as highly sensitive, the multimodal feature vectors are marked as moderately sensitive, and the feedback data are marked as low sensitive to obtain data sensitivity levels; anonymization processing is applied to data with high sensitivity levels (such as original EEG signals) to obtain de-identified data by removing identification information such as user names and IDs; a differential privacy algorithm is applied to the de-identified data, and Gaussian noise is added to the noise. The standard deviation σ=0.1×data range is used to obtain privacy-protected data. For data of different sensitivity levels, access permissions are set. High-sensitivity data only allows read-only permissions, medium-sensitivity data allows read and write permissions, and low-sensitivity data has no permission restrictions. The permission control rules are obtained. Privacy-protected data is encrypted and stored using the AES-256 algorithm with a key length of 256 bits. The encryption formula is C=AES_256_encrypt(M,K), where C is the ciphertext, M is the plaintext, and K is the key. AES_256_encrypt represents a function or operation that uses the AES-256 algorithm for encryption. In actual implementation, this is usually a function call in a programming language or encryption library to obtain an encrypted database; the access behavior of the encrypted database is recorded to generate an access log, which includes the operation time, operation type (read / write), and operation subject (user ID); abnormal access behavior is detected based on the access log, and abnormal access behavior is judged by the access frequency threshold. The threshold is 10 times per minute. If it exceeds the threshold, it is marked as abnormal. The encrypted database, permission control rules and access logs are integrated into a unified data management system to form a secure storage database for storing high-quality EEG basic data sets, multimodal feature vectors and feedback data.
[0045] In this embodiment, based on the feedback data in the secure storage database, the effect of the adaptive music output stream is continuously evaluated to obtain an effect evaluation report, including: extracting feedback data from the secure storage database to obtain an effect data set; performing time series analysis on the effect data set, using a time window method to calculate the change in emotional state and cognitive load at different time points, and applying a moving average method to obtain a smoothed change trend to obtain an effect trend report; based on the effect data set, calculating the improvement rate of the adaptive music output stream on the user's emotional state and cognitive load, the improvement rate being calculated by the difference between the target psychological state and the actual state to obtain an effect evaluation report; based on the effect evaluation report, extracting the music parameter combination with the highest improvement rate to obtain an optimized parameter recommendation; based on the optimized parameter recommendation, updating the training data set of the music-emotion mapping network, increasing the sample weights corresponding to the high improvement rate, and retraining the music-emotion mapping network to obtain an optimized mapping network; generating user personalized usage recommendations based on the effect evaluation report, the recommendations including recommended music parameters and usage scenarios to obtain a personalized guidance plan; integrating the effect trend report, the effect evaluation report, and the personalized guidance plan to obtain the effect evaluation report.
[0046] In this embodiment, feedback data, including real-time feedback feature vectors and effect difference values, are extracted from a secure storage database and sorted by timestamp to obtain an effect data set. Time series analysis is performed on the effect data set, and the changes in emotional state and cognitive load at different time points are calculated using a time window method (window length is 10 minutes). A moving average method (average window length is 3 time points) is applied to obtain a smoothed change trend, and an effect trend report is generated. Based on the effect data set, the improvement rate of the adaptive music output stream on the user's emotional state and cognitive load is calculated, and the improvement rate formula is: , where D_init is the difference between the initial state and the target psychological state, and D_final is the difference between the final state and the target psychological state. The differences are calculated using the Euclidean distance to obtain an effect evaluation report. Based on the effect evaluation report, the music parameter combination (rhythm, pitch, volume) with the highest improvement rate is extracted. For example, if the parameter combination with the highest improvement rate is BPM = 80, spectral centroid = 500Hz, and loudness = -10dB, it is marked as an optimized parameter recommendation. Based on the optimized parameter recommendation, the sample weights corresponding to the parameter combination with the high improvement rate are increased in the training dataset of the music-emotion mapping network. The amount of weight increase is proportional to the improvement rate. The music-emotion mapping network is retrained using the updated weighted training dataset to obtain an optimized mapping network. Based on the effect evaluation report, personalized user usage recommendations are generated. For example, if the scenario with the highest improvement rate is "relaxation", the recommendation is "use music with BPM = 80 in stressful scenarios" to obtain a personalized guidance plan. The effect trend report, effect evaluation report, and personalized guidance plan are integrated to generate a final effect evaluation report. The report content includes trend charts, improvement rate data, and usage recommendations.
[0047] The above describes the intelligent music control method based on multimodal feature recognition of EEG signals in the embodiment of the present application. The following describes the intelligent music control system based on multimodal feature recognition of EEG signals in the embodiment of the present application. Figure 2 In the embodiment of the present application, an embodiment of an intelligent music control system based on multimodal feature recognition of EEG signals includes: The EEG signal acquisition and processing module is used to perform anti-interference acquisition and adaptive filtering on EEG signals to obtain high-quality EEG basic data sets; A multimodal feature fusion module is used to perform multi-dimensional feature extraction and cross-modal information fusion based on the high-quality EEG basic data set to obtain a multimodal feature vector; A mental state recognition module, configured to perform emotional state recognition and cognitive load assessment based on the multimodal feature vector to obtain a representation of the user's current mental state; A music-emotion mapping module is used to determine a target mental state through user input or statistical analysis based on historical user data. Based on the difference between the target mental state and the user's current mental state representation, music features are screened in combination with a pre-trained music-emotion mapping network to obtain a personalized music control strategy. A personalized control strategy module, configured to adjust music parameters and generate an audio stream in real time based on the personalized music control strategy to obtain an adaptive music output stream; A music reconstruction output module is used to collect EEG feedback and evaluate the adjustment effect based on the adaptive music output stream, obtain feedback data and use it to optimize the music-emotion mapping network; A data security protection module, configured to perform data security classification and privacy protection processing on the high-quality EEG basic data set, the multimodal feature vector, and the feedback data, and store them in a secure storage database; A long-term tracking and evaluation module, configured to continuously evaluate the effect of the adaptive music output stream based on the feedback data in the secure storage database to obtain an effect evaluation report; The modules are connected to each other via wired and / or wireless means to achieve data transmission between modules.
[0048] In a specific embodiment of the present invention, an intelligent music control system and method based on multimodal feature recognition of EEG signals, through anti-interference acquisition and adaptive filtering of EEG signals, combined with multi-dimensional physiological feature extraction and cross-modal information fusion, achieves accurate identification of emotional state and cognitive load assessment, establishes a context-sensitive music-emotion mapping network, designs a personalized music control strategy, generates an adaptive music output stream, constructs a closed-loop optimization feedback mechanism and a safety specification database, and ultimately forms a comprehensive evaluation report for intelligent music control. This method realizes full-process intelligent processing from EEG signal acquisition to music emotion control, continuously optimizes the control effect through closed-loop feedback, and can provide users with accurate and personalized emotion and cognitive state regulation services.
[0049] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
[0050] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0051] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0052] In the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0053] In the description of the present invention, “several” means one or more, and “a large number” means two or more.
[0054] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0055] The formulas in this manual are all dimensionless and calculated using numerical values. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field based on actual conditions.
[0056] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. An intelligent music control method based on multimodal feature recognition of EEG signals, characterized in that: include: Step 1: Anti-interference acquisition and adaptive filtering of EEG signals are performed to obtain a high-quality EEG basic data set; Step 2: performing multi-dimensional feature extraction and cross-modal information fusion based on the high-quality EEG basic data set to obtain a multimodal feature vector; Step 3: performing emotional state recognition and cognitive load assessment based on the multimodal feature vector to obtain a representation of the user's current psychological state; Step 4: Determine the target mental state through user input or statistical analysis based on user historical data, and perform music feature screening based on the difference between the target mental state and the user's current mental state representation in combination with a pre-trained music-emotion mapping network to obtain a personalized music control strategy; Step 5: Perform real-time music parameter adjustment and audio stream generation based on the personalized music control strategy to obtain an adaptive music output stream; Step 6: Collecting EEG feedback and evaluating the adjustment effect based on the adaptive music output stream to obtain feedback data and use it to optimize the music-emotion mapping network; Step 7: Perform data security classification and privacy protection processing on the high-quality EEG basic data set, the multimodal feature vector, and the feedback data, and store them in a secure storage database; Step 8: Based on the feedback data in the secure storage database, continuously evaluate the effect of the adaptive music output stream to obtain an effect evaluation report.
2. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 1 is characterized in that: The step 1 comprises: Optimizing the spatial distribution of multi-lead EEG electrodes to obtain a full-brain coverage acquisition scheme, and applying active shielding technology to the full-brain coverage acquisition scheme to obtain an anti-interference acquisition circuit; Performing real-time sampling and quantization on the raw EEG data output by the anti-interference acquisition circuit to obtain a digitized EEG signal, and applying adaptive notch filtering to the digitized EEG signal to obtain a power frequency interference-free signal; Performing wavelet decomposition and reconstruction on the power frequency interference-removed signal to obtain a denoised EEG signal; Performing independent component analysis on the denoised EEG signal to obtain independent neural source signals; Performing artifact detection and removal on the independent neural source signal to obtain a pure EEG signal; Performing time-frequency analysis on the pure EEG signal to obtain enhanced EEG features; The enhanced EEG features are subjected to data standardization to obtain standardized EEG data, and the standardized EEG data are stored as a structured data set to obtain a high-quality EEG basic data set.
3. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 2 is characterized in that: The step of performing wavelet decomposition and reconstruction on the power frequency interference-removed signal to obtain a denoised EEG signal includes: Select the wavelet basis function suitable for EEG signals to obtain the optimized wavelet transform operator; Applying the optimized wavelet transform operator to perform multi-level decomposition on the power frequency interference-removed signal to obtain wavelet coefficient sets in different frequency ranges; Performing threshold denoising on the wavelet coefficient set to obtain denoised wavelet coefficients; Performing inverse wavelet transform on the denoised wavelet coefficients to obtain a denoised EEG signal.
4. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 2 is characterized in that: The performing artifact detection and removal on the independent neural source signal to obtain a pure EEG signal includes: Acquiring and labeling training sample neural source signals containing eye movement, heartbeat, and electromyographic artifact features to obtain artifact labeled signal samples; extracting features of the artifact labeled signal samples to obtain an artifact feature template set; Based on the artifact feature template set, a support vector machine algorithm is used to train an artifact recognition classifier to obtain an artifact detection model; Extracting the features of the independent neural source signal so that the features are consistent with the feature format of the artifact feature template set to obtain the neural source features to be detected; Using the artifact detection model to perform component classification on the neural source features to be detected to obtain artifact component labeling results; Based on the artifact component labeling results, the artifact energy ratio of each independent neural source signal is calculated to obtain an artifact impact score; Based on the artifact impact score, an artifact energy ratio threshold θ is set. The artifact energy ratio threshold θ is calculated by the formula θ=θ_base×(1+α×S), where θ_base is the base threshold, α is the adjustment coefficient, and S is the artifact impact score; Filtering the artifact components in the independent neural source signal according to the artifact energy ratio threshold θ to obtain a preliminary purified EEG signal; Applying principal component analysis to the preliminary purified EEG signal to obtain main neural source signals; Performing a signal integrity assessment on the main neural source signal, calculating the signal-to-noise ratio score and the spectral energy distribution balance score by weighted average to obtain a data quality score; Based on the data quality score, if the score is lower than a preset threshold, the main neural source signal is interpolated and compensated to obtain a pure EEG signal; if the score is not lower than the preset quality score threshold, the main neural source signal is directly used as a pure EEG signal.
5. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 1 is characterized in that: The step 2 includes: Performing time domain analysis on the high-quality EEG basic data set, extracting mean, variance and peak features, and obtaining a statistical feature set; Performing frequency domain analysis on the high-quality EEG basic data set, extracting power features of the δ, θ, α, β, and γ frequency bands, and obtaining a spectral feature set; Performing nonlinear analysis on the high-quality EEG basic data set, calculating sample entropy and fractal dimension, and obtaining a nonlinear feature set; Calculating the Pearson correlation coefficients between different brain regions for the high-quality EEG basic data set to obtain a functional connectivity matrix; Applying graph theory analysis to the functional connectivity matrix, calculating node degrees and clustering coefficients, and obtaining a network feature set; Based on pre-labeled training data, a random forest model is trained to obtain a feature importance evaluation model; the feature importance evaluation model is applied to calculate an importance score for each feature in the statistical feature set, the spectral feature set, the nonlinear feature set, and the network feature set to obtain an important feature set; Applying a principal component analysis algorithm to the important feature set to perform dimensionality reduction processing to obtain a feature vector after dimensionality reduction; collecting auxiliary physiological signals, the auxiliary physiological signals including heart rate variability and galvanic skin response signals, and performing feature extraction on the auxiliary physiological signals to obtain auxiliary physiological feature vectors; Time-aligning the dimension-reduced feature vector with the auxiliary physiological feature vector, achieving data synchronization through a linear interpolation method, and obtaining synchronized feature data; A weighted fusion method is applied to the synchronous feature data, and the fusion weight is calculated by Pearson correlation coefficient to obtain a multimodal feature vector.
6. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 5 is characterized in that: The Pearson correlation coefficients of different brain regions are calculated for the high-quality EEG basic data set to obtain a functional connectivity matrix, including: Calculate the Pearson correlation coefficient of the EEG signal time series between different brain regions to obtain the functional connectivity strength matrix; Applying a fixed threshold to the functional connectivity strength matrix, retaining connections with correlation coefficients greater than 0.6, and obtaining a significant connectivity matrix; Applying graph theory analysis to the significant connection matrix, calculating the node degree of each brain region, and defining brain regions with node degrees higher than the average node degree of the whole brain as key brain regions, thereby obtaining the distribution of key brain regions; Performing modular analysis on the significant connection matrix, dividing the functional modules using the Louvain algorithm, and obtaining a module division result; The average connection strength between modules in the module division result is calculated to obtain the inter-module interaction strength.
7. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 5 is characterized in that: The step 3 includes: Based on a pre-trained support vector machine classifier, the multimodal feature vector is subjected to emotional state classification to obtain an emotional state label of the user, wherein the emotional state label includes positive, negative, and neutral; extracting a power ratio of the θ frequency band and the α frequency band in the frequency spectrum characteristics of the multimodal feature vector to obtain a working memory load index; Extracting the power of the β frequency band in the spectral feature of the multimodal feature vector to obtain an attention allocation index; The emotional state label is converted into a numerical representation, where positive is assigned a value of 1, neutral is assigned a value of 0, and negative is assigned a value of -1, to obtain the emotional state value; the emotional state value, the working memory load index, and the attention allocation index are linearly normalized so that their value ranges are all [0,1], to obtain the representation of the user's current psychological state.
8. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 1 is characterized in that: The step 4 comprises: Extract acoustic features of music in the music library, including rhythm, pitch and volume, to obtain music feature vectors; Based on pre-labeled music emotion label data, the support vector regression algorithm is used to train the mapping relationship between music features and emotional states to obtain a music-emotion mapping network; Determine a target mental state through user input or statistical analysis based on user historical data, and represent the target mental state as a vector consistent with the representation format of the user's current mental state; Utilizing the music-emotion mapping network, querying a set of initial music feature vectors that matches the representation of the user's current mental state; Calculating the difference between the user's current mental state representation and the target mental state to obtain a state difference vector; Based on the state difference vector, adjusting the parameter value of each music feature in the initial music feature vector set to obtain an optimized music feature vector set; The optimized music feature vector set is mapped into specific music parameters to obtain a personalized music control strategy.
9. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 8, characterized in that: The step 5 comprises: Perform parameterization on the original music material library to extract adjustable music parameters, including rhythm, pitch and volume, to obtain an adjustable music element library; According to the personalized music control strategy, music materials that meet the optimized music feature vector set are selected from the controllable music element library to obtain an initial music material set; Based on the personalized music control strategy, music parameters of the initial music material set are adjusted in real time to obtain an adjusted music stream; Applying a smooth transition algorithm to the adjusted music stream to achieve seamless connection between different music materials through a linear interpolation method to obtain a smooth music stream; Performing sound quality optimization processing on the smooth music stream, the optimization method including dynamic range compression and equalization processing, to obtain a high-quality music stream; The high-quality music stream is transmitted to a terminal device for real-time playback to obtain an adaptive music output stream.
10. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 1, characterized in that: The step 6 comprises: During the playback of the adaptive music output stream, the user's EEG signals and auxiliary physiological signals are continuously collected to obtain real-time feedback data; Extracting and processing the real-time feedback data to obtain a real-time feedback feature vector consistent with a representation format of the user's current mental state; Performing Euclidean distance calculation on the real-time feedback feature vector and the target psychological state determined in step 4 to obtain an effect difference value; Based on the effect difference value, determining the adjustment effect, and if the effect difference value is greater than a preset threshold, adjusting the music parameters in the personalized music control strategy; Accumulating the real-time feedback data and the effect difference value to obtain an effect data time series; Perform trend analysis on the time series of the effect data, calculate the effect change trend using a linear regression method, and obtain an effect evaluation report; Based on the effect evaluation report, the effectiveness of the music parameters in the adaptive music output stream is analyzed. If the effect difference value of a certain music parameter is reduced by more than 10% after adjustment, it is marked as effective; otherwise, it is marked as invalid, thereby obtaining a parameter effectiveness classification result; Based on the parameter effectiveness classification results, the training data set of the music-emotion mapping network is updated, the sample weights corresponding to the effective parameters are increased, the sample weights corresponding to the invalid parameters are reduced, and the music-emotion mapping network is retrained to obtain an optimized mapping network.
11. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 1, characterized in that: The step 7 comprises: Performing a sensitivity assessment on the high-quality EEG basic data set, the multimodal feature vector, and the feedback data, where the assessment criteria include data type and usage, to obtain a data sensitivity level, where the data sensitivity level includes high sensitivity, medium sensitivity, and low sensitivity; Applying anonymization processing to the highly sensitive data to obtain de-identified data by removing user identification information; Applying a differential privacy algorithm to the de-identified data and adding Gaussian noise to obtain privacy-preserving data; For data of different sensitivity levels, set access permissions, which are divided into three levels: read-only, read-write, and no permission, and obtain permission control rules; Encrypting and storing the privacy-protected data using the AES-256 algorithm to obtain an encrypted database; Recording access behavior to the encrypted database to generate an access log, and detecting abnormal access behavior based on the access log, wherein abnormal access behavior is determined by an access frequency threshold; The encrypted database, the authority control rules and the access log are integrated to form a secure storage database for storing the high-quality EEG basic data set, the multimodal feature vector and the feedback data.
12. The intelligent music control method based on multimodal feature recognition of EEG signals according to claim 1, characterized in that: The step 8 comprises: Extracting the feedback data from the secure storage database to obtain an effect data set; Performing a time series analysis on the effect data set, using a time window method to calculate the changes in affective state and cognitive load at different time points, and applying a moving average method to obtain a smoothed change trend to obtain an effect trend report; Based on the effect data set, calculating the improvement rate of the adaptive music output stream on the user's emotional state and cognitive load, the improvement rate being calculated by the difference between the target psychological state and the actual state, to obtain an effect evaluation report; Based on the effect evaluation report, extract the music parameter combination with the highest improvement rate and obtain optimization parameter suggestions; According to the optimization parameter suggestions, the training data set of the music-emotion mapping network is updated, the weights of samples corresponding to high improvement rates are increased, and the music-emotion mapping network is retrained to obtain an optimized mapping network; Based on the effect evaluation report, generate personalized usage suggestions for the user, including recommended music parameters and usage scenarios, to obtain a personalized guidance plan; The effect trend report, the effect evaluation report and the personalized guidance plan are integrated to obtain an effect evaluation report.
13. An intelligent music control system based on multimodal feature recognition of EEG signals, which is used to implement the intelligent music control method based on multimodal feature recognition of EEG signals according to any one of claims 1 to 12, characterized in that: include: The EEG signal acquisition and processing module is used to perform anti-interference acquisition and adaptive filtering on EEG signals to obtain high-quality EEG basic data sets; A multimodal feature fusion module is used to perform multi-dimensional feature extraction and cross-modal information fusion based on the high-quality EEG basic data set to obtain a multimodal feature vector; A mental state recognition module, configured to perform emotional state recognition and cognitive load assessment based on the multimodal feature vector to obtain a representation of the user's current mental state; A music-emotion mapping module is used to determine a target mental state through user input or statistical analysis based on historical user data. Based on the difference between the target mental state and the user's current mental state representation, music features are screened in combination with a pre-trained music-emotion mapping network to obtain a personalized music control strategy. A personalized control strategy module, configured to adjust music parameters and generate an audio stream in real time based on the personalized music control strategy to obtain an adaptive music output stream; A music reconstruction output module is used to collect EEG feedback and evaluate the adjustment effect based on the adaptive music output stream, obtain feedback data and use it to optimize the music-emotion mapping network; A data security protection module, configured to perform data security classification and privacy protection processing on the high-quality EEG basic data set, the multimodal feature vector, and the feedback data, and store them in a secure storage database; A long-term tracking and evaluation module, configured to continuously evaluate the effect of the adaptive music output stream based on the feedback data in the secure storage database to obtain an effect evaluation report; The modules are connected via wired and / or wireless means to achieve data transmission between modules.
Citation Information
Patent Citations
MI electroencephalogram signal recognition method based on feature fusion and particle swarm optimization algorithm
CN111797674A
Electroencephalogram signal quality evaluation method based on characteristic wave detection and staging algorithm
CN114680904A
Multi-modal emotion recognition method and device, equipment and storage medium
CN114947852A
Learning state analysis system based on electroencephalogram wearable device
CN118526211A
Self-adaptive music regulation and control method and system based on electroencephalogram emotion recognition
CN118732847A
Cited By
Smell recognition and emotion analysis system and method based on multi-modal physiological signals, terminal and medium
CN120814823A
Non-invasive acupoint bio-electricity signal information acquisition method and system
CN120959686A
Postoperative brain function state evaluation method and system based on multi-modal fusion
CN120959693A
Thin-wall part milling state evaluation method and system based on sliding window-wavelet transform
CN121278364A
Pet physiological parameter detection and analysis method and system based on biosensor
CN121287154A