Intelligent music therapy system based on electroencephalogram emotion fluctuation index
By combining EEG emotional fluctuation indicators and computational musicology, a closed-loop feedback control and individual adaptive mechanism for music therapy system was established. This solved the problems of insufficient emotion quantification and individual differences in existing technologies, and enabled real-time assessment and individualized music matching for depression-related negative emotions, thereby improving the stability and applicability of therapeutic effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-01
AI Technical Summary
Existing music therapy systems based on physiological signals such as electroencephalograms lack objective and interpretable temporal quantitative indicators for intervention in depression-related negative emotions. Music matching lacks controllable variables based on music theory, and closed-loop feedback lacks individual adaptive mechanisms, resulting in unstable efficacy and large individual differences.
By extracting the negative emotion index and the emotion fluctuation index from the EEG emotion fluctuation index, and combining them with computational musicology modeling, a mapping relationship between music theory characteristics and emotion fluctuation is established. A closed-loop feedback control and individual adaptive iteration mechanism are constructed to achieve real-time music matching and strategy updates.
This enables objective quantitative assessment of music therapy, interpretable improvement in efficacy, stability and long-term applicability of individualized intervention strategies, and enhances the effectiveness and individual adaptability of music therapy.
Smart Images

Figure CN121944334A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of affective computing and intelligent closed-loop non-pharmacological intervention for depression. It integrates technologies such as EEG signal processing, emotion recognition, music information retrieval and recommendation, computational musicology modeling, and closed-loop feedback control to achieve temporal assessment of depression-related negative emotions under music intervention conditions. At the same time, it establishes a mapping relationship between music characteristics and emotional fluctuations, thereby driving music matching and closed-loop feedback updates in real time to form an individual adaptive music intervention strategy. Background Technology
[0002] The diagnosis and treatment of depressive disorders currently face severe challenges characterized by "three lows and one high": low clinical recognition rate (47.3%), low consultation rate (25%), low treatment efficiency (20%), and high relapse rate (75%). Patients often experience varying degrees of impairment in sleep, attention, motivation, pleasure experience, and social functioning. Traditional non-pharmacological interventions for depression have limitations in accessibility, long-term adherence, side effects, or cost, resulting in a lack of sustainable, quantifiable, and self-manageable mood regulation tools for some populations in non-medical settings. Music therapy, due to its non-invasiveness, accessibility, and ability to be continuously implemented in daily environments, has been widely used in applications such as psychological rehabilitation, mood regulation, and stress management. Furthermore, advancements in wearable sensors and mobile computing capabilities have made the "monitoring-intervention-feedback" closed-loop music therapy based on physiological signals feasible for engineering implementation.
[0003] Existing music therapy systems based on physiological signals such as electroencephalograms (EEGs) generally employ the following technical approach: EEG, heart rate, and skin conductance signals are collected via wearable devices; these signals are then filtered, segmented, and feature extracted; machine learning models are used to output emotion categories (such as pleasure, calmness, tension, and depression) or single-dimensional emotion intensity; subsequently, music is recommended based on the emotion recognition results, or tracks are selected from a pre-set music collection to achieve relaxation, soothing, or activation. Some systems incorporate user subjective feedback (such as like / dislike, skip, rating) as an update basis, but most still rely primarily on "tag recommendations" or "rules of thumb," making it difficult to explain from the perspective of the music's internal structure why it can improve mood.
[0004] However, the existing technologies mentioned above have significant shortcomings in music intervention scenarios for depression-related negative emotions. First, there is a lack of objective and interpretable temporal quantitative indicators of emotion. The core characteristics of depression-related emotions are not only "whether it is negative at the moment," but also include temporal dimensions such as "whether the negativity fluctuates repeatedly, whether it is difficult to stabilize, and whether it is overly sensitive to external stimuli." Existing systems mostly use short-term emotional classification or simple intensity scoring as evaluation, which cannot distinguish between two different states: "high but stable negativity" and "low but highly fluctuating negativity." It is even more difficult to assess the dynamic trajectory of emotion from fluctuation to convergence during music intervention. Therefore, music recommendation strategies are prone to instability and difficulty in reproducing therapeutic effects.
[0005] Secondly, the lack of controllable and calculable music theory variables in music matching leads to a lack of interpretability and controllability in recommendations. The impact of music on emotions is closely related to its internal structure, such as rhythmic tempo, beat stability, mode, harmonic tension, melodic fluctuations, timbre brightness, and loudness envelope. However, existing systems often match based on style tags, emotional tags, or subjective similarity, making it difficult to establish a mapping model of "music theory structure - emotional response." This makes it impossible to form a clear direction for regulation in closed-loop control. For example, when it is necessary to reduce fluctuations, which combination of rhythmic stability and harmonic tension characteristics should be prioritized? Or, when the negative index is consistently high, how should the activation level of the music be adjusted without inducing fluctuations?
[0006] Furthermore, closed-loop feedback is often insufficient and lacks individual adaptive mechanisms. Different individuals exhibit significant differences in their emotional responses to the same music, and the same person's responses can vary depending on different fatigue levels, situations, or intervention stages. Recommendations based solely on fixed rules or one-time calibration models are prone to "effective initially but ineffective later" or "instable results due to individual differences." While some systems incorporate feedback, it often relies on subjective ratings or skips behavioral updates, making it difficult to achieve rapid feedback regulation of emotional fluctuations on a short timescale and hindering the development of interpretable, individualized intervention strategies.
[0007] Therefore, it is necessary to propose a new technical solution that uses EEG emotional fluctuation indicators as a closed-loop control quantity to objectively quantify the degree of negativity and fluctuation of emotions under music intervention. Furthermore, by establishing a mapping relationship between music theory characteristics and emotional fluctuations through computational musicology, real-time closed-loop music matching and individual adaptive iterative updates can be achieved, thereby improving the effectiveness, stability, and individual adaptability of music therapy. Summary of the Invention
[0008] In view of this, the purpose of the present invention is to provide an intelligent music therapy system based on EEG emotional fluctuation indicators.
[0009] A smart music therapy system based on EEG emotional fluctuation indicators, characterized in that it includes: EEG acquisition unit, used to acquire EEG signals from subjects; The preprocessing unit is used to perform window processing on the subject's electroencephalogram (EEG) signals according to a set time window. The feature and index unit is used to: extract the static and dynamic spectral features of the EEG signals within each time window, and combine the two to obtain the window feature vector; and then extract the negative emotion index and emotion fluctuation index of the EEG signals within each time window. The music theory characteristic unit performs computational musicological analysis on music fragments, extracts music theory characteristics, and forms a music characteristic vector. The music library unit is used to store music fragments and their corresponding music feature vectors, and supports searching candidate music sets by feature. Model units are used for: (1) Construct a mapping model that uses the current state vector composed of the combination of the negative emotion index and the emotion fluctuation index as input and the candidate music characteristic vector as input, and outputs the regression prediction values of the change in the negative emotion index and the change in the emotion fluctuation index. (2) Construct an objective function based on the output of the mapping model, and select the most suitable music segment accordingly; The adaptive update unit is used to: construct a loss function based on the subject's feedback on the current music clip and the negative emotion index and emotion fluctuation index predicted by the mapping model, and update the parameters of the mapping model and the weight parameters of the objective function.
[0010] Preferably, the EEG acquisition unit is used to acquire EEG signals from the frontal lobe or prefrontal lobe.
[0011] Preferably, the preprocessing unit performs bandpass filtering, power frequency notch filtering, baseline correction, and eye movement and electromyography artifact suppression on the EEG signal.
[0012] Preferably, the static spectral features are used to describe the steady-state spectral structure within the window, and the dynamic spectral features are used to describe the trend of spectral variation over time between adjacent windows.
[0013] Preferably, the negative sentiment index is output by a random forest model; assuming the random forest model contains M parallel decision trees, each decision tree corresponds to the feature vector of the k-th window. Output negative probability The negative emotion index is defined as follows: ; The emotional fluctuation index is obtained by using the L2 norm of the difference between adjacent window state vectors as the fluctuation intensity and then smoothing it.
[0014] Preferably, in the model unit, the objective function as follows:
[0015] in, represents the state vector of the k-th window; i represents the sequence number of the selected music segment; ΔENI represents the change in the negative exponent, and ΔEVI represents the change in the volatility exponent; D is the smoothing penalty term, used to constrain abrupt changes in tempo, mode, and loudness between adjacent segments; Ω is the preference and safety constraint term, used to constrain user discomfort, overstimulation, or high-risk skipping situations; α, β, γ, and δ are the corresponding weight parameters. Select the music segment that minimizes the objective function value. Play it.
[0016] Preferably, the loss function is expressed as:
[0017] Wherein, ΔENI represents the change in the negative index, and ΔEVI represents the change in the volatility index; and These represent the changes in negative emotion and emotional fluctuations predicted by the mapping model, respectively.
[0018] Preferably, in the adaptive update unit, the parameter update of the mapping model is expressed as: ; in, The parameters represent the mapping model. This represents the hyperparameters of gradient descent. Represents the loss function Regarding mapping model parameters The differential.
[0019] Preferably, the weight parameter update of the objective function in the adaptive update unit is expressed as follows: ; in, , These represent the weight parameter vectors in the objective function for the (n+1)th and nth iterations, respectively; This represents the hyperparameters of gradient descent. Indicates the loss of function Regarding weight parameters The differential.
[0020] The present invention has the following beneficial effects: First, this invention proposes and engineered a brainwave-based emotional fluctuation index system based on static and dynamic spectral characteristics, including at least the Emotional Negativity Index (ENI) and the Emotional Fluctuation Index (EVI). This system can output data in real time during music intervention and can distinguish between two dimensions: "negativity level" and "fluctuation stability." This elevates the assessment of music therapy from a single-point classification or subjective feeling to a traceable, time-series quantitative indicator, thereby enhancing the objectivity and interpretability of the therapy evaluation.
[0021] Secondly, this invention uses computational musicology to perform computable modeling of the internal structure of music, transforming music theory characteristics such as rhythm, tempo, mode, harmonic tension, timbre, and loudness into musical characteristic vectors. It also establishes a mapping model between musical characteristics and emotional fluctuation indicators, elevating music matching from traditional label-based recommendations to a closed-loop regulation that is "predictable, controllable, and explainable." The system can not only answer "what music was recommended," but also explain "why this type of music theory structure is more suitable for reducing negativity or suppressing fluctuations."
[0022] Third, this invention constructs a closed-loop feedback mechanism with ENI and EVI as control variables, and supports online updates and individual adaptive iteration. The system can dynamically adjust music selection based on real-time user feedback, and learn from individual responses to different music characteristics in multiple sessions, thereby significantly improving the stability, long-term applicability, and individualization of the intervention effect, and reducing situations where "the same music has large differences in effect on different people" or "the same person's effect is unstable at different stages."
[0023] Fourth, this invention preferably employs a random forest model to estimate the negative emotion index (ENI) and predict the music-emotion mapping. Random forests exhibit good stability in scenarios with small samples, high noise levels, and a need for interpretability. They can explain the contributions of key frequency bands, key channels, and key music theory characteristics to decision-making through feature importance, tree path rules, and other methods, thereby supporting the integration of closed-loop strategies into clinical research.
[0024] Fifth, this invention introduces smoothing constraints and preference constraints into closed-loop decision-making, and can set safety thresholds and destimulation strategies, enabling continuous intervention while ensuring user experience and safety. It is applicable to various non-medical scenarios such as home, campus, and office, and is also easy to integrate with other health data such as scales, sleep, and activity to form a comprehensive digital therapy plan. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the overall structure of an intelligent music therapy system based on EEG emotional fluctuation indicators, showing the connection relationships between the EEG acquisition unit 1, preprocessing unit 2, feature and indicator unit 3, model unit 4, music playback unit 5, music theory characteristics unit 6, music library unit 7, adaptive update unit 8, and reporting and log unit 9.
[0026] Figure 2 This is a schematic diagram of the closed-loop music therapy method of the present invention, showing the process from EEG acquisition and baseline establishment, preprocessing and windowing, static and dynamic spectral feature extraction, emotion index calculation, music characteristic modeling and mapping prediction, closed-loop music selection and playback, to feedback update and individualized iteration.
[0027] Figure 3 This diagram illustrates the calculation of the Negativity Index and Emotional Fluctuation Index, showing the static and dynamic characteristics of EEG signals obtained through windowing and frequency / time-frequency analysis, and calculating the relationship between ENI and EVI in the time sequence.
[0028] Figure 4 A schematic diagram illustrating the modeling of music theory characteristics shows how features such as rhythm, mode, harmonic tension, timbre, and loudness are extracted from musical fragments and formed into music characteristic vectors. And the process of storing it in the music library.
[0029] Figure 5 This diagram illustrates the music-emotion mapping and closed-loop matching decision-making process. It shows the state and music characteristic input mapping model, predicts index changes, constructs an objective function to select the optimal music segment, and forms a closed loop through feedback updates.
[0030] Figure 6 This diagram illustrates how individual adaptive iterative updates are used to calculate feedback signals and update individual profiles, mapping models, or strategies, thereby influencing music matching in the next session. Detailed Implementation
[0031] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0032] The technical problem this invention aims to solve is: addressing the common issues in existing music therapy systems based on physiological signals such as electroencephalography (EEG), such as insufficient quantitative indicators of emotion, lack of temporal fluctuation characterization and interpretation, lack of music theory-controllable variables in music matching, and insufficient closed-loop feedback and individual adaptive capabilities. This invention provides an intelligent music therapy system and its implementation method with EEG emotional fluctuation indicators as the core, enabling the system to assess the subject's negative emotions and emotional fluctuation states in real time during music intervention, and to perform closed-loop music matching and strategy updates based on music theory characteristics and an emotional fluctuation mapping model.
[0033] At the technical implementation level, this invention proposes an intelligent music therapy system and its implementation method with EEG emotional fluctuation indicators as the core control quantity. By extracting and fusing the static and dynamic spectral characteristics of the subject's EEG signals under music intervention, interpretable negative emotion index and emotional fluctuation index are constructed to achieve temporal assessment of depression-related negative emotions. At the same time, musicological modeling is performed on the music theory characteristics of music, such as rhythm, mode, harmonic tension, timbre and loudness, to establish a mapping relationship between music characteristics and emotional fluctuations, thereby driving music matching and closed-loop feedback updates in real time and forming an individual adaptive music intervention strategy.
[0034] To achieve the above objectives, this invention needs to address at least the following key technical points. First, it requires the construction of an emotion fluctuation index and an emotion negativity index, enabling real-time calculation within a short time window and possessing clear physical meaning and interpretability. The emotion negativity index reflects the intensity or probability of current negative emotions, while the emotion fluctuation index reflects the amplitude and speed of emotion changes over time, and can distinguish between different states such as stable negativity and drastic fluctuations. Second, it requires establishing a music theory characteristic modeling method, representing music fragments as structured music characteristic vectors, and establishing a mapping model between music characteristics and emotion fluctuation indices, allowing the system to predict and interpret the direction and intensity of the influence of specific music characteristics on emotion negativity and emotion fluctuation. Third, it requires constructing a closed-loop feedback control and individual adaptive iterative mechanism to achieve a cycle of "collection—evaluation—matching—playback—feedback—update," continuously optimizing individualized intervention strategies while ensuring user experience, thereby improving the stability and long-term applicability of therapeutic effects.
[0035] In summary, the above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. To address the aforementioned technical problems, the present invention provides an intelligent music therapy system based on EEG emotional fluctuation indicators and its implementation method.
[0036] The overall structure of the system is as follows: Figure 1 As shown, it mainly includes: EEG acquisition unit used to acquire EEG signals from subjects; A preprocessing unit for filtering, artifact suppression, and windowing of EEG signals; Feature and indicator units used for static and dynamic spectral feature extraction and sentiment index calculation; Model unit used for emotion recognition and temporal emotion state estimation (random forest model); A music playback unit used to output music and receive user interaction information; Music theory feature unit used for computational musicological modeling of music; Music library units used to store music fragments and music feature vectors; An adaptive update unit used for online updating of mapping models or strategies; And reporting and log units for logging, generating assessment reports and audit information.
[0037] The units are connected via wired or wireless means to form a closed-loop feedback control link.
[0038] The method flow of the present invention is as follows: Figure 2 As shown, the process can be summarized as follows: First, EEG data is collected to establish an individual baseline. Then, the EEG data is preprocessed and windowed to extract static and dynamic spectral features. The negative emotion index and the emotion fluctuation index are calculated to form a temporal emotional state. Simultaneously, music fragments are modeled using music theory characteristics to construct music characteristic vectors. The expected impact of candidate music on emotional indicators is predicted through a music-emotion mapping model. Finally, music fragments are selected or switched for playback based on a closed-loop objective function. During playback, EEG feedback is continuously acquired, and the individual model and strategy are updated based on the feedback to iteratively form an individual adaptive intervention.
[0039] The key technical solutions of the present invention will be further described below.
[0040] (a) The EEG acquisition unit is used to acquire the EEG signals of the subject under music intervention.
[0041] To balance wearability with coverage of emotion-related brain regions, multi-channel electrodes can be deployed in the frontal or prefrontal cortex, with sampling frequencies set to 128Hz, 250Hz, or higher. Let the raw EEG signal be denoted as: (Equation 1) Where c is the channel index, c=1,2,…,C, C is the number of channels, and t is the time.
[0042] (ii) Preprocessing unit The raw EEG data is processed with bandpass filtering, power frequency notch filtering, baseline correction, and suppression of eye movement and electromyography artifacts to improve the stability of subsequent feature extraction and emotion recognition. The system then segments the continuous EEG data into sliding time windows, with the window length denoted as [missing information]. Step length is recorded as The EEG segment in the k-th window is denoted as: (Equation 2) Where k is the window number. By using sliding windows, the system can output sentiment indicators with a small delay, while taking into account both short-term changes and statistical stability.
[0043] (III) Features and Indicator Units (1) Perform frequency domain or time-frequency domain analysis on the EEG within each window to extract static and dynamic spectral features. Static spectral features are used to describe the steady-state spectral structure within the window, while dynamic spectral features are used to describe the trend of spectral variation between adjacent windows over time. Taking power spectral density as an example, let the Fourier transform of the signal in the k-th window be... Then the power spectral density can be expressed as: (Equation 3) The band power within the frequency band [f1, f2] can be used to reflect the energy distribution of different rhythmic components (such as δ, θ, α, β, γ frequency bands) within this window, and is defined as: (Equation 4) Based on this, static features such as relative power, bandwidth ratio, spectral entropy, spectral edge frequency, and brain region asymmetry can be constructed to enhance the characterization of emotional valence and arousal.
[0044] Dynamic spectral features can be used to characterize information such as energy changes and spectral shape changes between adjacent windows. Taking bandwidth power differential as an example, it can be defined as follows: (Equation 5) The system fuses static and dynamic features to form an EEG emotion feature vector for each window. Let the static feature sub-vector be... The dynamic feature sub-vector is Then the feature vector of the k-th window can be expressed as: (Equation 6) (2) Construction of the Negativity Index and Emotional Fluctuation Index. For example... Figure 3 As shown, this invention quantifies emotions into two complementary dimensions: an emotion negativity index to reflect the intensity or probability of current negative emotions, and an emotion fluctuation index to reflect the degree of emotional fluctuation over time. Through this dual-indicator system, the system can simultaneously perform closed-loop regulation of both "negativity level" and "stability," thus better reflecting the temporal manifestations of depression-related emotional states.
[0045] The Emotion Negativity Index (ENI) can be output by a random forest model. Taking a random forest containing M parallel decision trees as an example, each decision tree outputs a value for the feature vector of the k-th window. Output negative probability The negative emotion index is defined as follows: (Equation 7) in The value ranges from [0,1], with larger values indicating a stronger negative tendency. The random forest model, in addition to its output... In addition, the contribution of different frequency bands, different channels, and static / dynamic features to negative discrimination can be explained by feature importance or tree path rules, thereby improving the interpretability and adjustability of music therapy closed-loop control.
[0046] The Emotion Volatility Index (EVI) is used to quantify the intensity of changes in emotions over time. Preferably, [the EVI is used to]... The fluctuation intensity is represented by the L2 norm of the difference between the state vectors of adjacent windows. (Equation 8) To improve the stability of real-time output and suppress spikes caused by occasional noise, the system can perform exponential smoothing on the EVI to obtain a smoothing fluctuation index. : (Equation 9) Where λ is the smoothing coefficient, 0 < λ < 1. EVI or A higher value indicates more drastic or unstable emotional changes, while a lower value indicates that the emotion is stabilizing or entering a convergent adjustment phase. By combining ENI and EVI, the system can identify different intervention stages such as "high negativity but stable," "low negativity but drastic fluctuations," and "decreasing negativity and converging fluctuations," thus providing more reliable control values for closed-loop music matching.
[0047] (iv) Music Theory Characteristics Unit (1) Perform computational musicological analysis on music fragments to extract controllable and computable music theory characteristics and form music characteristic vectors for subsequent mapping prediction and matching decisions. For example Figure 4 As shown, the system can extract features such as rhythmic tempo and beat stability, modal and scale structure, harmonic complexity and tension, melodic undulation, timbre brightness, and loudness envelope from audio or symbolic musical scores (e.g., MIDI). Let the musical characteristic vector of the i-th musical segment be denoted as... ,but: (Equation 10) Where p represents the music characteristic dimension.
[0048] (v) The music library unit is used to store music segments and their corresponding music feature vectors, and supports searching candidate music sets by feature. To meet the requirements of closed-loop control for smooth transitions, music segments can be further divided into measure-level or paragraph-level segments with natural boundaries, and the connection relationship between segments can be recorded, so that the system can switch without significantly disrupting the auditory continuity.
[0049] (vi) Model Unit (1) Temporal emotion recognition and music-emotion mapping prediction under music intervention. In each window, the negative emotion index is... With the sentiment fluctuation index Combined to form a state vector Music-emotion mapping models are used to predict emotions in the current state. The expected changes in negative emotion and emotional fluctuation after playing candidate music segment i. Let the mapping model parameters be... Then it can be expressed as: (Equation 11) Where ΔENI represents the change in the negative exponent, and ΔEVI represents the change in the volatility exponent. The mapping model is preferably implemented using a regressive random forest: based on the current state vector... With candidate music feature vectors The concatenation is used as input, and the output is the regression prediction values of ΔENI and ΔEVI. The contribution of key music theory features and key EEG features to mood changes can be explained based on the feature importance of the tree model.
[0050] (2) Closed-loop feedback and music matching decision Based on the output of the mapping model, an objective function is constructed, with the comprehensive objectives of "reducing negativity, suppressing fluctuations, maintaining experience continuity, and balancing preferences" to select the most suitable music segment from the candidate music set. For example... Figure 5 As shown, the objective function can be constructed. as follows: (Equation 12) Where α, β, γ, and δ are weighting parameters; D is a smoothing penalty term, used to constrain abrupt changes in tempo, mode, loudness, etc., between adjacent segments; Ω is a preference and safety constraint term, used to constrain user discomfort, overstimulation, or high skipping risk. The system selects the segment that optimizes the objective function. Play: (Equation 13) C represents the candidate set. During playback, the system continuously monitors the EEG and updates ENI and EVI_s. When a significant increase in negativity or a significant increase in fluctuation is detected, a rapid destimulation strategy can be triggered. For example, music segments with more stable rhythms, lower harmonic tension, and gentler loudness are prioritized, and the switching frequency is limited to avoid frequent jumps that could have adverse effects.
[0051] (vii) Adaptive Update Unit The mapping model or strategy is updated online based on actual feedback, allowing the system to gradually learn the individual's sensitivity to different music theory characteristics and the optimal intervention path. For example... Figure 6As shown, the system can calculate the actual changes ΔENI and ΔEVI at the end of each segment or in each evaluation period, and construct a loss function based on the prediction error: (Equation 14) in, and These represent the changes in negative emotion and emotional fluctuations predicted by the mapping model, respectively. When the mapping model is a differentiable model, gradient descent can be used for parameter updates: (Equation 15) To improve individual adaptability, the system can also treat the weight parameter vector θ=[α,β,γ,δ] in the closed-loop decision objective function as learnable parameters, and apply them based on the loss function at the end of each segment or in each evaluation cycle. Update θ online: (Equation 16) When the mapping model uses a non-differentiable model such as random forest, incremental learning, sliding window retraining, or adding an individual calibration layer to the general model can be used for updates to balance real-time performance and stability. Through long-term iteration, the system can form a music-emotion response model that is more suitable for the individual, thereby gradually improving the stability and individual adaptability of the intervention effect in multiple sessions.
[0052] Example: The following is in conjunction with the appendix Figures 1 to 6 The specific embodiments of the present invention will be further described below. It should be noted that these embodiments are only used to explain the technical solutions of the present invention and do not constitute a limitation on the scope of protection; without departing from the spirit of the present invention, those skilled in the art can make equivalent substitutions or modifications to the parameter selection, module combination and implementation method, all of which should fall within the protection scope of the present invention.
[0053] A closed-loop music therapy session implementation process based on EEG emotional fluctuation indicators. This embodiment describes how the operator completes an intervention session after obtaining the system. The system structure is as follows: Figure 1 As shown, it mainly includes an EEG acquisition unit, a preprocessing unit, a feature and index unit, a model unit, a music playback unit, a music theory characteristics unit, a music library unit, an adaptive update unit, and a reporting and log unit. Each unit is connected via wired or wireless communication to form a closed-loop link.
[0054] First, the operator pairs and connects the EEG acquisition unit with the mobile terminal / edge computing device running the client program, and completes session parameter initialization and security parameter settings on the client. The initialization includes at least: number of channels C (preferably 4 to 8 channels), sampling frequency Fs (preferably 128 Hz, 250 Hz, or 500 Hz), bandpass filter bandwidth (preferably 0.5 Hz to 45 Hz), power frequency notch filter (50 Hz or 60 Hz), sliding window length Tw (preferably 2 s to 8 s, more preferably 4 s), and step size Ts (preferably 0.5 s to 2 s, more preferably 1 s). Simultaneously, model unit 4 loads the initial parameters of the pre-trained random forest emotion recognition model and music-emotion mapping model, and reporting and logging unit 9 generates a unique session identifier, records the model version number, and parameter snapshots for this session to meet traceability and reproducibility requirements.
[0055] Then, the operator imports candidate music files into the music library unit 7, and the system performs offline fragmentation and music theory feature modeling to obtain structured music feature vectors. Preferably, the system divides each piece of music into music segments of 20 to 30 seconds, and sets segmentation points at natural pauses, harmonic cadences, or section boundaries to reduce the abruptness of subsequent transitions. Subsequently, the music theory feature unit extracts music feature vectors for each segment. (corresponding to formula (10)), which includes at least: rhythmic speed and beat stability (speed and stability can be estimated by beat detection and autocorrelation / dynamic programming), mode and scale structure (major / minor keys can be identified by chromatic features and tonality template matching), harmonic complexity and harmonic tension (chord recognition and calculation based on tonality distance / tension measurement), melodic fluctuation (pitch profile statistics can be used), timbre brightness (spectral centroid and high-frequency energy ratio can be used as a metric), and loudness envelope (short-time energy or root mean square energy estimation can be used). Music library unit 7 will include fragments, And the connectable relationships between segments are written into the database, and can be based on Build vector indexes (such as KD trees or nearest neighbor indexes) to quickly recall the candidate fragment set later.
[0056] Subsequently, the operator guided the subject to enter a resting state before formal intervention to establish an individual baseline. The EEG acquisition unit 1 continuously acquired resting EEG for 1 to 3 minutes. The preprocessing unit 2 performed bandpass filtering, power frequency notch filtering, baseline correction, and rereference on the raw EEG signal xc(t) (Equation (1)). The artifact suppression algorithm was used to reduce the influence of eye movement and electromyography interference on the spectrum estimation. The artifact suppression can preferably be achieved by using independent component analysis (ICA) to separate and remove artifact components, or by using threshold detection combined with adaptive filtering for suppression. The feature and index unit 3 calculated the mean and variance of features such as relative power and power ratio of each frequency band within the baseline segment and stored them as individual standardized parameters (e.g., z-score standardized parameters) to reduce individual amplitude differences and improve cross-session stability during the online phase. At the same time, it can record electrode impedance and environmental noise levels to adaptively set artifact thresholds and quality thresholds.
[0057] Next, the operator initiates the intervention session, and the music playback unit 5 begins playing the initial music segment. Simultaneously, the EEG acquisition unit 1 acquires EEG data and sends it to the preprocessing unit 2 in real time. The initial segment can be selected based on safety and comfort constraints: preferably a segment with a relatively stable rhythm, low harmonic tension, and a smooth loudness envelope. In the absence of historical session information, the system can use rule-based filtering to determine the initial segment, or the music-emotion mapping model can provide candidate initial segments in the default state and select them according to the objective function.
[0058] Subsequently, preprocessing unit 2 performs sliding windowing on the continuous EEG. The system divides the data according to the window length. With step size The continuous signal is divided into the k-th window signal. (t) (Equation (2)), and performs consistent filtering, notch filtering and artifact suppression on each window to ensure the temporal consistency and comparability of subsequent feature extraction. Through this sliding window mechanism, the system can continuously output sentiment indicators with an update cycle of Ts, achieving low-latency closed-loop control.
[0059] Then, feature and index unit 3 extracts static and dynamic spectral features within each window. Preferably, the system uses Fast Fourier Transform combined with Welch averaging to estimate the power spectral density. (f) (Equation (3)) and calculate the power of each frequency band. , (Equation (4)) to further construct static spectral features This includes, but is not limited to, relative power, bandwidth power ratio, spectral entropy, spectral edge frequencies, and left and right frontal asymmetry; and simultaneously constructs dynamic spectral features. Including the power difference Δ between adjacent window frequency bands , (Equation (5)), spectral flux, and energy change statistics, etc. The system will... and The window feature vector is obtained by splicing and fusion. (Equation (6)), and based on the aforementioned individual baseline parameters, Implement standardization to improve the robustness of model inference.
[0060] Meanwhile, model unit 4 will use the window feature vector Input a random forest model to output a negative sentiment index Preferably, the random forest contains M decision trees (e.g., M=200), each tree outputs a probability for the negative class, and the system averages the probabilities of the outputs of each tree to obtain the probability. (Equation (7), with values ranging from 0 to 1). Based on this, the system constructs an emotion state vector. It contains at least In addition, the summarization volume related to dynamic features is used, and the sentiment fluctuation index is calculated using the L2 norm of the difference between adjacent window state vectors. =|| - ||2 (Equation (8)); To suppress spike fluctuations caused by occasional noise, the system further performs exponential smoothing on the EVI to obtain the smoothed fluctuation index. (Equation (9), where the smoothing coefficient λ satisfies 0 < λ < 1). Through the joint characterization of ENI and EVIs, the system can distinguish between different intervention stages such as "high negativeness but stable" and "low negativeness but drastic fluctuations", providing control for subsequent music matching.
[0061] Furthermore, the system generates a state vector based on the current emotional state. This allows for candidate music retrieval and music-emotion mapping prediction. For example... Figure 5 As shown, the state vector It can be composed of the mean, slope, fluctuation range, and individual preference parameters of the ENI and EVIs from the most recent N windows. The system first performs safety filtering and preference filtering on the music library to remove segments that are too loud, too stimulating, or explicitly disliked by the participants; then, based on the music characteristic vector... Nearest neighbor retrieval is performed to obtain a candidate set C. For each segment i in the candidate set, the system uses the mapping model fφ to calculate the state... The expected change [Δ] after playing the segment. , Δ ]=fφ( , (Equation (11)). The mapping model can preferably be implemented using a regression random forest or a gradient boosting regression tree to maintain prediction stability in small sample and noisy scenarios; and the contribution direction of music theory variables such as "rhythmic stability, harmonic tension, and loudness envelope" to the changes in the index can be explained based on the importance of tree model features.
[0062] Based on this, the system constructs a closed-loop decision objective function according to the mapping prediction results and selects the optimal music segment. As shown in equation (12), the system calculates the objective function J(i|x(k))=α· for candidate segment i. +β· +γ·D+δ·Ω, where D is the smoothing penalty term, used to measure the difference between the currently played segment and the candidate segment in dimensions such as tempo, mode, and loudness envelope (weighted Euclidean distance or weighted absolute difference can be used); Ω is the preference and safety constraint term, used to introduce penalties such as skip risk, adverse feedback, and stimulus cap; α, β, γ, and δ are weight parameters. The system solves i*=argmin_{i∈C}J(i| (Equation (13)) can be used as the next playback segment, and the minimum playback duration can be set. (For example, 10 s to 30 s) and a switching threshold ε (for example, switching is only allowed if the objective function improves by more than ε) are used to avoid excessively frequent switching that could damage the auditory experience. When switching segments, crossfade-in or fade-out or switching at preset natural boundaries can be used to ensure auditory continuity.
[0063] When the system detects that the ENI or EVIs continuously exceed a preset threshold, the system triggers a safety de-stimulation strategy to improve safety and comfort. Preferably, when the ENI is continuously higher than the first threshold for more than T1 seconds, or the EVIs are continuously higher than the second threshold for more than T2 seconds, the system limits the candidate recall range to a "set of low-stimulation segments" and increases the weight of the smoothing penalty term and the safety penalty term in the objective function (e.g., increasing γ and δ), prioritizing segments with more stable rhythm, lower harmonic tension, and smoother loudness, while reducing the switching frequency or limiting the loudness upper limit. The operator can also perform a "one-click stop / de-intensity / skip" interaction through the client. The system records this interaction as a feedback signal and writes it to the log for subsequent strategy updates.
[0064] Subsequently, the system performs feedback calculations and adaptive updates when a segment ends or an evaluation cycle is reached. For example... Figure 6 As shown, the system calculates the actual change in indicators within this period. and And compare it with the predicted value of the mapping model to form the loss function L(φ) (Equation (14)). When the mapping model is a differentiable model, gradient descent can be used to update the parameter φ (Equation (15)); when the mapping model is a non-differentiable model such as random forest, incremental learning or sliding window retraining can be used: , , , The data is added as a new sample to the individual data buffer, and the model is updated or an individual calibration layer is superimposed on the general model according to a preset period, so as to achieve individual adaptation while ensuring real-time performance. The system can also regard the objective function weight vector θ=[α,β,γ,δ] as a learnable parameter, and update θ online based on the loss function (Equation (16)) to gradually learn the individual's trade-off preference between "reducing negativity" and "suppressing fluctuations".
[0065] Finally, when the operator ends the session, the reporting and logging unit 9 generates and outputs an intervention report and audit log. The report includes at least the ENI and EVI curves, music segment playback and switching time points, summaries of the music theory characteristics of each segment, threshold trigger events, and user interaction events; it also saves the original EEG summary of the session, window feature summary, model version number and parameter snapshots, and decision logs to support long-term efficacy tracking, reproduction verification, and auditing. Through the process described in the preceding paragraphs, this embodiment achieves a closed-loop music intervention of "collection—assessment—prediction—decision—playback—feedback—update," enabling operators to complete deployment and use the system according to clearly defined steps after obtaining it. Each step corresponds to a specific algorithm that enables real-time quantification of negative emotions and emotional fluctuations, as well as interpretable closed-loop regulation of music matching.
Claims
1. An intelligent music therapy system based on EEG emotional fluctuation indicators, characterized in that, include: EEG acquisition unit, used to acquire EEG signals from subjects; The preprocessing unit is used to perform window processing on the subject's electroencephalogram (EEG) signals according to a set time window. The feature and index unit is used to: extract the static and dynamic spectral features of the EEG signals within each time window, and combine the two to obtain the window feature vector; and then extract the negative emotion index and emotion fluctuation index of the EEG signals within each time window. The music theory characteristic unit performs computational musicological analysis on music fragments, extracts music theory characteristics, and forms a music characteristic vector. The music library unit is used to store music fragments and their corresponding music feature vectors, and supports searching candidate music sets by feature. Model units are used for: (1) Construct a mapping model that uses the current state vector composed of the combination of the negative emotion index and the emotion fluctuation index as input and the candidate music characteristic vector as input, and outputs the regression prediction values of the change in the negative emotion index and the change in the emotion fluctuation index. (2) Construct an objective function based on the output of the mapping model, and select the most suitable music segment accordingly; The adaptive update unit is used to: construct a loss function based on the subject's feedback on the current music clip and the negative emotion index and emotion fluctuation index predicted by the mapping model, and update the parameters of the mapping model and the weight parameters of the objective function.
2. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 1, characterized in that, The EEG acquisition unit is used to acquire EEG signals from the frontal lobe or prefrontal lobe.
3. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 1, characterized in that, The preprocessing unit performs bandpass filtering, power frequency notch filtering, baseline correction, and suppression of eye movement and electromyography artifacts on the EEG signal.
4. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 1, characterized in that, The static spectral features are used to describe the steady-state spectral structure within a window, while the dynamic spectral features are used to describe the trend of spectral changes over time between adjacent windows.
5. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 1, characterized in that, The negative sentiment index is output by a random forest model; assuming the random forest model contains M parallel decision trees, each decision tree corresponds to the feature vector of the k-th window. Output negative probability The negative emotion index is defined as follows: ; The emotional fluctuation index is obtained by using the L2 norm of the difference between adjacent window state vectors as the fluctuation intensity and then smoothing it.
6. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 1, characterized in that, In the model unit, the objective function as follows: in, represents the state vector of the k-th window; i represents the sequence number of the selected music segment; ΔENI represents the change in the negative exponent, and ΔEVI represents the change in the volatility exponent; D is the smoothing penalty term, used to constrain abrupt changes in tempo, mode, and loudness between adjacent segments; Ω is the preference and safety constraint term, used to constrain user discomfort, overstimulation, or high-risk skipping situations; α, β, γ, and δ are the corresponding weight parameters. Select the music segment that minimizes the objective function value. Play it.
7. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 1, characterized in that, The loss function is expressed as: Wherein, ΔENI represents the change in the negative index, and ΔEVI represents the change in the volatility index; and These represent the changes in negative emotion and emotional fluctuations predicted by the mapping model, respectively.
8. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 7, characterized in that, In the adaptive update unit, the parameter update of the mapping model is represented as follows: ; in, The parameters represent the mapping model. This represents the hyperparameters of gradient descent. Represents the loss function Regarding mapping model parameters The differential.
9. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 7, characterized in that, The weight parameter update of the objective function in the adaptive update unit is expressed as follows: ; in, , These represent the weight parameter vectors in the objective function for the (n+1)th and nth iterations, respectively; This represents the hyperparameters of gradient descent. Indicates the loss of function Regarding weight parameters The differential.
10. The intelligent music therapy system based on EEG emotional fluctuation indicators as described in claim 1, characterized in that, It also includes reporting and logging units for logging, generating assessment reports and audit information.