A generative AI-driven physio-music following response intelligent emotion regulation system

By using a generative AI-driven physiological-music-following response system, combined with multimodal physiological signals and music psychology theory, personalized emotion regulation is achieved, which solves the shortcomings of emotion regulation in existing systems and provides a personalized emotion regulation and relaxation experience.

CN119896792BActive Publication Date: 2025-10-17SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510215895.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-10-17
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

Existing AI music generation systems lack real-time perception and feedback adjustment of individual emotions and physiological states, making it impossible to achieve personalized emotion regulation, and the complex relationships between multimodal physiological signals are not fully utilized.

Method used

The generative AI-driven physiological-music-responsive intelligent emotion regulation system uses multimodal physiological signal acquisition, preprocessing, real-time emotion decoding, and music-emotion response calibration. Combined with the theory of music expectation and tension release, it adaptively generates and adjusts music to achieve precise regulation of user emotions.

Benefits of technology

It achieves personalized and precise regulation of user emotions, guides users through rollercoaster-like emotional changes through music expectation theory, and provides relaxation and satisfaction. It also uses generative AI models to capture long-term dependence on music and the complementarity of physiological signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119896792B_ABST
    Figure CN119896792B_ABST
Patent Text Reader

Abstract

The application discloses a generative AI driven physiological-music following response intelligent emotion regulation system, mainly comprising a physiological signal acquisition module, a signal preprocessing module, an emotion real-time decoding module, a music-emotion response calibration module and a music generation and emotion adjustment module. Based on the generative AI technology, the application fuses music psychology theories such as music expectation and music tension release mode, accurately decodes the real-time emotion state of a user through multi-modal physiological signals, generates music capable of directionally regulating the emotion state of the user, gradually establishes the expectation of the user and releases at a high point, causes the roller coaster-like emotion change of the user, and enables the user to subjectively obtain positive emotions such as pleasure and satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a generative AI-based intelligent emotion regulation technology, belonging to the field of artificial intelligence and signal processing. BACKGROUND

[0002] Music is an art that arranges sounds according to time sequence by using elements such as melody, harmony, rhythm, and timbre. Unlike traditional art forms, music lacks direct form and semantic, so it is limited in expressing the material world. However, music is very successful in stimulating emotions. Composers can effectively stimulate people's emotional reactions by manipulating listeners' expectations.

[0003] According to the theory of music psychology, the results of musical events can be divided into two stages: before the results appear, especially when the results are uncertain, listeners mainly produce preparatory reactions related to stress or tension, and this emotional tension will continue to accumulate before the results appear, making the listener's emotional arousal and attention reach a peak; when the results appear, the listener's expectations are realized, and the emotional tension is released, resulting in strong feelings of relaxation and satisfaction. The principle of this process is similar to that of riding a roller coaster.

[0004] However, individual perception of music has strong heterogeneity, and single, fixed style of music has limited effect. In recent years, generative AI technology has made breakthroughs, enabling music generation models to understand and generate complex music structures, including melody, harmony, rhythm, etc., through deep neural networks and advanced sequence modeling techniques such as Transformer, generative adversarial network (GAN), and variational autoencoder (VAE). However, existing AI music generation systems often rely on static input or preset rules, lacking real-time perception and feedback adjustment of individual emotions and physiological states.

[0005] In the field of physiological measurement, research has found that various physiological parameters, such as respiratory rate, electrocardiogram, and skin electricity, are closely related to emotional state and tension level. As early as the early 20th century, polygraph experiments used physiological indicators such as breathing patterns and blood pressure changes. In recent years, the fusion of multi-modal physiological signals such as electroencephalogram, electrocardiogram, and respiration has shown great potential in affective computing, indicating that the complementarity between different signals can effectively improve the accuracy of emotion recognition. Therefore, it is feasible to fuse multi-modal physiological signals, adjust the attributes of AI-generated music, and then accurately and individually regulate user emotions.

[0006] Although some systems attempt to combine physiological feedback to achieve closed-loop regulation of emotions, these systems do not fully consider the complex relationship between multi-modal physiological signals, and the constraint mechanism is usually missing or relatively simple. In addition, in terms of design principles, existing systems fail to combine more advanced psychological theories to achieve targeted emotion regulation. SUMMARY

[0007] To overcome the deficiencies of the prior art, the present application aims to provide a generative AI-driven physiological-music following response intelligent emotion regulation system, which is applicable to computers and mobile devices, can be realized through data acquisition equipment or wearable acquisition equipment, is based on music psychology theories such as music expectancy and music tension release mode, and helps users obtain a roller coaster-like positive emotion regulation experience, creating an adaptive biological feedback system driven by maximizing user relaxation satisfaction.

[0008] TECHNICAL SOLUTION To achieve the above-mentioned application purpose, the present application adopts the following technical solutions:

[0009] In a first aspect, the present application provides a generative AI-driven physiological-music following response intelligent emotion regulation system, comprising:

[0010] A physiological signal acquisition module for acquiring multi-modal physiological signals of a user during music listening;

[0011] A signal preprocessing module for preprocessing various physiological signals to preliminarily filter out noise in the signals;

[0012] An emotion real-time decoding module for real-time decoding of user emotional state based on a pre-trained emotion feature recognition network;

[0013] A music-emotion response calibration module for generating a music-emotion response calibration curve by detecting the emotional response of an individual to music of different attributes;

[0014] A music generation and emotion adjustment module for generating and adjusting music based on a generative AI model, music expectancy and tension release theory, user emotional state and music-emotion response calibration curve.

[0015] Further, the multi-modal physiological signals include one or more of electroencephalogram signals, electrocardiogram signals, respiratory signals, heart rate signals, body surface temperature, facial expressions and electrodermal activity signals; and the denoising processing of each physiological signal in the signal preprocessing module includes a combination of one or more of high-pass filtering, low-pass filtering and wavelet threshold denoising methods.

[0016] Further, in the emotion real-time decoding module, one or more physiological signals including electroencephalogram signals, electrocardiogram signals, respiration signals, heart rate signals, body surface temperature, facial expressions and electrodermal activity signals are fused to decode the user's emotional state in real time; fine-grained features in one or more long-time complex physiological signals including electroencephalogram signals, electrocardiogram signals, respiration signals and electrodermal activity signals are extracted based on a CNN-LSTM network; deep fusion of multi-modal emotional features is performed through a Transformer, and other simple physiological signals are fused through an MLP to realize real-time decoding of the user's emotional state, and an emotional response E = {e1, e2,..., e s} is obtained, where s is the total dimension number of emotional evaluation.

[0017] Further, in the music-emotion response calibration module, a hierarchical clustering analysis is performed on multiple attributes (such as pitch, melody, harmony, sound intensity, etc.) in a large-scale music data set to determine typical music segments and generate a random test music group M = {F1, F2,..., F k}, where k is the cluster number of clustering; during the playing of the music group, the specific emotional response of the user to the nth music segment F n in the music group M is determined (s is the total dimension number of emotional evaluation), and a music-emotion response calibration curve

[0018] Further, in the music generation and emotion adjustment module, a Transformer model is used to capture the dynamics of music and the time-dependent characteristics between segments, and adaptive constraint conditions are constructed to generate music with consistent style, and embedding features of the music sequence and their position encodings are obtained, where T is the total number of time steps of the music, and d is the dimension of the music feature encoding.

[0019] Further, in the music generation and emotion adjustment module, based on a pre-trained music generation model, the music-emotion response calibration curve L emotion of the user is embedded into the encoding layer of the music generation model to realize the following response of the music attribute-emotion state:

[0020] C emotion = Embedding(L emotion ),

[0021]

[0022] where α = σ(W α [AvgPool(C emotion )]) ∈ (0, 1).

[0023] Cemotion a condition vector for music generation model, (k is the number of clustering clusters, T is the total number of time steps of music) is used to realize the soft alignment of music-emotion response calibration curve to music time step, ⊙ represents element-by-element multiplication, P music represents the music position encoding, C music represents the music sequence embedding feature, P emotion represents the position encoding of the music-emotion response calibration curve, X input is the model input, + represents concatenating features into a vector, ∥ represents merging music information and emotion information into a large input vector, α is an adaptive weight parameter, and σ is a sigmoid function, (T is the total number of time steps of music, and s is the total dimension number of emotion evaluation) is used to realize the weight control of the overall influence strength of the emotion condition.

[0024] Further, the music generation and emotion adjustment module, on the basis of the trained music generation model, generates music capable of directing the user's emotional state according to the user's current emotion E current and the set of previous several emotional states {E1, E2, …, E current} and music psychology theory:

[0025] C dynamic = Embedding(E1, E2, …, E current ),

[0026] p = f adjust (C dynamic , Θ),

[0027]

[0028] Wherein, f adjust is a learnable mapping function constructed based on music psychology theory; Θ is the specific parameter of the trained music generation model; p is the music attribute condition probability distribution generated according to the music psychology theory; is the loss function of generating music, T is the total number of time steps, s t is the expected emotional state at time step t, represents the probability of generating s t at time step t under the given previous state and dynamic condition C dynamic .

[0029] Secondly, the present application provides a physiological-music following response intelligent emotion regulation method driven by generative AI, comprising the following steps:

[0030] Collecting multi-modal physiological signals of a user during listening to music;

[0031] Pretreating various physiological signals to preliminarily filter out noise in the signals;

[0032] Real-time decoding of the emotional state of the user based on the pre-trained physiological signal analysis network;

[0033] Generating a music-emotion response calibration curve by detecting the emotional response of the individual to music of different attributes;

[0034] Adaptive generation and adjustment of music according to the emotional state of the user and the music-emotion response calibration curve based on the generative AI model, music expectation and tension release theory.

[0035] In a third aspect, the present application provides a computer system comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the generative AI driven physiological-music following response intelligent emotional regulation method of the second aspect.

[0036] In a fourth aspect, the present application provides a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of the generative AI driven physiological-music following response intelligent emotional regulation method of the second aspect.

[0037] Beneficial effects: The present application, in terms of design idea, refers to the music expectation theory, combines priori knowledge and data-driven, and proposes an intelligent emotional regulation system of physiological-music following response. Specifically, the music expectation theory is taken as the core theory of regulating emotions, the user's expectation is gradually established through music, and the user's emotions are changed like a roller coaster, so that the user subjectively obtains emotions such as pleasure and satisfaction. In terms of technical implementation, the present application proposes a new music processing and generation mode. Specifically, the music generation model captures the long-time dependence of music through a self-supervised learning module based on a generative AI model such as Transformer, which is used to generate complex melodies, harmonies and rhythms. In addition, the model can establish a constraint relationship between different segments of a music paragraph through the embedded attention mechanism, and generate music with coherent style. The present application refers to the current emotion of the user quantified by multi-modal physiological signals and the previous several emotional states, combines the personalized emotional response label, and if the user's tension state is still on the rise, selects to continue to increase the music tension and accumulate psychological expectation; if the user's tension state stops rising or rises relatively slowly, selects to quickly reduce the music tension and release the psychological expectation.

[0038] In summary, compared with the prior art, the beneficial effects of the present application are: (1) the present application uses music, which is easy to obtain and can bring aesthetic experience, as a stimulus to regulate the user's mood in the form of a passive biofeedback system, helping the user to intuitively experience the process of expected accumulation and release, and obtaining relaxation and satisfaction; (2) guided by the expectation theory of music, based on real-time multi-modal physiological signals and generative AI large models, the corresponding music is adaptively generated and adjusted to realize the precise establishment and release of user expectations. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 It is a schematic diagram of the basic principle of the present application.

[0040] Figure 2 It is a schematic diagram of the system structure of the embodiment of the present application.

[0041] Figure 3 It is a system flowchart of the embodiment of the present application.

[0042] Figure 4 It is a schematic diagram of the emotion recognition network training process in the embodiment of the present application.

[0043] Figure 5 It is a schematic diagram of the music generation model training process in the embodiment of the present application. DETAILED DESCRIPTION

[0044] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings of the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0045] The present embodiment is implemented according to the design of the present application, and is in the form of an automated biofeedback system. The working principle is that the system generates music based on generative AI technology according to the current and previous emotional states of the user, to induce and guide the user's emotions. The regulation of emotions is based on the music expectation theory as the core theory, and the emotions here refer to positive emotions such as stress, tension, arousal and expectation. As shown in Figure 1 The target of the biofeedback system is to gradually increase the emotional intensity to the highest point by letting the user listen to music, and then quickly release these emotions, so that the user can obtain feelings of relaxation and satisfaction. At the same time, the user can further understand the psychological and physiological process of his own relaxation through the interaction module, which helps the user to learn the process of relaxation and judge the effectiveness of the relaxation means.

[0046] The present embodiment realizes a generative AI driven physiological-music following response intelligent emotion regulation system, and the system structure is as followsFigure 2 As shown, the system consists of a sound output module, a physiological signal acquisition module, a signal preprocessing module, a real-time emotion decoding module, a music-emotion response calibration module, a music generation and emotion adjustment module, and a real-time interaction and storage module.

[0047] like Figure 3 As shown, the working principle of this embodiment is as follows: the sound output module interacts directly with the user and is used for transmission, conversion and playback of music signals; the physiological signal acquisition module is used to collect multimodal physiological signals of the user in the process of listening to music in real time, including but not limited to EEG signals, ECG signals, respiratory signals, heart rate signals, body surface temperature, facial expressions and skin electrical activity signals, or other types of contact or non-contact physiological signals that can represent emotional states, and transmit the data to the signal preprocessing module; the signal preprocessing module performs denoising processing on various physiological signals such as EEG signals, ECG signals, respiratory signals, body surface temperature and skin electrical activity signals transmitted from the physiological signal acquisition module, so as to reduce the interference of signal noise on subsequent processing and analysis, and preliminarily filter The de-noised signal is transmitted to the real-time emotion decoding module. The real-time emotion decoding module, with a multimodal deep learning network at its core, uses a large-scale emotion recognition physiological database for pre-training. Based on the pre-trained emotion feature recognition network, it integrates various physiological signals to decode the user's emotional state in real time. The music-emotion response calibration module detects an individual's emotional response to music of different attributes and generates a music-emotion response calibration curve for specific musical attributes, providing a basis for music-based automatic emotion regulation methods. The music generation and emotion adjustment module, based on a generative AI model and music expectation and tension release theory, adaptively generates and adjusts music according to the user's emotional state and the music-emotion response calibration curve to achieve a cumulative increase in the user's emotional intensity or release satisfaction. The real-time interaction and storage module is used to display and store physiological signal waveforms and signal quality in real time, user emotional state assessment results in real time, and music parameter adjustments in real time.

[0048] The sound output module converts the music electrical signal into vibration through a transducer. In some embodiments, the transducer can be a speaker built into a smart device (smartphone, computer, watch, AR or VR device), a wired speaker, a wireless speaker, a wired headset, a wireless headset, a bone conduction headset, a vibration sound transmission device, a hearing aid, or any other device or apparatus that can interact with the user to generate acoustic information.

[0049] The physiological signal acquisition module acquires physiological signals of a user in real time through a sensing device, including but not limited to electroencephalogram signals, electrocardiogram signals, respiration signals, heart rate signals, body surface temperature, electrodermal activity signals, or any one or combination of other physiological signals that can represent emotional states. In some embodiments, for a certain type of physiological signal, single-point single-form measurement or multi-point multi-form combined measurement can be used. The electroencephalogram signal acquisition device can be any device of arbitrary form or measurement principle, such as an electroencephalogram cap, a portable electroencephalogram headband, a distributed electrode patch, an ear clip, etc. The electrocardiogram signal, heart rate signal, and respiration signal acquisition device can be any device of arbitrary form or measurement principle, such as a distributed electrode patch, smart fabric, a smart bracelet, a smart watch, a chest strap, a waistband, a millimeter wave radar, etc. The body surface temperature acquisition device can be any device of arbitrary form or measurement principle, such as a thermistor, a thermocouple, an infrared sensor, a MEMS temperature sensor, etc. The electrodermal activity (EDA) can be any device of arbitrary form or measurement principle, such as a bracelet (e.g., Empatica E4 bracelet), a finger clip, a patch, a chest strap, a glove, an ear clip, smart fabric, etc.

[0050] The signal preprocessing module is used to filter out electromagnetic noise and physiological artifacts in the acquisition process to achieve accurate emotional decoding. In some embodiments, high-pass filtering and low-pass filtering can be performed on each physiological signal to filter out baseline, power frequency interference, and high-frequency noise, and to retain signals in the desired frequency band. If necessary, wavelet threshold denoising is used to filter out other physiological artifacts. The module can be implemented based on any programming language, such as C, matlab, python, java, etc. The specific preprocessing operations performed on each physiological signal can be:

[0051] (1) Electroencephalogram signal: first, perform 0.5Hz high-pass filtering and 80Hz low-pass filtering on the electroencephalogram signal to filter out baseline drift and high-frequency noise in the signal; then perform band-stop filtering at 50Hz to filter out power frequency noise; finally, perform wavelet threshold denoising to remove electrooculogram and electromyogram artifacts in the electroencephalogram signal.

[0052] (2) Electrocardiogram signal and respiration signal: first, perform band-pass filtering on the electrocardiogram signal and respiration signal at 3-40Hz, and then perform band-stop filtering at 50Hz.

[0053] (3) Electrodermal activity and body surface temperature: mainly perform smoothing filtering to filter out high-frequency noise, with a window length of 3 points.

[0054] The filters used in the signal preprocessing module can all be FIR digital filters, which perform double-pass filtering to achieve zero-phase shift filtering by inputting the signal forward and reverse into the digital filter.

[0055] The emotion real-time decoding module performs real-time analysis on the pre-processed multi-modal physiological signals based on a pre-trained emotion feature recognition network model, to realize accurate decoding of the user's emotional state (including tension, arousal, valence, and other indicators). In some embodiments, model pre-training can use public physiological signal emotion databases (such as DEAP, AMIGOS, DREAMER, etc.) and self-built databases as training data sources.

[0056] Preferably, the emotion real-time decoding module extracts fine-grained features in one or more long-time complex physiological signals including electroencephalogram signals, electrocardiogram signals, respiration signals, and electrodermal activity signals based on a CNN-LSTM network; performs deep fusion of multi-modal emotion features through a Transformer, and fuses other simple physiological signals through an MLP, to realize real-time decoding of the user's emotional state. As shown in Figure 4 , the specific training process of this module in one preferred embodiment can be:

[0057] (1) Segment the input signal with a 1-second window and 50% overlap to generate data frames x of fixed length N t,w , and perform Z standardization on each data frame:

[0058]

[0059] where x t,w is the vector form of the original data (e.g., time series signal of EEG), μ w is the mean of the time series x t,w , and σ w is the standard deviation of the time series x t,w .

[0060] (2) Use a convolutional neural network (CNN) to extract features from multi-modal complex physiological signals such as electroencephalogram, electrocardiogram, respiration, and electrodermal activity. Specifically, the convolution kernel size of the CNN layer can be set to 3x3, with 32-64 filters per layer, ReLU activation, and Batch Normalization to maintain training stability and accelerate convergence. Each signal branch outputs a 128-dimensional feature vector.

[0061] (3) Use a long short-term memory network (LSTM) to capture the time sequence dependence of complex physiological signals such as electroencephalogram, electrocardiogram, respiration, and electrodermal activity. The LSTM unit has 128 hidden units per layer, and Dropout is set to 0.3 to prevent overfitting.

[0062] (4) In terms of multi-modal fusion, firstly, the spatio-temporal feature vectors Feature_dim of all complex physiological signals N_modalities are spliced into a multi-modal feature matrix (N_modalities, Feature_dim). Subsequently, the multi-modal features are weighted and distributed using the Transformer architecture to capture the complex dependence between different signals, grasp the correlation and complementarity between different modal signals, and achieve accurate evaluation of user tension, arousal, valence and other emotional indicators. Specifically, the number of attention heads is set to 4, each head has a dimension of 32, the inter-modal dependence relationship is modeled, and the attention weight of each modality is calculated; the feedforward neural network uses 2 layers of fully connected layers (MLP), the hidden layer is 128 dimensions, and the activation function is ReLU; Layer Normalization is used to maintain the stability of the network.

[0063] (5) Finally, the body temperature data is further included, and all physiological signal derived feature vectors are passed through a layer of fully connected layer to output a multi-dimensional feature vector E = {e1, e2, …, e s} related to emotion. Wherein s is the total dimension number of emotional evaluation.

[0064] (6) In terms of model training and optimization, the large-scale emotional data set is divided into a training set of 80%, a validation set of 10%, and a test set of 10%, the optimizer uses the Adam optimization algorithm, the initial learning rate is set to 0.001, and the learning rate scheduler is used for dynamic adjustment; the loss function uses mean square error (MSE) loss.

[0065] In some embodiments, the music-emotion response calibration module determines the real-time emotional response of the user to different types of music segments by traversing typical music segments that can mobilize the user's emotions in the Lakh MIDI Dataset and other music databases and self-built databases, and uses it as a multi-dimensional fine-grained emotional label of the same type or specific attribute music. The specific processing flow can be as follows:

[0066] (1) For music segments M = {M1, M2, …, M N} from a large-scale music database, N is the total number of music segments, and each segment M i contains the following music attributes:

[0067] f music,i = {f pitch ,f rhythm ,f harmony ,f loudness ,f duration},

[0068] Where f pitch is the pitch attribute, f rhythmFor melody attribute, f harmony For harmony attribute, f loudness For intensity attribute, f duration For duration.

[0069] (2) In order to eliminate the dimensional differences of different features, Z standardization is used to process each feature, and for the ith music segment:

[0070]

[0071] Wherein, μ j and σ j are the mean and standard deviation of the jth music feature, respectively.

[0072] (3) According to the hierarchical clustering analysis of music attributes, the average linkage is used for the aggregation strategy, and k clusters (classes) are obtained, and for the nth cluster P n , according to the distance minimum measurement principle, the music m j ∈P n in the cluster is searched to obtain the typical music closest to the center point attribute:

[0073]

[0074] Wherein, d(m j ,P n ) represents the distance between m j and the cluster P n ; a test music group M = {F1, F2, …, F k} is generated, and k is the number of clustering clusters.

[0075] The test music group M is traversed, the specific emotional response I n of the user to the nth typical music segment F n is determined, and a music-emotional response calibration curve is constructed:

[0076]

[0077] Wherein, k is the number of clustering clusters, and s is the total dimension number of emotional evaluation.

[0078] The music generation and emotion adjustment module takes the music expectation theory as the core of adjusting emotion, and based on the pre-trained music generation model, through the joint analysis of the current and past emotional states of the user and the individualized music-emotional response label, the music that can significantly mobilize the emotional state of the user is generated, and the user expectation is established to cause the roller coaster type emotional change of the user. Specifically, in some embodiments, as shown in Figure 5 , the module is divided into three processes of upstream, midstream and downstream. Among them,

[0079] (1) The upstream task mainly involves music generation. The music dynamics and temporal dependencies between segments can be captured based on the Transformer model, and adaptive constraints can be constructed to generate music with consistent style. For example, the process parses the music MIDI file into discrete tokens, encodes the music event sequence into embedding vectors, and based on the Transformer model and the embedded multi-head attention mechanism, it self-supervises the capture and learning of the short and long-range dynamics and generation rules of music attributes in music databases such as Lakh MIDI Dataset and self-built databases, to achieve continuous generation, real-time adjustment and smooth transition of music, and obtain embedding features of music sequences and their position encodings where T is the total number of time steps of music, and d is the dimension of music feature encoding.

[0080] (2) The midstream task is used to build a precise mapping between music and emotion. Based on the music generation Transformer model trained in the upstream task, the individualized music-emotion response calibration curve L emotion is input as a supplement, and a conditional vector C emotion is added emotion (Embedding operation converts input data L music into a low-dimensional dense vector representation so that the Transformer model can perform subsequent sequence modeling), which is dynamically aligned with the music position encoding P music , and embedded into the Transformer encoding layer of the upstream model:

[0081]

[0082] α=σ(W α [AvgPool(C emotion )])∈(0,1).

[0083] where LayerNorm is a layer normalization operation used to normalize and aggregate features in space; (k is the number of clustering clusters, and T is the total number of time steps of music) is used to achieve soft alignment of the music-emotion response calibration curve to the music time step, ⊙ represents element-wise multiplication, P emotion represents the position encoding of the music-emotion response calibration curve, X input is the model input, + represents concatenating features into a vector, || represents combining music information and emotion information into a large input vector, α is an adaptive weight parameter used to prevent emotion conditions from dominating too much; AvgPool is an average pooling operation used to extract the global emotion intensity in the time dimension; σ is the Sigmoid function, (T is the total number of time steps of music, s is the total dimension number of emotion evaluation) is used to realize the weight control of the overall influence strength of the emotional condition.

[0084] This task enables the model to adaptively analyze the music (group) that mobilizes the user's emotions, and fine-tune the model based on the analysis result to achieve precise follow-up response of music attributes-emotion state.

[0085] (3) The downstream task analyzes the user's emotional state based on music psychology theory, and determines the adjustment direction of the music. This process takes the user's current emotion E current and the set of previous several emotional states {E1, E2, …, E current} transmitted by the emotion real-time decoding module as input, and obtains the control vector C dynamic of the Transformer model through Embedding operation:

[0086] C dynamic = Embedding (E1, E2, …, E current ).

[0087] Subsequently, based on the model constructed in steps (1) and (2), the model parameters are adaptively adjusted according to the music psychology theory to generate music that can direct the user's emotional state:

[0088] p = f adjust (C dynamic , Θ),

[0089]

[0090] Where f adjust is a learnable mapping function constructed based on music psychology theory, and its main functions and implementation effects are shown in Figure 1 , specifically: when it is determined that the user's tension state is still on the rise, fine-tune the model output to continue accumulating expectations for the corresponding music; when the user's tension state stops rising or rises relatively slowly, convert the music generation parameters to release the user's expectations; Θ is the specific parameters of the Transformer model obtained after training in steps (1) and (2); p is the music attribute condition probability distribution generated according to the music psychology theory; is the loss function of generating music, T is the total number of time steps, s t is the emotional state of expectation at time step t, specifically, represents the probability of generating s t at time step t under the given previous state and dynamic condition C dynamic .

[0091] The real-time interaction and storage module includes a user interaction and operation interface, and a storage medium for saving physiological data and music parameters generated during use. The functions implemented by the module include real-time display and storage of physiological signal waveforms and signal quality, real-time display and storage of user emotional state evaluation results, and real-time display and storage of music parameter adjustment conditions. Specifically, the module can be implemented based on various smart terminals such as mobile phones, tablets, computers, watches, AR or VR headsets, etc., or can be configured in a remote cloud server.

[0092] The embodiment of the present application also discloses a generative AI driven physiological-music following response intelligent emotional regulation method, comprising the following steps: collecting multi-modal physiological signals of a user during listening to music; pre-processing various physiological signals to preliminarily filter out noise in the signals; decoding the emotional state of the user in real time based on a pre-trained physiological signal analysis network; generating a music-emotion response calibration curve by detecting the emotional response of the individual to music of different attributes; and generating and adjusting music according to the emotional state of the user and the music-emotion response calibration curve based on a generative AI model, music expectation and tension release theory. For specific implementation details of each step, refer to the implementation of the system module of the aforementioned embodiments, which will not be repeated here.

[0093] The embodiment of the present application also discloses a computer system comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the computer program is executed by the processor to implement the steps of a generative AI driven physiological-music following response intelligent emotional regulation method.

[0094] The embodiment of the present application also discloses a computer program product comprising a computer program, wherein the computer program is executed by the processor to implement the steps of a generative AI driven physiological-music following response intelligent emotional regulation method.

[0095] The embodiments are only for illustrating the technical idea of the present application, and cannot limit the protection scope of the present application. Any modification made according to the technical idea of the present application on the basis of the technical solution falls within the protection scope of the present application.

Claims

1. A generative AI-driven physiological-music-response intelligent emotion regulation system, characterized by: include: Physiological signal acquisition module, used to collect multimodal physiological signals of users while listening to music; The signal preprocessing module is used to preprocess various physiological signals and initially filter out noise in the signals; The real-time emotion decoding module is used to decode the user's emotional state in real time based on the pre-trained emotion feature recognition network; The music-emotion response calibration module is used to generate a music-emotion response calibration curve by detecting the individual's emotional response to music with different attributes; hierarchical clustering analysis of multiple attributes in a large-scale music dataset is used to determine typical music clips and generate random test music groups. ,in is the number of clusters; in the process of playing the music group, the user's Middle Music clips Specific emotional response , constructing music-emotion response calibration curve ,in is the total number of dimensions for emotion evaluation; The music generation and emotion adjustment module is used to adaptively generate and adjust music according to the user's emotional state and the music-emotion response calibration curve based on the generative AI model, music expectation and tension release theory; based on the pre-trained music generation model, the user's music-emotion response calibration curve is used to generate and adjust music. Embedded into the encoding layer of the music generation model to achieve follow-up response of music attribute-emotional state: ; ; ; in, , is the conditional vector for the music generation model, Used to achieve soft alignment of music-emotion response calibration curves to music time steps, represents element-wise multiplication, is the total number of time steps of the music, Indicates the music position code, represents the music sequence embedding feature, Represents the positional encoding of the music-emotion response calibration curve, is the model input, Indicates concatenating features into vectors, Indicates combining music information and emotional information into a large input vector, is the adaptive weight parameter, is the Sigmoid function, Used to achieve weight control of the overall impact strength of emotional conditions; Based on the trained music generation model, according to the user's current mood and the collection of several previous emotional states Based on the theory of music psychology, the model parameters are adaptively adjusted to generate music that can specifically mobilize the user's emotional state: ; ; ; in, It is a learnable mapping function based on music psychology theory; The specific parameters of the music generation model obtained by training; is the conditional probability distribution of music attributes generated according to music psychology theory; is the loss function for generating music, is the time step Expecting emotional state, Represents a given previous state and dynamic conditions Next, at time step generate probability.

2. A generative AI-driven physiological-music-response intelligent emotion regulation system according to claim 1, characterized in that: The multimodal physiological signals include one or more of EEG signals, ECG signals, respiratory signals, heart rate signals, body surface temperature, facial expressions and skin electrical activity signals; the denoising processing of each physiological signal in the signal preprocessing module includes a combination of one or more of high-pass filtering, low-pass filtering and wavelet threshold denoising methods.

3. The generative AI-driven physiological-music-response intelligent emotion regulation system according to claim 1, characterized in that: The real-time emotion decoding module integrates one or more physiological signals including EEG signals, ECG signals, respiratory signals, heart rate signals, body surface temperature, facial expressions and skin electrode activity signals to decode the user's emotional state in real time; based on the CNN-LSTM network, fine-grained features are extracted from one or more long-term complex physiological signals including EEG signals, ECG signals, respiratory signals and skin electrode activity signals; Through the deep fusion of multimodal emotional features through Transformer and the fusion of other simple physiological signals through MLP, the real-time decoding of the user's emotional state is achieved to obtain the emotional response. .

4. The generative AI-driven physiological-music-response intelligent emotion regulation system according to claim 1, characterized in that: The music generation and emotion adjustment module captures the dynamic properties of music and the temporal dependencies between segments based on the Transformer model, adaptively constructs constraints, and generates music with coherent style.

5. A generative AI-driven physiological-music-response intelligent emotion regulation method, characterized by: The steps include: Collect multimodal physiological signals of users while listening to music; Pre-process various physiological signals and initially filter out noise in the signals; Decode the user's emotional state in real time based on the pre-trained physiological signal analysis network; By detecting the emotional response of individuals to music of different attributes, a music-emotion response calibration curve is generated; hierarchical clustering analysis of multiple attributes in a large-scale music data set is performed to determine typical music clips and generate random test music groups. ,in is the number of clusters; in the process of playing the music group, the user's Middle Music clips Specific emotional response , constructing music-emotion response calibration curve ,in is the total number of dimensions for emotion evaluation; Based on the generative AI model, music expectation and tension release theory, music is adaptively generated and adjusted according to the user's emotional state and the music-emotion response calibration curve; based on the pre-trained music generation model, the user's music-emotion response calibration curve is Embedded into the encoding layer of the music generation model to achieve follow-up response of music attribute-emotional state: ; ; ; in, , is the conditional vector for the music generation model, Used to achieve soft alignment of music-emotion response calibration curves to music time steps, represents element-wise multiplication, is the total number of time steps of the music, Indicates the music position code, represents the music sequence embedding feature, Represents the positional encoding of the music-emotion response calibration curve, is the model input, Indicates concatenating features into vectors, Indicates combining music information and emotional information into a large input vector, is the adaptive weight parameter, is the Sigmoid function, Used to achieve weight control of the overall impact strength of emotional conditions; Based on the trained music generation model, according to the user's current mood and the collection of several previous emotional states Based on the theory of music psychology, the model parameters are adaptively adjusted to generate music that can specifically mobilize the user's emotional state: ; ; ; in, It is a learnable mapping function based on music psychology theory; The specific parameters of the music generation model obtained by training; is the conditional probability distribution of music attributes generated according to music psychology theory; is the loss function for generating music, is the time step Expecting emotional state, Represents a given previous state and dynamic conditions Next, at time step generate probability.

6. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the generative AI-driven physiological-music-response intelligent emotion regulation method according to claim 5 are implemented.

7. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the generative AI-driven physiological-music-response intelligent emotion regulation method according to claim 5 are implemented.