Emotional Music Adjustment Method and Terminal Based on Dynamic Perception and Multimodal Fusion
Through dynamic perceptual field and multimodal fusion technology, the problems of insufficient user feature modeling and inaccurate emotional mapping in existing music therapy are solved, and the stability of personalized music generation and continuous optimization of treatment effects are achieved, improving the user experience.
Patent Information
- Application Number
- CN202510600232.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing music therapy technology has insufficient user feature modeling, low emotional-music mapping accuracy and weak control mechanism of the treatment process, resulting in poor personalized treatment effects.
A emotional music adjustment method based on dynamic perception and multimodal fusion is constructed. Through dynamic perception field, emotion-music bridge adapter and adaptive feedback control, the full process of intelligent processing from user feature capture to emotional music generation is realized, including multi-dimensional feature data processing, facial emotional feature mapping and music generation optimization.
It significantly improves the comprehensiveness of feature perception and the accuracy of emotional mapping, realizes the stability of personalized music generation and continuous optimization of therapeutic effects, and improves user experience and treatment safety.
Smart Images

Figure CN120114726B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and image processing, and in particular to an emotional music adjustment method and terminal based on dynamic perception and multimodal fusion. Background Art
[0002] Music therapy, as a non-invasive psychological intervention, influences the body's nervous and endocrine systems through adjustable musical sound waves, thereby regulating emotions and improving physiological states. Studies have shown that music can activate the prefrontal cortex, amygdala, and other brain regions associated with emotion regulation, promoting the secretion of neurotransmitters such as dopamine and serotonin, thereby influencing the body's autonomic nervous and endocrine systems and improving physiological indicators such as heart rate variability and cortisol levels.
[0003] At present, music therapy technology mainly includes the following three technical paradigms: (1) Traditional rule-based music therapy mainly relies on the therapist's experience to select appropriate music clips and perform manual intervention. Although it has strong pertinence, it is inefficient and difficult to apply on a large scale. (2) Intelligent recommendation systems based on machine learning use collaborative filtering and content-based recommendation methods, combined with user profiles for music matching, which improves the degree of system automation, but is often limited to the selection of existing music libraries, making it difficult to achieve truly personalized treatment. (3) Generative therapy systems based on deep learning use models such as LSTM and Transformer to achieve music generation and emotion mapping. Although they have certain creativity, they still need to be improved in terms of music quality and controllability of treatment effects.
[0004] With the development of artificial intelligence (AI), music therapy systems have made significant progress in feature extraction, emotion mapping, and feedback control. Researchers have begun experimenting with introducing attention mechanisms to achieve adaptive adjustment of feature weights, using variational autoencoders to map emotion to musical features, and optimizing treatment strategies through reinforcement learning. However, existing technologies still have shortcomings in the following areas:
[0005] 1. Inadequate User Feature Modeling: Existing systems often use static or fixed-weight feature extraction methods, making it difficult to effectively perceive dynamic changes in user status. Although attention mechanisms have been introduced to enhance feature selection, the lack of modeling of deep inter-feature connections can still lead to inaccurate judgments of feature importance. Furthermore, existing systems fail to consider users' multidimensional background characteristics (such as personal experience, cultural background, and emotional tendencies), which hinders the effectiveness of personalized treatment.
[0006] 2. Low emotion-music mapping accuracy: Existing methods generally use a single mapping model to convert user features into music generation parameters, which makes it difficult to characterize complex nonlinear mapping relationships. Although some studies have attempted to transform the feature space using variational autoencoders, the lack of effective semantic constraints and music quality control mechanisms has resulted in poor performance in terms of emotional expression accuracy and artistic quality of the generated music.
[0007] 3. Weak control mechanisms for the treatment process: Existing systems often rely on threshold judgments to adjust music parameters, lacking a systematic control strategy. While reinforcement learning methods have been introduced to optimize the control process, imperfect feedback mechanisms and a lack of a multimodal signal-based evaluation system make it difficult to accurately assess and continuously optimize treatment outcomes. Furthermore, the system lacks an effective mechanism for accumulating, extracting, and transferring treatment experience, limiting its overall adaptability and intelligence.
[0008] Therefore, there is an urgent need for a personalized music generation solution that can fully perceive the user's multi-dimensional state characteristics, accurately establish the mapping relationship between emotions and music, and have perfect feedback control capabilities to improve treatment effects and user experience. Summary of the Invention
[0009] In order to solve the technical problems existing in the prior art, the present invention provides an emotional music adjustment method and terminal based on dynamic perception and multimodal fusion. The present invention constructs a complete end-to-end generation framework, which mainly includes a user feature input layer, a dynamic perception field, an emotion-music bridging adapter, a music generation module and dynamic feedback. Through the organic connection formed by the feature flow, mapping flow and feedback flow, the whole process from user feature capture to emotional music generation is realized, thereby improving the treatment effect and user experience.
[0010] To achieve the above object, the present invention provides the following technical solutions:
[0011] The present invention discloses an emotional music adjustment method based on dynamic perception and multimodal fusion, comprising the following steps:
[0012] S1. Obtain multi-dimensional characteristic data of users, including user profile, personal preferences, historical experience and physiological data;
[0013] S2. Dynamically assigning weights to the multi-dimensional feature data through a dynamic perception field to generate fusion features;
[0014] S3 real-time collection of user facial image data, facial emotion features extracted by the ViT model, through the emotion - music bridge adapter facial emotion feature space mapped to the music feature space; the facial emotion features include expression coding, emotion dimension, emotion intensity and temporal characteristics;
[0015] S4. Inputting the fused features and the facial emotion features mapped to the music feature space into the music generation module, and generating personalized therapeutic music through feature fusion encoding, music structure planning, and hierarchical decoding architecture in the module;
[0016] S5. Monitor the user's facial emotional features and generate music features in real time, adjust the music generation parameters through a dynamic feedback mechanism, and form a closed-loop control.
[0017] As a further improvement of the above scheme, in step S1, four LlaMA encoders are used to process the text descriptions of the four dimensions of user portrait, personal hobbies, historical experience and physiological data respectively, and convert the text descriptions into standardized encoded feature vectors, thereby forming the multi-dimensional feature data.
[0018] As a further improvement of the above solution, step S2 includes the following specific steps:
[0019] S21. Concatenate the four types of coded feature vectors: user profile, personal hobbies, historical experience, and physiological data, to construct a unified feature representation space and achieve preliminary feature fusion. The expression formula is:
[0020] ;
[0021] Where, E is the fused feature vector with a dimension of n×4, where n is the dimension of the single-class feature vector; 、 、 and These are the encoding feature vectors of user portrait, personal hobbies, historical experiences, and physiological data respectively;
[0022] S22. Construct a dynamic perception field to achieve dynamic perception of feature importance;
[0023] Among them, the weighted module in the dynamic perception field is used to initialize the weights of the four types of features, evenly distribute the feature weights in each dimension, and the weight set W Defined as:
[0024] ;
[0025] Where, The weight coefficients corresponding to the four types of features are: , , ;
[0026] The expression formula of the perception field intensity function of the dynamic perception field is as follows:
[0027] ;
[0028] Where, is the perceptual field strength function; t For time; and are direct impact terms and interactive impact terms respectively, and the expression formula is:
[0029] ;
[0030] ;
[0031] Where, is the time decay factor; ; For feature correlation, cosine similarity is used to quantify the encoded feature vector and The degree of correlation between them ranges from [0,1]. The larger the value, the stronger the feature correlation.
[0032] The dynamic coupling coefficient is the gradient information used to characterize the direction of the feature's influence on the perception field, the gradient product is used to reflect the degree of coordination of feature changes, and the sigmoid function is used to normalize the value range of the dynamic coupling coefficient to [0, 1].
[0033] S23. Update the feature weights using an improved momentum gradient descent method, and adaptively fuse the multi-dimensional features using the updated feature weights to generate fused features. The weight update formula is:
[0034] ;
[0035] Where, and are the feature weights before and after updating respectively; is the learning rate; is the momentum factor; is the second-order gradient, is the first-order gradient.
[0036] As a further improvement of the above scheme, in step S3, the encoder of the ViT model receives the facial image sequence as input and outputs the encoded facial emotion feature vector. The expression formula of the feature extraction process is:
[0037] ;
[0038] Where, is the facial emotion feature vector; Represents the encoding process of the ViT encoder; is a facial image sequence; is the first facial emotion feature vector feature components, each of which represents a specific aspect in the emotional feature space, together forming a complete representation of the four dimensions of expression coding, emotional dimension, emotional intensity and temporal characteristics. , m is the total number of characteristic components.
[0039] As a further improvement of the above solution, in step S3, the emotion-music bridge adapter adopts a dual-channel design, including two processing channels: uplink adaptation and downlink adaptation;
[0040] Among them, the uplink adaptation channel is used to convert the facial emotion features into the music feature space through the projection layer, ReLU activation, and layer normalization. The expression formula is:
[0041] ;
[0042] Where, Feature representation for uplink adaptation channel output; Normalization of the representation layer; ReLU activation processing; is the uplink adaptation channel projection weight matrix, , is the music feature space dimension, It is the dimension of facial emotion characteristics; Adapt channel bias parameters for uplink; is the set of real numbers;
[0043] The downlink adaptation channel is used to compress the music feature dimensions through layer normalization, projection layer, and gating unit, and extract the core features related to emotion. The expression formula is:
[0044] ;
[0045] Where, Feature representation for downlink adaptation channel output; It is processed by gated units to selectively retain key features; is the representation of the music feature space; is the downlink adaptation channel projection weight matrix, ; Adapt channel bias parameters for downlink;
[0046] The uplink adaptation channel and the downlink adaptation channel work together and are integrated through residual connections. The expression formula is:
[0047] ;
[0048] Where, is a bridging feature;
[0049] As a further improvement to the above solution, in step S3, the emotion-music bridge adapter adopts the following optimization and constraint mechanism:
[0050] (1) Computational efficiency goal: Minimize the delay introduced by the emotion-music bridge adapter, that is:
[0051] ;
[0052] In the formula, the subscript A set of trainable parameters for the emotion-music bridge adapter, including upstream channel parameters , downlink channel parameters and internal parameters of the gate control unit; represents the expected value operator, which is used to calculate the average performance of a random variable; Delay time introduced for the Emotion-Music Bridge Adapter;
[0053] (2) Emotional mapping accuracy goal: maximize the consistency between emotional expression and musical characteristics, that is:
[0054] ;
[0055] Where, Represents facial emotion feature vector Music parameter characteristics Semantic consistency measurement of
[0056] (3) Music quality goal: maximize the music quality under the delay constraint, that is:
[0057] ;
[0058] Where, To generate music quality metrics; is the upper threshold of the delay time;
[0059] Among them, music quality measurement It is defined as a multi-factor comprehensive score, and the calculation formula is:
[0060] ;
[0061] Where, is the weight coefficient; Scoring the musical structural integrity, a measure of the logic of musical form; Scoring melodic fluency, which assesses melodic naturalness and memorability; Scoring harmonic coordination to evaluate the balance and richness of harmonic progressions; Scoring emotional expressiveness measures how well the music matches the target emotion;
[0062] The emotion-music bridge adapter adopts the following multi-objective balancing strategy:
[0063] (1) Hierarchical priority strategy, dynamically adjust the optimization target weight according to user scenario requirements:
[0064] ;
[0065] Where, Optimize the objective function for the whole; Indicates the A single optimization goal; A set of parameters for the emotion-music bridge adapter; is the context-related weight coefficient, ; The current interaction scene type;
[0066] When in a real-time treatment scenario, > > , giving priority to ensuring computational efficiency;
[0067] When quality is the priority, > > , giving priority to music quality;
[0068] When in the emotional response field, > > ,give priority to ensuring the accuracy of emotion mapping;
[0069] (2) Constraint relaxation strategy allows some constraints to be softened in extreme scenarios. The expression formula is:
[0070] ;
[0071] Where, is a soft constraint penalty term; is the constraint relaxation coefficient; For the Constraints, , corresponding to the three constraints of computational efficiency, emotion mapping accuracy, and music quality;
[0072] When the system load exceeds the preset load threshold hour, Decrease, relax non-critical constraints;
[0073] When the emotional change exceeds the threshold hour, Incrementally, strengthen the emotion mapping constraints.
[0074] As a further improvement of the above solution, in step S4, the method for generating personalized therapeutic music includes the following specific steps:
[0075] S41. Using a feature fusion encoder, integrating the fusion feature and the facial emotion feature, and constructing a basic feature representation for music generation;
[0076] The feature fusion encoder is composed of N identical computational blocks stacked together. Each computational block includes: residual connections and layer normalization operations, a feedforward neural network layer, and a multi-head attention layer. The residual connections and layer normalization operations are used to implement feature normalization and residual connections; the feedforward neural network layer is used to improve feature representation capabilities through nonlinear transformations; and the multi-head attention layer is used to capture long-range dependencies between features.
[0077] The attention mechanism is used to inject facial emotion features between the encoding layers of the feature fusion encoder to achieve progressive fusion of emotion features. The attention mechanism obtains the query vector Q, key vector K, and value vector V through three independent linear transformations. The calculation formula is:
[0078] ;
[0079] ;
[0080] ;
[0081] Where, For the l The hidden state of the encoding layer, are the weight matrices for query, key, and value respectively, for The facial emotion feature representation obtained after processing through the linear layer, normalization layer and activation function:
[0082] ;
[0083] Where, and are the weight matrix and bias vector of the alignment transformation, is the normalization layer, is the activation function;
[0084] S42. Based on the basic feature representation output by the feature fusion encoder, music structure planning is performed to form a structural framework of the music;
[0085] The structural framework of the music The expression formula is:
[0086] ;
[0087] Where, Generate functions for structures; It is the basic feature representation output by the feature fusion encoder; is the style constraint; for a set of music theory rules;
[0088] The structural planning adopts a multi-dimensional design, which can be expressed as:
[0089] ;
[0090] Where, For the musical form and structure; to frame harmonic progressions; Organize the structure for the rhythm;
[0091] S43. Using a three-layer cascade decoding architecture, the basic feature representation Z and the structural framework To decode, the expression formula is:
[0092] ;
[0093] ;
[0094] ;
[0095] Where, 、 and Respectively represent the structure layer decoding process, content layer decoding process and detail layer decoding process, 、 and They are the outputs of each layer respectively;
[0096] S44. Output features of detail layer Input pre-trained EnCodec decoder to generate high-fidelity audio waveform W’ , the process is expressed as:
[0097] ;
[0098] In the formula, Quantize is a vector quantization module that converts continuous features Mapped to discrete audio tokens;
[0099] Φ is the decoder parameter, which restores the audio waveform through deconvolution and upsampling; A pre-trained EnCodec decoder that converts quantized audio tags into high-fidelity waveforms W’ ;
[0100] The music generation process in step S4 adopts the following optimization and constraint mechanism:
[0101] (1) Musical continuity constraints, namely:
[0102] ;
[0103] Where, is the music coherence measure at time t, is the minimum coherence threshold;
[0104] (2) Emotional expression constraints, namely:
[0105] ;
[0106] Where, is the L2 norm; for t The expression changes at each moment. is the maximum allowable variation.
[0107] As a further improvement of the above solution, step S5 includes the following specific steps:
[0108] S51. Using the same ViT model as in step S3 to monitor the user's facial emotional features in real time, thereby calculating the expression change vector, emotional state vector and attention focus;
[0109] The calculation formula of the expression change vector is:
[0110] ;
[0111] Where, and Continuous frame features extracted by the encoder of the ViT model;
[0112] The calculation formula of the emotional state vector is:
[0113] ;
[0114] Where, for t The emotional state vector at the moment; for t The pleasure value of the moment, ranging from [-1, 1], where positive values indicate positive emotions and negative values indicate negative emotions; for t The awakening value at the moment, ranging from [-1, 1], where positive values indicate high energy state and negative values indicate low energy state;
[0115] The formula for calculating attention concentration is:
[0116] ;
[0117] Where, for t The attention score at the moment, ranging from [0,1]; For the attention component, the user's concentration on music is evaluated by analyzing eye movements, head posture and micro-expression changes; for t The set of key facial features at the moment;
[0118] S52. Perform feature analysis on the generated music in real time, and the expression formula is:
[0119] ;
[0120] Where, for t The music feature vector at the moment, for t The audio signal generated at each moment; A music analyzer is used to analyze the rhythm dynamics, harmonic complexity and timbre characteristics of the generated music;
[0121] S53. The intensity of the input emotional features and the output music parameters is adjusted through a bidirectional gating mechanism. The calculation formula is:
[0122] ;
[0123] ;
[0124] Where, and They are t The characteristic input control gate and parameter output modulation gate at each moment; is the input weight matrix, is the output weight matrix, is the input bias term; is the output bias term; The sigmoid activation function is used to normalize the gate value to the [0,1] interval;
[0125] S54. Use a PID controller to dynamically adjust the music parameters according to the emotional deviation. The calculation formula is:
[0126] ;
[0127] Where, is the output control signal of the PID controller; is the error between the target emotional state and the current state; are proportional term, integral term and differential term respectively; among them, when When the preset threshold is exceeded, the control intensity is increased to speed up the emotional adjustment, and vice versa, the control intensity is reduced to maintain a stable state;
[0128] S55. Comprehensively evaluate the effect of music therapy through multi-dimensional indicators to form a quantitative treatment effect function:
[0129] ;
[0130] Where, express t The therapeutic effect of music at all times; 、 and are the weight coefficients of expression change, emotional state and attention level, , adjust the weight according to different treatment stages and goals; among them, increase , increases during the emotional expression period , increased during the focused training period ;
[0131] S56. The music generation parameters are updated according to the treatment effect. The expression formula of the update process is:
[0132] ;
[0133] Where, and They are and The music generation parameters at each moment, is the learning rate, for The gradient of treatment effect at each moment.
[0134] As a further improvement to the above solution, the following optimization and constraint mechanism is adopted in the closed-loop control of step S5:
[0135] (1) Minimize treatment bias, that is:
[0136] ;
[0137] Where, For the target treatment effect;
[0138] (2) Musical continuity constraints, namely:
[0139] ;
[0140] Where, for ta measure of musical coherence at each moment; is the minimum coherence threshold;
[0141] (3) Emotional smoothness, namely:
[0142] ;
[0143] Where, is the change in emotional state; is the maximum allowable variation;
[0144] (4) Response delay, that is:
[0145] ;
[0146] Where, is the response time; is the maximum allowed response delay.
[0147] The present invention also discloses a computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the emotional music adjustment method based on dynamic perception and multimodal fusion as described above are implemented.
[0148] Compared with the prior art, the present invention has the following beneficial effects:
[0149] 1. The emotional music adjustment method proposed in this invention, based on dynamic perception and multimodal fusion, primarily addresses issues in existing music therapy methods, such as incomplete feature perception, inaccurate emotion mapping, and unstable treatment effects. Through innovative design of core technologies such as a dynamic perception field mechanism, an emotion-music bridge adapter, and adaptive feedback control, this invention achieves intelligent processing of the entire process, from user feature perception to emotional music generation. This invention adopts a modular design approach, progressively optimizing aspects such as feature processing, spatial mapping, and feedback control, ultimately constructing a complete end-to-end treatment framework.
[0150] 2. In terms of feature perception, the present invention uses the LLaMA encoder to process four types of features, namely user portraits, personal hobbies, historical experiences, and physiological data, and introduces a dynamic perception field mechanism to realize adaptive calculation of feature importance, significantly improving the integrity of feature expression and the accuracy of feature weight allocation.
[0151] 3. Regarding emotion-music mapping, this invention utilizes an innovative bridge adapter structure to achieve precise conversion from feature space to music parameter space through facial expression feature extraction and emotion state mapping. Compared to existing technologies, facial expression feature extraction achieves higher accuracy, emotion state mapping provides more comprehensive dimensional coverage, and the precision of music parameter mapping is significantly improved. Through strict parameter constraints and adaptive temperature control, the standardization and creative performance of the generated music are significantly improved.
[0152] 4. Regarding treatment effectiveness, this invention incorporates a PID controller and adaptive feedback regulation strategy to achieve precise control and continuous optimization of treatment parameters. This significantly improves control accuracy, parameter optimization efficiency, and treatment effect assessment. The optimized configuration of feature processing parameters and control parameters makes system operation more stable, effectively reduces volatility, significantly enhances treatment safety, and provides patients with a more reliable treatment experience.
[0153] 5. Through the organic integration of the aforementioned technological innovations, this invention achieves breakthroughs in feature perception, emotion mapping accuracy, and therapeutic efficacy, significantly enhancing overall performance. This invention not only provides a reliable music therapy technical solution, but its modular design also facilitates functional expansion and performance optimization, thus possessing significant application value and promotional significance.
[0154] 6. The computer terminal disclosed in the present invention can produce the same effect as the above method by applying the above method, and will not be described in detail. BRIEF DESCRIPTION OF THE DRAWINGS
[0155] Figure 1 This is a flowchart of the emotional music adjustment method based on dynamic perception and multimodal fusion in Example 1 of the present invention.
[0156] Figure 2 This is a system block diagram of the emotional music adjustment method based on dynamic perception and multimodal fusion in Example 1 of the present invention. DETAILED DESCRIPTION
[0157] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0158] Example 1
[0159] In response to the technical problems existing in the prior art, the present invention proposes an innovative solution: First, a multi-level feature extraction architecture based on dynamic perception fields is designed, which realizes adaptive analysis and dynamic adjustment of feature importance by integrating multi-dimensional features such as user portraits, historical experiences, personal preferences and real-time physiological data. Secondly, an emotion-music bridge based on a hierarchical memory network is constructed, which combines a multimodal cross-attention mechanism to achieve precise mapping between emotional features and music generation parameters while ensuring the artistry of music. Finally, a complete feedback control loop is established, which realizes real-time adjustment of treatment strategies and experience accumulation through multimodal signal processing and parameter dynamic optimization mechanism, thereby significantly improving the personalized treatment effect.
[0160] See also Figure 1 and Figure 2 This embodiment provides an emotional music adjustment method based on dynamic perception and multimodal fusion, comprising the following steps:
[0161] S1. Obtain multi-dimensional characteristic data of users, including user profile, personal preferences, historical experience and physiological data;
[0162] S2. Dynamically assigning weights to the multi-dimensional feature data through a dynamic perception field to generate fusion features;
[0163] S3 real-time collection of user facial image data, facial emotion features extracted by the ViT model, through the emotion - music bridge adapter facial emotion feature space mapped to the music feature space; the facial emotion features include expression coding, emotion dimension, emotion intensity and temporal characteristics;
[0164] S4. Inputting the fused features and the facial emotion features mapped to the music feature space into the music generation module, and generating personalized therapeutic music through feature fusion encoding, music structure planning, and hierarchical decoding architecture in the module;
[0165] S5. Monitor the user's facial emotional features and generate music features in real time, adjust the music generation parameters through a dynamic feedback mechanism, and form a closed-loop control.
[0166] In its overall architectural design, this invention innovatively employs a dual-path feature input mechanism, achieving deep synergy between dynamic perception field features and real-time facial expression features. The dynamic perception field module is responsible for comprehensively processing the user's multi-dimensional features, including user profiles (such as basic information, psychological characteristics, and medical background), personal interests (such as art and culture, sports and health, and cognitive exploration), historical experiences (such as growth and development, emotional interaction, and musical experience), and physiological data (such as vital signs, neuroendocrine function, and physical function). Through processing with a deep learning model, these features are converted into a unified representation consisting of a fused feature vector, feature weight states, importance scores, and an interaction matrix. After feature alignment, normalization, and temporal supplementation, the basic feature representation required for music generation is formed. Regarding facial expression feature processing, video analysis continuously captures the user's facial expression changes, extracting a complete expression feature set that includes expression coding, emotional dimensions (pleasure and arousal), emotional intensity, and temporal changes. Based on the user's multi-dimensional features and real-time facial expression characteristics, an innovative bridge adapter is designed to inject these facial expression features between the various blocks of the music generation module encoder. Through feature alignment and attention mechanisms, the adapter ensures that facial emotion features are fully utilized during the encoding process, thereby achieving progressive feature fusion and enhancement. This design not only ensures the input perception of perceptual field features, but also enables the dynamic injection of facial expression features, which can help improve emotional expression capabilities.
[0167] To achieve real-time optimization and dynamic adjustment of emotional music generation, the present invention constructs a feedback mechanism based on bidirectional gating. This mechanism establishes a complete feedback loop by continuously monitoring changes in facial emotional features (including expression amplitude, emotional state vector, and attention focus) and the characteristics of the generated music (including rhythmic dynamics, harmonic complexity, and timbre). The present invention innovatively designs input gating and output gating, dynamically adjusting the strength of the feature stream through a learnable weight matrix, ensuring real-time matching of the generated music with the user's emotional state. In specific implementations, input gating adjusts the input feature stream based on the degree of match between facial emotional and musical features, while output gating optimizes music generation parameters based on real-time monitoring of expression changes, emotional state, and attention level. When a significant deviation between the emotional state and the target is detected, a rapid response is achieved by adjusting the gate opening and parameter adjustment step size. When the deviation is within an acceptable range, the parameters are kept stable and fine-tuned to ensure the coherence of the generated music. The entire closed-loop control process is precisely regulated by a PID controller, ensuring both real-time responsiveness and smooth transitions in emotional expression.
[0168] Through the integration of dynamic perceptual field features, feature injection via the bridging adapter, and feedback regulation via the gating mechanism, we achieve a deep fusion of emotional features and music generation, providing powerful technical support for emotional music therapy. The specific implementation schemes for each module are detailed below.
[0169] 1. Dynamic perception field feature capture
[0170] The main function of the dynamic perception field is to realize multi-dimensional capture of user characteristics, dynamic importance perception and adaptive feature fusion. Its main innovation lies in processing unstructured text based on a large language model to achieve in-depth understanding of the user's all-round characteristics and dynamic weight allocation.
[0171] 1.1 Multi-dimensional feature vector capture
[0172] In order to achieve a comprehensive perception of user characteristics, the present invention uses four LLaMA encoders to process the text descriptions of the four dimensions of user portrait, personal hobbies, historical experience and physiological data, and convert them into standardized feature vectors. Figure 2 As shown in Figure 2, the LLaMA encoder receives four types of input information and outputs corresponding feature representations. The basic mapping process of each encoder can be expressed as:
[0173] ;
[0174] Where, is the encoded feature vector, X To ensure computational efficiency, this paper uses a lightweight version of LLaMA to convert the original text into a 768-dimensional feature vector.
[0175] The four types of features are: User portrait information is encoded as , personal hobby information is encoded as , historical experience information is encoded as , physiological data information is encoded as The specific contents of the four types of features are as follows:
[0176] (1) User profile information:
[0177] It mainly covers the following four dimensions: basic information: including age, gender, occupation, education level and marital status; psychological characteristics: including personality type, depression tendency score, anxiety level score, stress tolerance level and emotional stability score; treatment background: including medical history, treatment stage, treatment goals, expected course of treatment and treatment frequency; social background: including cultural background, language habits, living environment, family structure and social activity.
[0178] (2) Personal hobbies:
[0179] It mainly covers the following four dimensions: Art and culture hobbies: including music appreciation preferences, visual art interests, performing arts tendencies, literary creation enthusiasm and traditional cultural preferences; Sports and health hobbies: including sports habits, health and wellness methods, meditation and relaxation preferences, outdoor activity interests and physical training tendencies; Cognitive exploration hobbies: including knowledge learning direction, scientific and technological innovation interests, thinking challenge preferences, skill training enthusiasm and professional research tendencies; Life and entertainment hobbies: including social activities, leisure and entertainment methods, hand-made interests, collection appreciation preferences and life experience tendencies.
[0180] (3) Historical experience:
[0181] It mainly covers the following four dimensions: growth and development experience: including family education experience, school growth trajectory, important life turning points, stress coping experience and changes in self-cognition; emotional interaction experience: including important emotional relationships, interpersonal communication patterns, social support networks, emotional trauma experiences and emotional regulation strategies; music experience: including music learning experience, participation in music activities, music emotional memory, music therapy experience and music scene experience; life adaptation experience: including career development process, environmental adaptability, habit formation process, changes in interests and hobbies and evolution of lifestyle.
[0182] (4) Physiological data:
[0183] It mainly covers the following four dimensions:
[0184] Basic vital signs: including heart rate and blood pressure levels, respiratory rhythm changes, body temperature and metabolic indicators, height-to-weight ratio, and physical activity status;
[0185] Neuroendocrine: including brain wave activity, neurotransmitter levels, hormone secretion, stress response level and physiological rhythm performance;
[0186] Physical function: including immune system function, organ function, motor coordination, sensory system sensitivity, and musculoskeletal condition;
[0187] Health status: including disease diagnosis records, degree of chronic diseases, sleep quality level, rehabilitation treatment progress and complication risk assessment.
[0188] 1.2. Dynamic Perception Field Construction
[0189] After completing the feature coding of the four dimensions of user portrait, personal hobbies, historical experience and physiological data, the present invention realizes dynamic perception of the importance of user features through feature fusion and dynamic perception field technology. Figure 2 As shown in Figure 2, the four types of features after LLaMA encoding are processed and fused through the weighting module.
[0190] 1.2.1 Multi-dimensional feature fusion
[0191] The present invention first performs a preliminary fusion of multi-dimensional features, and then constructs a dynamic perception field to realize dynamic perception of feature importance.
[0192] In the initial feature fusion stage, the user portrait ( )、Personal hobbies( )、Historical Experience( ) and physiological data ( ) four types of coded feature vectors are spliced together to build a unified feature representation space and achieve preliminary feature fusion. Where E is the fused feature vector, , the dimension is n×4, where n is the dimension of the single-class feature vector (n=768).
[0193] Then, a dynamic perception field is constructed to achieve dynamic perception of user features. The Weighting Module corresponds to the core component of the dynamic perception field. First, the weights of the four types of features are initialized. The feature weights in each dimension are evenly distributed. The weight set W is defined as:
[0194] ;
[0195] in The weight coefficients corresponding to the four types of features are: , .
[0196] It's important to note that this initial fusion stage involves a low-level, structural feature fusion using simple vector concatenation. This leaves the original features intact, their weights unchanged, and the dimensionality expanded from N to N*4. This can be thought of as a "physical concatenation" of features, similar to simply stacking different information sources together. This provides the foundational data structure for subsequent dynamic weight assignment and deep fusion.
[0197] The adaptive fusion of dynamic perception fields is a high-level, semantic feature fusion. It uses adaptive weighted fusion based on dynamic weights to dynamically adjust feature weights according to the results of preliminary fusion: the weight distribution is continuously optimized through the gradient descent method.
[0198] 1.2.2 Perception Field Strength Function
[0199] Dynamic perception field is achieved by designing the perception field intensity function To dynamically perceive the importance of each dimension feature. This function includes the direct impact item and interaction terms There are two core parts, which are used to capture the independent effects and interactions of features:
[0200] ;
[0201] Among them, the direct impact Reflects the direct contribution of each feature to the perception field, and the expression formula is:
[0202] ;
[0203] This method comprehensively considers three key elements: feature weights, feature vectors, and time decay. It uses the weighted combination of feature weights and feature vectors to reflect the importance of each feature. It also introduces a time decay factor to dynamically adjust the strength of feature influence over time. The sum of all weighted features yields the overall direct impact strength.
[0204] Interaction terms It focuses on describing the interaction between features, and the expression formula is:
[0205] ;
[0206] Where, is the time decay factor; ; For feature correlation, cosine similarity is used to quantify the encoded feature vector and The degree of correlation between them ranges from [0,1]. The larger the value, the stronger the feature correlation.
[0207] is the dynamic coupling coefficient, which uses gradient information to characterize the direction of the feature's influence on the perception field, uses gradient product to reflect the degree of coordination of feature changes, and uses the sigmoid function to normalize the value range of the dynamic coupling coefficient to [0,1].
[0208] The interaction term is modeled by two factors: feature correlation and dynamic coupling coefficient. Cosine similarity is used to quantify the degree of correlation between features. The value range is limited to the interval [0,1]. The larger the value, the stronger the feature correlation. In the calculation process, special consideration is given to the directional consistency of the feature vector to eliminate the influence of dimensional differences. Dynamic coupling coefficient Gradient information is used to characterize the direction of a feature's influence on the perceptual field. Gradient products are used to reflect the degree of coordination among feature changes. The coefficients are normalized to the range [0, 1] using a sigmoid function. Finally, the interaction effects of all feature pairs are accumulated to capture the complex interactions and combined effects between features.
[0209] 1.2.3 Dynamic Weight Optimization
[0210] To achieve dynamic optimization of feature weights, this paper adopts an improved momentum-based gradient descent method for weight update. This method improves the efficiency and stability of weight optimization by introducing momentum terms and second-order gradient information:
[0211] ;
[0212] Where, and are the feature weights before and after updating respectively; is the learning rate; is the momentum factor; is the second-order gradient; the learning rate Controls the step size of weight update, the default value is 0.01, momentum factor Use historical gradient information to accelerate convergence, the default value is 0.9, by introducing the second-order gradient , to accurately grasp the optimization direction and effectively avoid local optimal solutions. is the first-order gradient.
[0213] Through the above design, the dynamic perception field can capture the user's multi-dimensional features and perform adaptive fusion, providing optimized feature representation for the subsequent emotion-music bridge, and is the basic module of the entire music therapy method. Figure 2 The weighted module in
[15] is the core innovation to achieve this dynamic feature weight distribution.
[0214] The dynamic perception field mechanism is a core innovation and main protection content of the present invention. This mechanism processes four types of features, namely user portraits, personal hobbies, historical experiences and physiological data, through the LLaMA model to realize the vectorized expression of multi-dimensional features. On this basis, through feature fusion and dynamic perception field construction methods, including feature splicing, weight initialization and field strength calculation technologies, unified expression of features is achieved. At the same time, a dynamic evaluation mechanism of feature importance is introduced, which comprehensively considers direct and indirect influences, and adopts the gradient descent method with momentum for adaptive weight update, thus realizing accurate calculation and dynamic adjustment of feature importance.
[0215] 2. Emotion-Music Bridge Adapter
[0216] The emotion-music bridge adapter is the second innovation of the present invention, and its main function is to establish the mapping relationship between facial emotion features and music generation space. Figure 2 As shown in the figure, this module connects the facial emotion feature extraction and music generation modules to achieve deep integration of facial emotion feature extraction and music generation.
[0217] 2.1 Facial emotion feature extraction
[0218] 2.1.1 ViT model architecture and parameter configuration
[0219] Facial emotion is the most direct and richest carrier of human emotional expression. This paper uses the Vision Transformer (ViT) model to process the input video frame sequence, such as Figure 2 As shown in Figure 2, the ViT encoder receives a sequence of facial images in different emotional states as input and outputs the encoded emotional feature vector. The feature extraction process can be formally expressed as:
[0220] ;
[0221] Where, is the facial emotion feature vector; Represents the encoding process of the ViT encoder; is a facial image sequence; is the first facial emotion feature vector feature components, each of which represents a specific aspect in the emotional feature space, together forming a complete representation of the four dimensions of expression coding, emotional dimension, emotional intensity and temporal characteristics. , m is the total number of characteristic components.
[0222] The model is configured with the following parameters:
[0223] Basic architecture: ViT-B / 16 (12-layer Transformer, 16×16 image block); Input resolution: 224×224 pixels, 32-bit color image; Time sequence window: 15 consecutive frames (about 0.5 seconds, sampling rate 30fps); Output dimension: 768-dimensional emotion feature vector; Number of attention heads: 12, each head has 64 dimensions; Hidden layer dimension: 3072.
[0224] 2.1.2 Multi-dimensional analysis of facial features
[0225] To fully capture the user's emotional state, this paper analyzes facial features from the following four dimensions:
[0226] Expression coding: Capturing micro-expression features such as facial muscle changes and eye changes is the basis of expression recognition;
[0227] Emotional dimension: constructing a two-dimensional emotional space through pleasure and arousal to achieve precise emotional positioning;
[0228] Emotional intensity: quantifies the degree and depth of emotional expression and determines the weight of the influence of emotion on music generation;
[0229] Temporal features: Characterize the temporal dynamic characteristics of emotional changes to ensure that the generated music can smoothly transition with emotional changes.
[0230] Each feature dimension is represented by an independent sub-vector and merged into a unified feature vector through linear projection, thereby retaining multiple aspects of information.
[0231] 2.2 Bridge Adapter Design
[0232] The emotion-music bridge adapter is a key component that connects facial emotion features with the music generation space. The core task of the emotion-music bridge adapter is to solve the heterogeneity problem between the emotion feature space and the music feature space. Figure 2 As shown in FIG, the adapter adopts a dual-channel design, including two processing paths: upstream adaptation and downstream adaptation.
[0233] (1) Uplink adaptation channel ( )
[0234] The facial emotion features are converted to the music feature space through projection layer, ReLU activation and layer normalization:
[0235] ;
[0236] Where, is the characteristic representation of the uplink adaptation channel output; Normalization of the representation layer; ReLU activation processing; is the uplink adaptation channel projection weight matrix, , is the music feature space dimension, It is the dimension of facial emotion characteristics; Adapt channel bias parameters for uplink; is a set of real numbers. In this embodiment , .
[0237] (2) Downlink Adaptation Channel ( )
[0238] The music feature dimensions are compressed through layer normalization, projection layer, and gated unit features to extract core features related to emotions.
[0239] ;
[0240] Where, Feature representation for downlink adaptation channel output; It is processed by gated units to selectively retain key features; is the representation of the music feature space; is the downlink adaptation channel projection weight matrix, ; Adapt channel offset parameters for downlink.
[0241] The two channels work together and are integrated through residual connections:
[0242] ;
[0243] Where, It is a bridging feature, a bidirectional fusion bridging feature. It should be noted that: It is the result of converting facial emotion features into music feature space. It is to extract the core features related to emotions from the music feature space. The purpose is to achieve bidirectional mapping and information fusion between the emotional feature space and the music feature space; It is a residual connection process used to preserve feature information and optimize gradient propagation.
[0244] 2.3 Optimization and Constraint Mechanism
[0245] The present invention considers two optimization goals simultaneously: computational efficiency and generation quality:
[0246] (1) Computational efficiency goal: Minimize the delay introduced by the emotion-music bridge adapter
[0247] ;
[0248] In the formula, the subscript A set of trainable parameters for the emotion-music bridge adapter, including upstream channel parameters , downlink channel parameters and internal parameters of the gate control unit; represents the expected value operator, which is used to calculate the average performance of a random variable; The delay introduced for the Emotion-Music Bridge Adapter.
[0249] (2) Emotional mapping accuracy goal: maximize the consistency between emotional expression and musical characteristics
[0250] ;
[0251] Where, Represents facial emotion feature vector Music parameter characteristics Semantic consistency measurement of . It should be noted that is the output of the encoder of the music generation module in the overall architecture. The training goal of the model is to make the recognized facial emotion features as similar as possible to the corresponding music parameter features. It can be understood as in the following text.
[0252] (3) Music quality goal: maximize music quality under delay constraints
[0253] ;
[0254] Where, To generate music quality metrics; The upper threshold of the delay time. The full name in mathematics is "subject to", which means "subject to".
[0255] Music quality metrics Defined as a multi-factor composite score:
[0256] ;
[0257] Where, is the weight coefficient, : Musical structure integrity score (0-1), measuring the logic of musical form; : Melodic fluency score (0-1), assessing the naturalness and memorability of the melody; : Harmonic coordination score (0-1), evaluating the balance and richness of the harmonic progression; : Emotional expressiveness score (0-1), measuring the match between the music and the target emotion.
[0258] The emotion-music bridge adapter adopts the following multi-objective balancing strategy:
[0259] (1) Hierarchical priority strategy, dynamically adjust the optimization target weight according to user scenario requirements:
[0260] ;
[0261] Where, Optimize the objective function for the whole; Indicates the A single optimization goal; A set of parameters for the emotion-music bridge adapter; is the context-related weight coefficient, ; The current interaction scene type;
[0262] When in a real-time treatment scenario, > > , giving priority to ensuring computational efficiency;
[0263] When quality is the priority, > > , giving priority to music quality;
[0264] When in the emotional response field, > > ,give priority to ensuring the accuracy of emotion mapping;
[0265] (2) Constraint relaxation strategy allows some constraints to be softened in extreme scenarios. The expression formula is:
[0266] ;
[0267] Where, is a soft constraint penalty term; is the constraint relaxation coefficient; For the Constraints, , which correspond to the three constraints mentioned above: computational efficiency, emotion mapping accuracy, and music quality;
[0268] When the system load exceeds the preset load threshold hour, Decrease, relax non-critical constraints; it should be noted that the system load index is calculated by weighting CPU usage, memory usage and response delay time. When the system load exceeds When performing the optimization, the parameter weights of non-critical constraints such as music quality are attenuated to prioritize system responsiveness.
[0269] When the emotional change exceeds the threshold hour, Incrementally, strengthen the emotional mapping constraints, that is, ensure that the music responds to emotional changes in a timely manner.
[0270] Through the emotion-music bridging adapter, the present invention realizes the mapping from user facial expressions to music generation, so that the generated music can accurately reflect the user's emotional state and provide a solid technical foundation for personalized music therapy.
[0271] The emotion-music bridge adapter is another key feature of this invention. This module first extracts and encodes facial emotion features using the ViT model, acquiring core information such as expression encoding, emotion dimension, emotion intensity, and temporal features. On this basis, a multi-layer feature fusion mechanism is designed. Through feature alignment, attention calculation, and fusion, it enables the layer-by-layer injection of emotion features. Deep integration with the music generation module ensures the precise mapping of emotion features to music parameters.
[0272] 3. Music Generation Module
[0273] The music generation module is the core execution unit of the present invention, responsible for converting the user characteristics and emotional state acquired in the early stage into personalized therapeutic music. Figure 2 As shown in the music generation module, this module adopts the "encoder-decoder" architecture to achieve accurate mapping from feature space to music space.
[0274] 3.1 Feature Fusion Encoder
[0275] The feature fusion encoder is the first stage of music generation, responsible for integrating the two feature inputs and constructing the basic representation of music generation.
[0276] (1) Feature input
[0277] The encoder receives two feature inputs:
[0278] User historical feature information from the dynamic perception field, including the fusion representation of user portraits, personal preferences, historical experiences, and physiological data;
[0279] Real-time facial emotion features from the emotion-music bridge adapter are injected layer by layer during encoding through the adapter.
[0280] (2) Encoder structure
[0281] like Figure 2 As shown on the left side of the music generation module, the encoder is composed of N computational blocks with the same structure, each of which contains:
[0282] Add&Norm layer (residual connection and layer normalization operation): implements feature normalization and residual connection;
[0283] Feed-Forward layer (feedforward neural network layer): improves feature representation capabilities through nonlinear transformation;
[0284] Multi-Head Attention layer: captures long-distance dependencies between features.
[0285] (3) Feature fusion
[0286] By injecting facial emotion features between the layers of the feature fusion encoder, we achieve progressive fusion of emotion features. This process is mainly accomplished through the attention mechanism, which specifically includes the calculation of Q, K, and V and the fusion of attention weights.
[0287] The attention calculation mechanism obtains the query vector (Q), key vector (K) and value vector (V) through three independent linear transformations. The query vector is obtained from the hidden state of the current encoding layer. The key vector and value vector are generated from the aligned sentiment features derived.
[0288] ;
[0289] ;
[0290] ;
[0291] in, is the hidden state of the Lth layer, are the weight matrices for query, key, and value respectively, for The facial emotion feature representation obtained after processing through the linear layer, normalization layer and activation function:
[0292] ;
[0293] Where, and are the weight matrix and bias vector of the alignment transformation, is the normalization layer, is the activation function.
[0294] The standard scaled dot product attention mechanism is used to achieve selective fusion of features by calculating the attention weight matrix.
[0295] ;
[0296] Where, For the l +1 hidden state of the encoding layer; Normalization of the representation layer; Represents a feedforward neural network for nonlinear mapping and dimension transformation of features; is the activation function; d is the scaling factor for the dot-product attention mechanism.
[0297] 3.2 Music Structure Planning
[0298] Based on the feature representation Z output by the encoder, the present invention first performs music structure planning and designs the overall framework of the music. This process is similar to the conception stage before the creation of a composer, which determines the overall skeleton of the music:
[0299] ;
[0300] in is the structure generating function, is a style constraint, is a set of music theory rules.
[0301] The structural planning adopts a multi-dimensional design to form a complete music organization structure:
[0302] ;
[0303] in:
[0304] : Musical form and structure, including the design of sections such as theme presentation, development and reappearance;
[0305] : Harmonic progression framework, including tonal center, chord progression and mode changes;
[0306] : Rhythmic organizational structure, including beat type, rhythmic density and rhythmic pattern.
[0307] 3.3 Layered Decoding Generation
[0308] Music decoding uses a three-layer cascade decoding architecture, gradually refining the music content from macro to micro:
[0309] ;
[0310] ;
[0311] ;
[0312] (1) Structure layer decoding ):
[0313] Input: Encoded feature Z and structural frame ;
[0314] Function: Determine the basic structure of the music, including paragraph length, theme position and tonality planning;
[0315] Output: A representation of the musical structure, similar to a "sketch" of the music.
[0316] (2) Content layer decoding ( )
[0317] Input: Structure layer output ;
[0318] Function: Generate specific musical materials, including melodic lines, harmonic progressions, and rhythmic patterns;
[0319] Emotional mapping: different emotional states are mapped to specific musical parameters:
[0320] Pleasantness → pitch distribution, harmony type (e.g., high pleasantness corresponds to ascending melodies and bright chords);
[0321] Arousal → rhythm density, note duration (e.g. high arousal corresponds to tight rhythm and short notes).
[0322] (3) Detail layer decoding ):
[0323] Input: content layer output ;
[0324] Function: Improve the expressive details of music, including dynamic changes, playing techniques and timbre adjustment.
[0325] To balance predictability and creativity, the present invention introduces a temperature control mechanism. The temperature parameter T dynamically adjusts the randomness of the decoder. When stabilizing emotions is needed, the temperature is lowered to enhance predictability. When emotional exploration is needed, the temperature is increased to increase variability. The temperature parameter T at time t is expressed as :
[0326] ;
[0327] in, is the basic temperature parameter, For creative regulation.
[0328] 3.4 Music Generation
[0329] Output features of the detail layer Input pre-trained EnCodec decoder to generate high-fidelity audio waveform W’ , the process is expressed as:
[0330] ;
[0331] In the formula, Quantize is a vector quantization module that converts continuous features Mapped to discrete audio tokens; Φ is the decoder parameter, and the audio waveform is restored through deconvolution and upsampling; A pre-trained EnCodec decoder that converts quantized audio tags into high-fidelity waveforms W’ .
[0332] 3.5 Optimization and Constraint Mechanism
[0333] To ensure the therapeutic effect and artistic quality of the generated music, the present invention sets multiple constraints:
[0334] (1) Musical coherence constraints ;
[0335] in, is the music coherence measure at time t, is the minimum coherence threshold, set to 0.7.
[0336] (2) Emotional expression constraints
[0337] This constraint ensures the gradual expression of emotions and avoids sudden emotional changes during the treatment process, which is in line with the theory of "psychological safety zone".
[0338] ;
[0339] Where, is the L2 norm; for t The expression change vector at each moment; is the maximum allowable variation, set to 0.3.
[0340] 4. Dynamic feedback mechanism
[0341] The dynamic feedback mechanism is the closed-loop control core of the present invention, which enables it to respond to user emotional changes in real time and adjust the music generation strategy. Figure 2 As shown in the dynamic feedback mechanism module, this mechanism connects the emotion feature extraction and music generation modules to form a complete "perception-generation-feedback" cycle, achieving precise emotion regulation and personalized music therapy experience.
[0342] 4.1. Emotional Feature Monitoring
[0343] This paper adopts a unified facial emotion feature extraction architecture to ensure consistency in feature representation. Although the feedback mechanism and the bridging adapter use the same ViT encoder to extract basic features, there are significant differences in feature processing and application between the two. The bridging adapter focuses on feature mapping transformation, converting facial emotion features into a representation space suitable for music generation, and focuses on the transformation and adaptation of the feature space. Unlike the emotion-music bridging adapter, the dynamic feedback mechanism focuses more on the temporal changes and comparative analysis of facial expressions rather than single-frame features.
[0344] (1) Expression change vector: measures the magnitude of expression changes between consecutive frames and reflects the degree of emotional fluctuation ;
[0345] Where, and It is the continuous frame features extracted by the encoder of the ViT model, i.e., time t and time t-1.
[0346] The system uses the same continuous frame features extracted by the ViT encoder in Section 2.1 of this embodiment and , calculates the amount of expression change. This indicator is designed in the hope of capturing subtle changes in facial expressions, such as raised eyebrows, upturned corners of the mouth, or widened eyes, and providing instant feedback on emotional changes.
[0347] (2) Emotional state vector: Construct a two-dimensional emotional space to accurately locate the emotional state at the current time t :
[0348] ;
[0349] The present invention extracts pleasure and arousal values from the 768-dimensional feature vector output by the ViT encoder through a specialized mapping layer, constructing a real-time updated two-dimensional emotional space representation. Valence(t) is the pleasure value at time t, ranging from -1 to 1, with positive values indicating positive emotions and negative values indicating negative emotions; arousal(t) is the arousal value at time t, ranging from -1 to 1, with positive values indicating high energy and negative values indicating low energy.
[0350] (3) Attention: Evaluate the user's concentration on music:
[0351] ;
[0352] The present invention reuses the facial features extracted by ViT and uses the additional Attention component to specifically analyze eye movements, head posture and micro-expression changes to assess the user's concentration on music. is the attention score at time t, ranging from [0, 1]. The attention level is assessed by analyzing eye movement patterns, head posture stability, and facial micro-expression changes. A high attention score indicates that the user is actively engaged in the current music, while a low score may mean distraction or waning interest. for t The collection of key facial features at that moment.
[0353] 4.2 Music Feature Analysis
[0354] Simultaneously with emotion monitoring, the characteristic analysis of the generated music is carried out. This process enables understanding the musical characteristics of its own output and evaluating its effects, ensuring that the musical output matches the emotional needs:
[0355] ;
[0356] in, is the music feature vector at time t, is the audio signal generated at time t. represents the processing of the feature analyzer.
[0357] The music analyzer extracts and analyzes the following key music parameters in real time:
[0358] (1) Rhythmic dynamic characteristics
[0359] Rhythmic density: the number of notes per minute, reflecting the activity of the music;
[0360] Accent distribution: the location and intensity of the strong beats, shaping the rhythm of the music;
[0361] Speed change: The stability or gradual change of speed affects emotional stability.
[0362] (2) Harmonic complexity index
[0363] Harmonic texture: chord density and number of voices;
[0364] Tonal stability: the degree of certainty or ambiguity of the tonal center;
[0365] Dissonance: The frequency and intensity of the use of dissonant intervals.
[0366] (3) Characteristics of timbre expression
[0367] Bright timbre: the proportion of high-frequency energy affects the emotional color;
[0368] Pitch differentiation: pitch range and main activity area;
[0369] Dynamics Contour: Volume change curve that shapes musical expression.
[0370] This multi-dimensional music analysis allows for a precise understanding of the characteristics of the generated music, which can then be compared with the user's emotional response to assess the therapeutic effects of the music. When specific musical features are found to trigger positive emotional responses, these features are strengthened; when certain musical elements lead to negative reactions, these elements are weakened or modified.
[0371] 4.3 Adaptive Control Mechanism
[0372] Adaptive control is the core link of dynamic feedback. By designing a bidirectional gating mechanism, precise adjustment of emotional characteristics and music parameters can be achieved.
[0373] 4.3.1 Bidirectional Gating Mechanism
[0374] The present invention designs two control units, input gate and output gate, such as Figure 2 The gating module is shown in Figure 1.
[0375] (1) Input gate: Controls the influence of emotional features on music generation ;
[0376] (2) Output gate: controls the adjustment intensity of music parameters ;
[0377] Where, and They are t The feature input control gate and parameter output modulation gate at each moment, the former controls the flow of feature information into the control system, and the latter controls the degree of influence of the control signal on the music parameters; is the input weight matrix, is the output weight matrix, is the input bias term; is the output bias term; The sigmoid activation function is used to normalize the gate value to the [0,1] interval.
[0378] The practical application of the gate value is reflected in the regulation of feature flow:
[0379] ;
[0380] ;
[0381] in, is the original extracted emotional feature vector, is the emotional feature vector after Gate adjustment, is the generated original music feature vector, is the music feature vector after Gate adjustment.
[0382] ⊙ represents element-wise multiplication, which enables dimension-wise adjustment of features. This design can dynamically control the flow of information. For example:
[0383] When users show a positive reaction: As the value increases, the influence of emotional features increases; Value adjustment to optimize current music parameters;
[0384] When a user displays emotional instability: The value decreases, reducing the impact of emotional fluctuations; As the value increases, the stability of the music is enhanced;
[0385] When user attention is waning: Adjust the value and introduce novel musical elements to attract attention again.
[0386] 4.3.2 PID Controller
[0387] To achieve more precise emotional state regulation, the present invention introduces a PID (proportional-integral-derivative) controller, which is a classic feedback control mechanism.
[0388] ;
[0389] Where, is the output control signal of the PID controller; is the error between the target emotional state and the current state; are proportional term, integral term and differential term respectively; among them, when When the preset threshold is exceeded, the control intensity is increased to speed up emotional adjustment, and vice versa, the control intensity is reduced to maintain a stable state.
[0390] The corresponding relationship between the functions of each part of the PID controller and the treatment situation:
[0391] (1) Proportional term ( ): Provides immediate adjustments based on the gap between the current emotional state and the target state;
[0392] For example: when anxiety is detected, the speed and complexity of the music are immediately reduced.
[0393] (2) Integral term ( ): Resolve long-term emotional deviations and deal with persistent emotional states;
[0394] For example: For prolonged periods of low mood, gradually introduce bright tones and ascending melodies.
[0395] (3) Differential term ( ): Predict emotional trends and make adjustments in advance to prevent emotional overshoot;
[0396] For example: when you detect a rapid shift in emotion, you can smooth out changes in musical characteristics to avoid overstimulation.
[0397] The present invention dynamically adjusts the PID parameters according to the size of the emotional deviation:
[0398] When |e(t)|>threshold: increase control intensity and speed up emotional adjustment;
[0399] When |e(t)|≤threshold: reduce the control intensity to maintain a stable state.
[0400] This design simulates the working method of professional music therapists: closely observing the patient's response, adjusting the intervention strategy in a timely manner, and balancing short-term response and long-term effect.
[0401] 4.4 Treatment Effect Evaluation
[0402] The present invention comprehensively evaluates the effect of music therapy through multi-dimensional indicators to form a quantitative treatment effect function. :
[0403] ;
[0404] Where, express t The therapeutic effect of music at all times; 、 and are the weight coefficients of expression change, emotional state and attention level, , adjust the weight according to different treatment stages and goals; among them, increase , increases during the emotional expression period , increased during the focused training period .
[0405] Treatment effect evaluation drives parameter update process:
[0406] ;
[0407] in, Generate parameters for music, is the learning rate, The present invention uses this parameter updating mechanism to continuously optimize the music generation strategy, so that the treatment effect gradually approaches the ideal goal.
[0408] To quantitatively evaluate treatment progress, the present invention designs the following specific indicators:
[0409] (1) Mood Improvement Index (EMI) ;
[0410] Where V is the pleasure value, a positive value indicates an improvement in mood, and a negative value indicates a deterioration in mood. is the final happiness value, is the initial happiness value.
[0411] (2) Emotional Stability (ESS) ;
[0412] The value range is [0,1], and the higher the value, the more stable the emotion. To find the standard deviation function, To find the distance function.
[0413] (3) Treatment Engagement Index (TEI) ;
[0414] Take into account the average attention level and the degree of attention fluctuation. is the averaging function.
[0415] 4.5 Optimization Constraints
[0416] To ensure the effectiveness and safety of treatment, the present invention sets multiple optimization constraints:
[0417] (1) Minimize treatment bias: Ensure that the actual treatment effect is close to the target level ;
[0418] in, Target treatment effect.
[0419] (2) Musical coherence constraints: ensuring the fluency of music generation ;
[0420] in, is the measure of music coherence at time t, is the minimum coherence threshold.
[0421] (3) Emotional smoothness: Avoiding drastic fluctuations in emotional states and ensuring gradual adjustment ;
[0422] in, is the change in emotional state, is the maximum allowable variation.
[0423] (4) Response delay: ensuring timely response ;
[0424] in, is the response time, is the maximum allowed response delay.
[0425] Through this dynamic feedback mechanism, the present invention enables precise monitoring, timely adjustment, and smooth transition of emotional states, providing users with a personalized music therapy experience. The various modules work together to ensure both therapeutic efficacy and the naturalness and coherence of the generated music.
[0426] The dynamic feedback control mechanism establishes a complete feedback control system, including real-time monitoring of emotional and musical characteristics, capable of accurately capturing key indicators such as facial expressions, emotional states, and attention levels. By designing an adaptive control mechanism based on a bidirectional gate and combining it with the precise adjustment capabilities of a PID controller, real-time optimization of treatment parameters is achieved. Simultaneously, under multiple constraints, including musical coherence, emotional smoothness, and response latency, the feedback control is ensured to be stable and reliable.
[0427] Example 2
[0428] This embodiment provides a computer terminal, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor.
[0429] The computer terminal can be a smart phone, tablet computer, laptop computer, etc. that can execute programs. In some embodiments, the processor can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data. When the processor executes the program, the steps of the emotional music adjustment method based on dynamic perception and multimodal fusion in Example 1 are implemented.
[0430] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A computer terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the following steps of the emotional music adjustment method based on dynamic perception and multimodal fusion are implemented: S1. Obtain multi-dimensional characteristic data of users, including user profile, personal preferences, historical experience and physiological data; S2. Dynamically assigning weights to the multi-dimensional feature data through a dynamic perception field to generate fusion features; S3 real-time collection of user facial image data, facial emotion features extracted by the ViT model, through the emotion - music bridge adapter facial emotion feature space mapped to the music feature space; the facial emotion features include expression coding, emotion dimension, emotion intensity and temporal characteristics; S4. Inputting the fused features and the facial emotion features mapped to the music feature space into the music generation module, and generating personalized therapeutic music through feature fusion encoding, music structure planning, and hierarchical decoding architecture in the module; S5. Monitor the user's facial emotional features and generate music features in real time, adjust the music generation parameters through a dynamic feedback mechanism, and form a closed-loop control.
2. The computer terminal according to claim 1, wherein: In step S1, four LlaMA encoders are used to process the text descriptions of the four dimensions of user portrait, personal hobbies, historical experience and physiological data respectively, and convert the text descriptions into standardized encoded feature vectors, thereby forming the multi-dimensional feature data.
3. The computer terminal according to claim 2, wherein: Step S2 includes the following specific steps: S21. Concatenate the four types of coded feature vectors: user profile, personal hobbies, historical experience, and physiological data, to construct a unified feature representation space and achieve preliminary feature fusion. The expression formula is: And=[And P ;AND H ;AND T ;AND G ] Where E is the fused feature vector with a dimension of n×4, where n is the dimension of the single-class feature vector; E P 、E H 、E T and E G These are the encoding feature vectors of user portrait, personal hobbies, historical experiences, and physiological data respectively; S22. Construct a dynamic perception field to achieve dynamic perception of feature importance; Among them, the weighting module in the dynamic perception field is used to initialize the weights of the four types of features and evenly distribute the feature weights in each dimension. The weight set W is defined as: W={w P ,In H ,In T ,In G } Where w P 、w H 、w T 、w G The weight coefficients corresponding to the four types of features, ∑w i =1,0≤w i ≤1, i∈{P,H,T,G}; The expression formula of the perception field intensity function of the dynamic perception field is as follows: F(E,t)=F direct (E,t)+F interact (E,t) Where F(E,t) is the perceptual field strength function; t is time; F direct (E,t) and F interact (E, t) are direct influence terms and interactive influence terms respectively, and the expression formula is: F direct (E,t)=∑[w i ·E i ·λ(t)] F interact (E,t)=∑∑[α ij ·R(E i ,E j )] Where λ(t) is the time decay factor; j∈{P,H,T,G}; R(E i ,E j ) is the feature correlation, and cosine similarity is used to quantify the encoding feature vector E i and E j The degree of correlation between them is in the range of [0,1]. The larger the value, the stronger the feature correlation. ij The dynamic coupling coefficient is the gradient information used to characterize the direction of the feature's influence on the perception field, the gradient product is used to reflect the degree of coordination of feature changes, and the sigmoid function is used to normalize the value range of the dynamic coupling coefficient to [0, 1]. S23. Update the feature weights using an improved momentum gradient descent method, and adaptively fuse the multi-dimensional features using the updated feature weights to generate fused features. The weight update formula is: Where w(t) and w(t+1) are the feature weights before and after the update, respectively; η is the learning rate; θ is the momentum factor; is the second-order gradient, is the first-order gradient.
4. The computer terminal according to claim 1, wherein: In step S3, the encoder of the ViT model receives the facial image sequence as input and outputs the encoded facial emotion feature vector. The expression formula of the feature extraction process is: E face =ViT encoder (V frames )=[e1,e2,…,e m ] Where, E face is the facial emotion feature vector; Vi T encoder (·) indicates the encoding process of the ViT encoder; V frames is a facial image sequence; e i’ is the i'th feature component in the facial emotion feature vector. Each component represents a specific aspect in the emotion feature space, and together they constitute a complete representation of the four dimensions of expression coding, emotion dimension, emotion intensity, and temporal characteristics. i'∈[1,m], where m is the total number of feature components.
5. The computer terminal according to claim 4, wherein: In step S3, the emotion-music bridge adapter adopts a dual-channel design, including two processing channels: uplink adaptation and downlink adaptation; Among them, the uplink adaptation channel is used to convert the facial emotion features into the music feature space through the projection layer, ReLU activation, and layer normalization. The expression formula is: Where H up is the feature representation of the uplink adaptation channel output; LayerNorm(·) represents the layer normalization process; ReLU(·) represents the ReLU activation process; W up is the uplink adaptation channel projection weight matrix, d music is the dimension of music feature space, d face is the facial emotion feature dimension; b up Adapt channel bias parameters for uplink; is the set of real numbers; The downlink adaptation channel is used to compress the music feature dimensions through layer normalization, projection layer, and gating unit, and extract the core features related to emotion. The expression formula is: H down =GetUnit(W down ·LayerNorm(H music )+b down ) Where H down is the feature representation of the downlink adaptation channel output; GateUnit(·) is the gate unit processing, which is used to selectively retain key features; H music is the representation of the music feature space; W down is the downlink adaptation channel projection weight matrix, b down Adapt channel bias parameters for downlink; The uplink adaptation channel and the downlink adaptation channel work together and are integrated through residual connections. The expression formula is: H bridge =H up +SkipConnection(H down ) Where H bridge is the bridging feature; SkipConnection(·) is the residual connection processing, which is used to retain feature information and optimize gradient propagation.
6. The computer terminal according to claim 5, wherein: In step S3, the emotion-music bridge adapter adopts the following optimization and constraint mechanisms: (1) Computational efficiency goal: Minimize the delay introduced by the emotion-music bridge adapter, that is: Where, subscript Φ bridge is the trainable parameter set of the emotion-music bridge adapter, including the upstream channel parameter W up 、b up , downlink channel parameter W down 、b down and internal parameters of the gate control unit; E[·] represents the expected value operator, which is used to calculate the average performance of random variables; Δt adapter Delay time introduced for the Emotion-Music Bridge Adapter; (2) Emotional mapping accuracy goal: maximize the consistency between emotional expression and musical characteristics, that is: Where, Sim(E face ,M param ) represents the facial emotion feature vector E face and music parameter characteristics M param Semantic consistency measurement of (3) Music quality objective: maximize the music quality under the delay constraint, that is: Where Q music is the generated music quality metric; ε is the upper threshold of the delay time; Among them, the music quality metric Q music It is defined as a multi-factor comprehensive score, and the calculation formula is: Q music =w 1 ·Q structure +w 2 ·Q melody +w 3 ·Q harmony +w 4 ·Q emotion Where w 1 ~w 4 is the weight coefficient; Q structure Scoring the integrity of musical structure, used to measure the logic of musical form; Q melody Scoring melodic fluency to assess melodic naturalness and memorability; Q harmony Scoring the harmony, used to evaluate the balance and richness of the harmony; Q emotion Scoring emotional expressiveness measures how well the music matches the target emotion; The emotion-music bridge adapter adopts the following multi-objective balancing strategy: (1) Hierarchical priority strategy, dynamically adjust the optimization target weight according to user scenario requirements: Where, J total Optimize the objective function for the whole; Indicates the i * A separate optimization objective; Φ is the parameter set of the emotion-music bridge adapter; α i* (c) is the context-related weight coefficient, c is the current interaction scene type; When in a real-time treatment scenario, α1>α2>α3, giving priority to ensuring computational efficiency; When the quality is prioritized, α3>α2>α1, giving priority to music quality. When in the emotional response field, α2>α3>α1, giving priority to ensuring the accuracy of emotional mapping; (2) Constraint relaxation strategy allows some constraints to be softened in extreme scenarios. The expression formula is: Where C soft is a soft constraint penalty term; is the constraint relaxation coefficient; For the jth * Constraints, j * =1, 2, 3, corresponding to the three constraints of computational efficiency, emotion mapping accuracy, and music quality, respectively; When the system load exceeds the preset load threshold τ load hour, Decrease, relax non-critical constraints; When the emotion change exceeds the threshold τ emotion hour, Incrementally, strengthen the emotion mapping constraints.
7. The computer terminal according to claim 6, wherein: In step S4, the method for generating personalized therapeutic music includes the following specific steps: S41. Using a feature fusion encoder, integrating the fusion feature and the facial emotion feature, and constructing a basic feature representation for music generation; The feature fusion encoder is composed of N identical computational blocks stacked together. Each computational block includes: residual connections and layer normalization operations, a feedforward neural network layer, and a multi-head attention layer. The residual connections and layer normalization operations are used to implement feature normalization and residual connections; the feedforward neural network layer is used to improve feature representation capabilities through nonlinear transformations; and the multi-head attention layer is used to capture long-range dependencies between features. The attention mechanism is used to inject facial emotion features between the encoding layers of the feature fusion encoder to achieve progressive fusion of emotion features. The attention mechanism obtains the query vector Q, key vector K, and value vector V through three independent linear transformations. The calculation formula is: Q=W q h l K=W k H align V=W v H align Where h l is the hidden state of the lth encoding layer, W q ,W k ,W v are the weight matrices for query, key, and value respectively, H align H bridge The facial emotion feature representation obtained after processing through the linear layer, normalization layer and activation function: H align =LayerNorm(ReLU(W align ·H bridge +b align )) Where W align and b align They are the weight matrix and bias vector of the alignment transformation, LayerNorm is the normalization layer, and ReLU is the activation function; S42. Based on the basic feature representation output by the feature fusion encoder, music structure planning is performed to form a structural framework of the music; Among them, the structural framework of music S frame The expression formula is: S frame =G struct (Z,C style ,R rules ) Where G struct is the structure generation function; Z is the basic feature representation output by the feature fusion encoder; C style is the style constraint; R rules for a set of music theory rules; The structural planning adopts a multi-dimensional design, which can be expressed as: S frame =[S sect ;S harm ;S rhyt ] Where S sect For the musical form structure; S harm To provide a framework for the harmonic progression; S rhyt Organize the structure for the rhythm; S43. Using a three-layer cascade decoding architecture, the basic feature representation Z and the structural framework S frame To decode, the expression formula is: L struct =DecoderBlock struct (Z,S frame ) L content =DecoderBlock content (L struct ) L detail =DecoderBlock detail (L content ) Where, DecoderBlock struct 、DecoderBlock content and DecoderBlock detail Respectively represent the structure layer decoding process, content layer decoding process and detail layer decoding process, L struct , L content and L detail They are the outputs of the corresponding layers respectively; S44. Output feature L of detail layer detail Input the pre-trained EnCodec decoder to generate a high-fidelity audio waveform W'. The process is expressed as: W’=Decoder En1odec (Quantize(L detail ),Φ) In the formula, Quantize is a vector quantization module that converts the continuous feature L detail Mapped to discrete audio tags; Φ is the decoder parameter, which restores the audio waveform through deconvolution and upsampling; Decoder EnCodec The pre-trained EnCodec decoder is responsible for restoring the quantized audio tokens to a high-fidelity waveform W'; The music generation process in step S4 adopts the following optimization and constraint mechanism: (1) Musical continuity constraints, namely: C music (t)≥C min Where C music (t) is the music coherence measure at time t, C min is the minimum coherence threshold; (2) Emotional expression constraints, namely: ‖ΔE(t)‖≤δ 1max Where ‖·‖ is the L2 norm; ΔE(t) is the expression change vector at time t, δ 1max is the maximum allowable variation.
8. The computer terminal according to claim 7, wherein: Step S5 includes the following specific steps: S51. Using the same ViT model as in step S3 to monitor the user's facial emotional features in real time, thereby calculating the expression change vector, emotional state vector and attention focus; The calculation formula of the expression change vector is: ΔE(t)=F face (t)-F face (t-1) Where, F face (t) and F face (t-1) is the continuous frame feature extracted by the encoder of the ViT model; The calculation formula of the emotional state vector is: S e (t)=[valence(t),arousal(t)] Where S e (t) is the emotional state vector at time t; valence(t) is the pleasure value at time t, ranging from [-1, 1], with positive values indicating positive emotions and negative values indicating negative emotions; arousal(t) is the arousal value at time t, ranging from [-1, 1], with positive values indicating high energy states and negative values indicating low energy states; The formula for calculating attention concentration is: A f (t)=attention score (F r (t)) Where a f (t) is the attention score at time t, ranging from [0,1]; attention scorre (·) is the evaluation processing of the attention component, which evaluates the user's concentration on music by analyzing eye movements, head posture and micro-expression changes; F r (t) is the set of key facial features at time t; S52. Perform feature analysis on the generated music in real time, and the expression formula is: M r (t)=MusicAnalyzer(Audio gen (t)) Where M r (t) is the music feature vector at time t, Audio gen (t) is the audio signal generated at time t; MusicAnalyzer(·) is a music analyzer, which is used to analyze the rhythm dynamics, harmonic complexity and timbre characteristics of the generated music; S53. The intensity of the input emotional features and the output music parameters is adjusted through a bidirectional gating mechanism. The calculation formula is: G in (t)=σ(W g [F r (t);M r (t)]+b g ) G out (t)=σ(W o [ΔE(t);S e (t);A f (t)]+b o ) Where G in (t) and G out (t) are the characteristic input control gate and parameter output modulation gate at time t; W g is the input weight matrix, W o is the output weight matrix, b g is the input bias term; b o is the output bias term; σ is the sigmoid activation function, which is used to normalize the gate value to the [0,1] interval; S54. Use a PID controller to dynamically adjust the music parameters according to the emotional deviation. The calculation formula is: u(t)=K p e(t)+K i ∫e(t)dt+K d de(t) / dt Where u(t) is the output control signal of the PID controller; e(t) is the error between the target emotional state and the current state; K p ,K i ,K d are proportional term, integral term and differential term respectively; when e(t) exceeds the preset threshold, the control intensity is increased to speed up the emotional adjustment, otherwise the control intensity is reduced to maintain a stable state; S55. Comprehensively evaluate the effect of music therapy through multi-dimensional indicators to form a quantitative treatment effect function: Where, E therapy (t) represents the effect of music therapy at time t; and are the weight coefficients of expression change, emotional state and attention level, The weight is adjusted according to different treatment stages and goals; among them, it is increased during the emotional stability period. Increased during emotional expression Increased during focused training S56. The music generation parameters are updated according to the treatment effect. The expression formula of the update process is: Where θ music (t+1) and θ music (t) are the music generation parameters at time t+1 and t, η is the learning rate, is the gradient of treatment effect at time t.
9. The computer terminal according to claim 8, wherein: The following optimization and constraint mechanisms are used in the closed-loop control of step S5: (1) Minimize treatment bias, namely: Where, E target For the target treatment effect; (2) Musical continuity constraints, namely: C music (t)≥C min Where C music (t) is the music coherence measure at time t; C min is the minimum coherence threshold; (3) Emotional smoothness, namely: ‖ΔS e (t)‖≤δ 2max Where, ΔS e (t) is the change in emotional state; δ 2max is the maximum allowable variation; (4) Response delay, that is: Δt response ≤τ max Where, Δt response is the response time; τ max is the maximum allowed response delay.
Citation Information
Patent Citations
Smart home equipment music recommendation method based on continuous emotion of user
CN118245629A
Immersive music experience through synchronized auditory and haptic feedback tailored to cognitive and emotional states
WO2024180549A1