Emotional music adjusting method and terminal based on dynamic perception and multi-modal fusion
Through the combination of dynamic perception field and emotion-music bridge adapter and dynamic feedback mechanism, intelligent processing from user characteristics to emotional music generation is achieved, solving the shortcomings of user feature modeling, emotion-music mapping accuracy and treatment process control mechanism in the existing technology, and significantly improving the personalization and effect of music therapy.
Patent Information
- Application Number
- CN202510600232.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing music therapy technology has shortcomings in user feature modeling, emotion-music mapping accuracy and treatment process control mechanism, making it difficult to achieve personalized and efficient music therapy.
The emotional music adjustment method based on dynamic perception and multimodal fusion is adopted, and the full process of intelligent processing from user characteristics capture to emotional music generation is realized through dynamic perception field, emotion-music bridge adapter and dynamic feedback mechanism.
It significantly improves the integrity of user feature expression and the accuracy of emotion-music mapping, enhances the control ability of the treatment process and personalized treatment effects, and improves the user experience.
Smart Images

Figure CN120114726A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and image processing, and specifically, to an emotional music adjustment method and terminal based on dynamic perception and multimodal fusion. Background Art
[0002] As a non-invasive psychological intervention means, music therapy affects the human nervous system and endocrine system through adjustable music sound wave stimulation, so as to achieve the purpose of regulating emotions and improving physiological states. Research shows that music can activate brain regions related to emotion regulation such as the prefrontal cortex and amygdala of the brain, promote the secretion of neurotransmitters such as dopamine and serotonin, thereby affecting the human autonomic nervous system and endocrine system, and improving physiological indicators such as heart rate variability and cortisol level.
[0003] Currently, music therapy technologies mainly include the following three technical paradigms: (1) Traditional music therapy based on rules mainly relies on therapists' experience to select appropriate music segments and perform manual intervention. Although it has strong pertinence, it is inefficient and difficult to be applied on a large scale. (2) The intelligent recommendation system based on machine learning uses collaborative filtering and content-based recommendation methods to perform music matching in combination with user portraits, improving the automation degree of the system, but it is often limited to the selection of existing music libraries and is difficult to achieve true personalized treatment. (3) The generative therapy system based on deep learning uses models such as LSTM and Transformer to achieve music generation and emotion mapping. Although it has a certain degree of creativity, there is still room for improvement in terms of the controllability of music quality and treatment effect.
[0004] With the development of artificial intelligence technology, music therapy systems have made remarkable progress in aspects such as feature extraction, emotion mapping, and feedback control. Researchers have begun to try to introduce attention mechanisms to achieve adaptive adjustment of feature weights, use variational autoencoders to construct mapping relationships between emotions and music features, and optimize treatment strategies through reinforcement learning methods. However, the existing technologies still have the following deficiencies: 1. Insufficient user feature modeling: Existing systems mostly adopt static or fixed-weight feature extraction methods, which are difficult to effectively perceive the dynamic changes of user states. Although attention mechanisms are introduced to enhance feature selection capabilities, due to the lack of modeling of deep associations between features, inaccurate judgments of feature importance may still occur. In addition, existing systems lack consideration of users' multidimensional background features (such as personal experiences, cultural backgrounds, emotional tendencies, etc.), affecting the effect of personalized treatment.
[0005] 2. Low emotional-music mapping accuracy: Existing methods generally adopt a single mapping model to convert user features into music generation parameters, making it difficult to depict complex non-linear mapping relationships. Although some studies have attempted to use variational autoencoders to transform the feature space, they lack effective semantic constraints and music quality control mechanisms, resulting in poor performance in terms of the accuracy of emotional expression and artistry of the generated music.
[0006] 3. Weak treatment process control mechanism: Existing systems mostly rely on threshold judgment methods to adjust music parameters and lack systematic control strategies. Although reinforcement learning methods have been introduced to optimize the control process, due to the imperfect feedback mechanism and the lack of an evaluation system based on multi-modal signals, it is difficult to achieve accurate evaluation and continuous optimization of the treatment effect. In addition, the system has not yet formed an effective mechanism for the accumulation, extraction, and transfer application of treatment experience, which limits the overall adaptability and intelligence level.
[0007] Therefore, there is an urgent need for a personalized music generation solution that can comprehensively perceive the multi-dimensional state characteristics of users, accurately establish the mapping relationship between emotions and music, and have a perfect feedback control ability to improve the treatment effect and user experience. Summary of the Invention
[0008] To solve the technical problems existing in the prior art, the present invention provides an emotional music adjustment method and terminal based on dynamic perception and multi-modal fusion. The present invention constructs a complete end-to-end generation framework, mainly including a user feature input layer, a dynamic perception field, an emotion-music bridging adapter, a music generation module, and dynamic feedback. Through the organic connection of the feature flow, the mapping flow, and the feedback flow, it realizes the full-process intelligent processing from user feature capture to emotional music generation, improving the treatment effect and user experience.
[0009] To achieve the above object, the present invention provides the following technical solutions: The present invention discloses an emotional music adjustment method based on dynamic perception and multi-modal fusion, including the following steps: S1. Obtain multi-dimensional feature data of the user, including user portraits, personal hobbies, historical experiences, and physiological data; S2. Perform dynamic weight allocation on the multi-dimensional feature data through the dynamic perception field to generate fused features; S3. Real-time collect the user's facial image data, extract facial emotion features through the ViT model, and map the facial emotion feature space to the music feature space through the emotion-music bridging adapter; the facial emotion features include expression encoding, emotion dimension, emotion intensity, and temporal features; S4. Input the fused features and the facial emotion features mapped to the music feature space into the music generation module, and generate personalized therapeutic music through the feature fusion encoding, music structure planning, and hierarchical decoding architectures in this module; S5. Monitor the user's facial emotion features and the generated music features in real time, and adjust the music generation parameters through a dynamic feedback mechanism to form a closed-loop control.
[0010] As a further improvement of the above solution, in step S1, through four LlaMA encoders, the text descriptions in the four dimensions of the user portrait, personal hobbies, historical experiences, and physiological data are processed respectively, and the text descriptions are converted into standardized encoded feature vectors, so as to form the multi-dimensional feature data.
[0011] As a further improvement of the above solution, step S2 includes the following specific steps: S21. Perform a splicing operation on the four types of encoded feature vectors of the user portrait, personal hobbies, historical experiences, and physiological data to construct a unified feature representation space, realize the preliminary fusion of features, and the expression formula is: ; In the formula, E is the fused feature vector, with a dimension of n×4, and n is the dimension of a single type of feature vector; , , and are the encoded feature vectors of the user portrait, personal hobbies, historical experiences, and physiological data respectively; S22. Construct a dynamic perception field to realize the dynamic perception of feature importance; Among them, the weighting module in the dynamic perception field is used to initialize the weights of the four types of features, evenly distribute the feature weights within each dimension, and the weight set W is defined as: ; In the formula, correspond to the weight coefficients of the four types of features respectively, , , ; The expression formula of the perception field intensity function of the dynamic perception field is as follows: ; In the formula, is the perception field intensity function; t is time; and are the direct influence term and the interaction influence term respectively, and the expression formulas are: ; ; In the formula, is the time decay factor; ; is the feature correlation. The cosine similarity is used to quantify the encoded feature vector and The degree of association between them ranges from [0, 1]. The larger the value, the stronger the feature correlation; is the dynamic coupling coefficient. The gradient information is used to characterize the influence direction of the feature on the perception field. The gradient product is used to reflect the degree of cooperation of feature changes, and the sigmoid function is used to normalize the value range of the dynamic coupling coefficient to [0, 1]; S23. An improved momentum gradient descent method is used to update the feature weights, and the updated feature weights are used to adaptively fuse the multi-dimensional features to generate fused features; among them, the expression formula for weight update is: ; In the formula, and are the feature weights before and after update, respectively; is the learning rate; is the momentum factor; is the second-order gradient, is the first-order gradient.
[0012] As a further improvement of the above solution, in step S3, the encoder of the ViT model receives the facial image sequence as input and outputs the encoded facial emotion feature vector. The expression formula for the feature extraction process is: ; In the formula, is the facial emotion feature vector; represents the encoding process of the ViT encoder; is the facial image sequence; is the th feature component in the facial emotion feature vector. Each component represents a specific aspect in the emotion feature space, and together they constitute a complete representation of the four dimensions of expression encoding, emotion dimension, emotion intensity, and temporal sequence features, , m is the total number of feature components.
[0013] As a further improvement of the above solution, in step S3, the emotion-music bridging adapter adopts a dual-channel design, including an uplink adaptation and a downlink adaptation processing channel; Among them, the up - adaptation channel is used to elevate and transform facial emotion features to the music feature space through a projection layer, ReLU activation, and layer normalization. The expression formula is: ; In the formula, is the feature representation output by the up - adaptation channel; represents layer normalization processing; represents ReLU activation processing; is the projection weight matrix of the up - adaptation channel, , is the dimension of the music feature space, is the dimension of facial emotion features; is the bias parameter of the up - adaptation channel; is the set of real numbers; The down - adaptation channel is used to compress the music feature dimension through layer normalization, projection layer, and gated unit, and extract the core features related to emotions. The expression formula is: ; In the formula, is the feature representation output by the down - adaptation channel; is the gated unit processing, which is used to selectively retain key features; is the representation of the music feature space; is the projection weight matrix of the down - adaptation channel, ; is the bias parameter of the down - adaptation channel; The up - adaptation channel and the down - adaptation channel work together and are integrated through residual connection. The expression formula is: ; In the formula, is the bridging feature; As a further improvement of the above - mentioned solution, in step S3, the emotion - music bridging adapter adopts the following optimization and constraint mechanism: (1) Computational efficiency objective: Minimize the latency introduced by the emotion - music bridging adapter, that is: ; In the formula, the subscript is the set of trainable parameters of the emotion - music bridging adapter, including the up - channel parameters , the down - channel parameters and the internal parameters of the gated unit; represents the expected - value operator, which is used to calculate the average performance of a random variable; is the latency time introduced by the emotion - music bridging adapter; (2) Emotional mapping accuracy objective: Maximize the consistency between emotional expression and musical features, i.e.: ; In the formula, represents the facial emotional feature vector and the semantic consistency measure of the musical parameter features ; (3) Musical quality objective: Maximize the musical quality under the delay constraint, i.e.: ; In the formula, is the generated musical quality measure; is the upper limit threshold of the delay time; Among them, the musical quality measure is defined as a multi-factor comprehensive score, and the calculation formula is: ; In the formula, is the weight coefficient; is the musical structure integrity score, used to measure the musical form logic; is the melody fluency score, used to evaluate the melody naturalness and memorability; is the harmony coordination score, used to evaluate the balance and richness of the harmony progression; is the emotional expressiveness score, measuring the matching degree between the music and the target emotion; The emotion-music bridging adapter adopts the following multi-objective balance strategy: (1) Hierarchical priority strategy, dynamically adjusting the weights of the optimization objectives according to the user scenario requirements: ; In the formula, is the overall optimization objective function; represents the th individual optimization objective; is the parameter set of the emotion-music bridging adapter; is the context-related weight coefficient, ; is the current interaction scenario type; When in the real-time treatment scenario, > > , giving priority to ensuring the calculation efficiency; When in the quality-priority scenario, > > , giving priority to ensuring the musical quality; When in the emotional response field, > > , prioritize ensuring the accuracy of emotional mapping; (2) The constraint relaxation strategy allows for softening of some constraint conditions in extreme scenarios. The expression formula is: ; In the formula, is the soft constraint penalty term; is the constraint relaxation coefficient; is the th constraint condition, , corresponding to the three constraint conditions of computational efficiency, emotional mapping accuracy, and music quality respectively; When the system load exceeds the preset load threshold , decreases, relaxing non-critical constraints; When the emotional change amplitude exceeds the threshold , increases, strengthening the emotional mapping constraint.
[0014] As a further improvement of the above solution, in step S4, the method for generating personalized therapeutic music includes the following specific steps: S41. Use a feature fusion encoder to integrate the fusion features and the facial emotional features, and construct a basic feature representation for music generation; Among them, the feature fusion encoder is composed of N identical computational blocks stacked. Each computational block includes: a residual connection and layer normalization operation, a feedforward neural network layer, and a multi-head attention layer. The residual connection and layer normalization operation are used to achieve feature normalization and residual connection; the feedforward neural network layer is used to improve the feature representation ability through non-linear transformation; the multi-head attention layer is used to capture the long-range dependence relationship between features; Inject the facial emotional features between the encoding layers of the feature fusion encoder through the attention mechanism to achieve progressive fusion of emotional features; among them, the attention mechanism obtains the query vector Q, key vector K, and value vector V through three independent linear transformations. The calculation formula is: ; ; ; In the formula, is the hidden state of the l th encoding layer, are the weight matrices of the query, key, and value respectively, is the facial emotional feature representation obtained after processing through a linear layer, a normalization layer, and an activation function: ; In the formula, and are the weight matrix and bias vector for alignment conversion respectively, is the normalization layer, is the activation function; S42. Based on the basic feature representation output by the feature fusion encoder, perform music structure planning to form the structure framework of the music; Among them, the structure framework of the music The expression formula is: ; In the formula, is the structure generation function; is the basic feature representation output by the feature fusion encoder; is the style constraint condition; is the music theory rule set; The structure planning adopts a multi-dimensional design, expressed as: ; In the formula, is the musical form structure; is the harmonic progression framework; is the rhythm organization structure; S43. Adopt a three-level cascaded decoding architecture to decode the basic feature representation Z and the structure framework The expression formula is: ; ; ; In the formula, , and respectively represent the decoding processing of the structure layer, the decoding processing of the content layer and the decoding processing of the detail layer, , and are the outputs corresponding to each layer respectively; S44. Input the detail layer output feature into the pre-trained EnCodec decoder to generate a high-fidelity audio waveform W’ , and the process is expressed as: ; In the formula, Quantize is the vector quantization module, which maps the continuous feature to discrete audio tokens; Φ is the decoder parameter, which restores the audio waveform through transposed convolution and upsampling; It is a pre-trained EnCodec decoder, responsible for restoring the quantized audio tokens to high-fidelity waveforms W’ ; Among them, the music generation process in step S4 adopts the following optimization and constraint mechanisms: (1), Music coherence constraint, that is: ; In the formula, is the music coherence metric at time t, is the minimum coherence threshold; (2), Emotional expression constraint, that is: ; In the formula, is the L2 norm; is t the expression change vector at time is the maximum allowable change amplitude.
[0015] As a further improvement of the above solution, step S5 includes the following specific steps: S51. Use the same ViT model as in step S3 to monitor the user's facial emotional features in real time, so as to calculate the expression change vector, emotional state vector and attention focus; Among them, the calculation formula of the expression change vector is: ; In the formula, and are the continuous frame features extracted by the encoder of the ViT model; The calculation formula of the emotional state vector is: ; In the formula, is t the emotional state vector at time is t the pleasure value at time, ranging from [-1,1], positive value indicates positive emotion, negative value indicates negative emotion; is t the arousal value at time, ranging from [-1,1], positive value indicates high energy state, negative value indicates low energy state; The calculation formula of attention focus is: ; In the formula, is t the attention score at time, ranging from [0,1]; is the evaluation process of the attention component, which evaluates the user's attention to music by analyzing eye movement, head posture and micro-expression changes; is t the set of facial key features at a moment; S52. Analyze the features of the generated music in real time, and the expression formula is: ; In the formula, is t the music feature vector at a moment, is t the audio signal generated at a moment; is a music analyzer, which analyzes the rhythm dynamics, harmony complexity and timbre features of the generated music; S53. Adjust the intensity of the input emotional features and the output music parameters through a bidirectional gating mechanism, and the calculation formula is: ; ; In the formula, and are respectively t the feature input control gate and the parameter output modulation gate at a moment; is the input weight matrix, is the output weight matrix, is the input bias term; is the output bias term; is the sigmoid activation function, which is used to normalize the gating value to the interval [0,1]; S54. Use a PID controller to dynamically adjust the music parameters according to the emotional deviation, and the calculation formula is: ; In the formula, is the output control signal of the PID controller; is the error between the target emotional state and the current state; are respectively the proportional term, the integral term and the differential term; among them, when exceeds the preset threshold, increase the control intensity to accelerate the emotional adjustment, otherwise reduce the control intensity to maintain the stable state; S55. Comprehensively evaluate the music therapy effect through multi-dimensional indicators to form a quantified therapy effect function: ; In the formula, represents t the music therapy effect at a moment; , and are respectively the weight coefficients of the expression change, the emotional state and the attention level, , adjust the weights according to different treatment stages and goals; among them, increase during the emotional stability period , increase during the emotional expression period , increase during the focused training period ; S56. Update the music generation parameters driven by the treatment effect, and the expression formula of the update process is: ; In the formula, and are respectively and the music generation parameters at time , is the learning rate, is the treatment effect gradient at time
[0016] As a further improvement of the above solution, the following optimization and constraint mechanism is adopted in the closed-loop control of step S5: (1), Minimize the treatment deviation, that is: ; In the formula, is the target treatment effect; (2), Music coherence constraint, that is: ; In the formula, is t the music coherence metric at time is the minimum coherence threshold; (3), Emotional smoothness, that is: ; In the formula, is the emotional state change amount; is the maximum allowable change amplitude; (4), Response delay, that is: ; In the formula, is the response time; is the maximum allowable response delay.
[0017] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the emotional music adjustment method based on dynamic perception and multi-modal fusion as described above are implemented.
[0018] Compared with the prior art, the beneficial effects of the present invention are: 1. The emotional music adjustment method based on dynamic perception and multimodal fusion proposed by the present invention mainly aims at the problems of incomplete feature perception, inaccurate emotional mapping, and unstable treatment effects in existing music therapy methods. By innovatively designing core technologies such as a dynamic perception field mechanism, an emotion-music bridging adapter, and an adaptive feedback control, the present invention realizes the full-process intelligent processing from user feature perception to emotional music generation. The present invention adopts a modular design idea, gradually optimizing in aspects such as feature processing, spatial mapping, and feedback control, and finally constructs a complete end-to-end treatment framework.
[0019] 2. In terms of feature perception, the present invention processes four types of features, namely user portrait, personal hobbies, historical experiences, and physiological data, through a LLaMA encoder respectively, and introduces a dynamic perception field mechanism to realize the adaptive calculation of feature importance, significantly improving the integrity of feature expression and remarkably enhancing the accuracy of feature weight allocation.
[0020] 3. In terms of emotion-music mapping, the present invention innovatively designs a bridging adapter structure, and through facial expression feature extraction and emotional state mapping, realizes the precise conversion from the feature space to the music parameter space. Compared with the prior art, the facial expression feature extraction has a higher accuracy rate, the dimension coverage of the emotional state mapping is more comprehensive, and the accuracy of the music parameter mapping is significantly improved. Through strict parameter constraints and adaptive temperature control, the normativity and creativity performance of the generated music are also significantly improved.
[0021] 4. In terms of treatment effects, the present invention introduces a PID controller and an adaptive feedback adjustment strategy to realize the precise control and continuous optimization of treatment parameters. The control accuracy is significantly improved, the parameter optimization efficiency is obviously improved, and the treatment effect evaluation is more accurate. The optimized configuration of the feature processing parameters and control parameters makes the system operation more stable, effectively reduces the volatility, significantly improves the treatment safety, and provides a more reliable treatment experience for patients.
[0022] 5. Through the organic combination of the above technological innovations, the present invention has made breakthrough progress in aspects such as feature perception ability, emotion mapping accuracy, and treatment effects, and the overall performance has been significantly improved. The present invention not only provides a reliable music therapy technology solution, but its modular design also creates conditions for function expansion and performance optimization, and has important application value and promotion significance.
[0023] 6. The computer terminal disclosed by the present invention can produce the same effects as the above method by applying the above method, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of the emotional music adjustment method based on dynamic perception and multimodal fusion in Embodiment 1 of the present invention.
[0025] Figure 2 This is the system block diagram of the emotional music adjustment method based on dynamic perception and multi-modal fusion in Embodiment 1 of the present invention. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] Embodiment 1 In view of the technical problems existing in the prior art, the present invention proposes an innovative solution: First, a multi-level feature extraction architecture based on a dynamic perception field is designed to achieve adaptive analysis and dynamic adjustment of feature importance by fusing multi-dimensional features such as user portraits, historical experiences, personal preferences, and real-time physiological data. Second, an emotion-music bridge based on a hierarchical memory network is constructed, combined with a multi-modal cross-attention mechanism, to achieve an accurate mapping between emotion features and music generation parameters while ensuring the artistry of music. Finally, a complete feedback control loop is established to achieve real-time adjustment and experience accumulation of treatment strategies through multi-modal signal processing and parameter dynamic optimization mechanisms, thereby significantly improving the personalized treatment effect.
[0028] Please refer to Figure 1 and Figure 2 , this embodiment provides an emotional music adjustment method based on dynamic perception and multi-modal fusion, including the following steps: S1. Obtain multi-dimensional feature data of the user, including user portraits, personal hobbies, historical experiences, and physiological data; S2. Perform dynamic weight allocation on the multi-dimensional feature data through a dynamic perception field to generate fused features; S3. Real-time collect user facial image data, extract facial emotion features through a ViT model, and map the facial emotion feature space to the music feature space through an emotion-music bridge adapter; the facial emotion features include expression encoding, emotion dimensions, emotion intensity, and temporal features; S4. Input the fused features and the facial emotion features mapped to the music feature space into a music generation module, and generate personalized treatment music through the feature fusion encoding, music structure planning, and hierarchical decoding architectures in this module; S5. Real-time monitor the user's facial emotion features and the generated music features, and adjust the music generation parameters through a dynamic feedback mechanism to form a closed-loop control.
[0029] In the overall architecture design, the present invention innovatively adopts a dual-channel feature input mechanism to achieve deep collaboration between dynamic perception field features and real-time facial expression features. Among them, the dynamic perception field module is responsible for comprehensively processing the multi-dimensional features of users, including user portraits (such as basic information, psychological characteristics, treatment background), personal hobbies (such as art and culture, sports and health, cognitive exploration), historical experiences (such as growth and development, emotional interaction, music experience), and physiological data (such as vital signs, neuroendocrine, physical functions), etc. Through the processing of deep learning models, these features are converted into a unified representation containing fused feature vectors, feature weight states, importance scores, and interaction matrices, and after feature alignment, normalization processing, and temporal supplementation, a basic feature representation required for music generation is formed. In terms of facial expression feature processing, the continuous changes in the user's facial expressions are captured through video analysis, and complete expression features including expression encoding, emotional dimensions (pleasure and arousal), emotional intensity, and temporal changes are extracted. Based on the user's multi-dimensional features and real-time facial expression features, a special bridging adapter is innovatively designed to inject these facial expression features between the blocks of the music generation module encoder (Encoder). The adapter ensures that the facial emotional features can be fully utilized during the encoding process through feature alignment and attention mechanisms, thereby achieving progressive fusion and enhancement of features. This design can not only ensure the input perception of perception field features but also realize the dynamic injection of facial expression features, which helps to improve the emotional expression ability.
[0030] To achieve real-time optimization and dynamic adjustment of emotional music generation, the present invention constructs a feedback mechanism based on bidirectional gating. This mechanism establishes a complete feedback loop by continuously monitoring the changes in facial emotional features (including the amplitude of expression changes, emotional state vectors, and attention focus) and the features of the generated music (including rhythm dynamics, harmony complexity, and timbre features). The present invention innovatively designs an input gate and an output gate, and realizes the dynamic adjustment of the feature flow intensity through a learnable weight matrix, ensuring the real-time matching of the generated music with the user's emotional state. In specific implementation, the input gate is responsible for adjusting the input feature flow according to the matching degree between the facial emotional features and the music features, while the output gate optimizes the music generation parameters based on the real-time monitored expression changes, emotional states, and attention levels. When a significant deviation between the emotional state and the target is detected, a quick response is achieved by adjusting the gate opening and the parameter adjustment step size; when the deviation is within an acceptable range, the parameters are kept stable and fine-tuned to ensure the coherence of music generation. The entire closed-loop control process is precisely adjusted through a PID controller, which not only ensures the real-time response ability but also ensures the smooth transition of emotional expression.
[0031] Through the feature integration of the dynamic perception field, the feature injection of the bridging adapter, and the feedback regulation of the gating mechanism, the deep integration of emotional features and music generation is achieved, providing strong technical support for emotional music therapy. The specific implementation solutions of each module will be described in detail below.
[0032] 1. Feature capture of the dynamic perception field The main function of the dynamic perception field is to achieve multi-dimensional capture of user features, dynamic importance perception, and adaptive feature fusion. Its main innovation lies in processing unstructured text based on large language models to achieve in-depth understanding and dynamic weight allocation of users' comprehensive features.
[0033] 1.1 Multi-dimensional feature vector capture To achieve a comprehensive perception of user features, the present invention uses four LLaMA encoders to process text descriptions in four dimensions: user profile, personal hobbies, historical experiences, and physiological data, and converts them into standardized feature vectors. As Figure 2 shown, the LLaMA encoder receives four types of input information and outputs corresponding feature representations. The basic mapping process of each encoder can be expressed as: ; In the formula, is the encoded feature vector, X is the original text input. To ensure computational efficiency, the present invention uses a lightweight version of LLaMA to convert the original text into 768-dimensional feature vectors.
[0034] The four types of features are: user profile information is encoded as , personal hobby information is encoded as , historical experience information is encoded as , and physiological data information is encoded as . The specific content of the four types of features is as follows: (1) User profile information: It mainly covers the following four dimensions: basic information: including age, gender, occupation, education level, and marital status; psychological characteristics: including personality type, depression tendency score, anxiety level score, stress tolerance level, and emotional stability score; treatment background: including past medical history, treatment stage, treatment goal, expected treatment course, and treatment frequency; social background: including cultural background, language habits, living environment, family structure, and social activity level.
[0035] (2) Personal hobbies: It mainly covers the following four dimensions: Art and cultural hobbies: including music appreciation preferences, visual art interests, performing art inclinations, literary creation enthusiasm, and traditional culture preferences; Sports and health hobbies: including sports habits, health preservation methods, meditation and relaxation preferences, outdoor activity interests, and physical training inclinations; Cognitive exploration hobbies: including knowledge learning directions, scientific and technological innovation interests, thinking challenge preferences, skill cultivation enthusiasm, and professional research inclinations; Life and entertainment hobbies: including social activity forms, leisure and entertainment methods, handicraft interests, collection and appreciation preferences, and life experience inclinations.
[0036] (3)Historical experiences: It mainly covers the following four dimensions: Growth and development experiences: including family education experiences, school growth trajectories, important life turning points, stress coping experiences, and self - awareness changes; Emotional interaction experiences: including important emotional relationships, interpersonal communication patterns, social support networks, emotional trauma experiences, and emotion regulation strategies; Music experience experiences: including music learning experiences, music activity participation, music emotional memories, music therapy experiences, and music scene experiences; Life adaptation experiences: including career development processes, environmental adaptation abilities, habit formation processes, interest and hobby changes, and lifestyle evolutions.
[0037] (4)Physiological data: It mainly covers the following four dimensions: Basic vital signs: including heart rate and blood pressure levels, respiratory rhythm changes, body temperature and metabolic indicators, height - to - weight ratios, and physical energy and vitality states; Neuroendocrinology: including brain wave activities, neurotransmitter levels, hormone secretion conditions, stress response degrees, and physiological rhythm manifestations; Physical functions: including immune system functions, organ operation states, movement coordination abilities, sensory system sensitivities, and musculoskeletal conditions; Health status: including disease diagnosis records, chronic disease degrees, sleep quality levels, rehabilitation treatment progress, and complication risk assessments.
[0038] 1.2. Dynamic perception field construction After the present invention completes the feature encoding in the four dimensions of user portrait, personal hobbies, historical experiences, and physiological data, through feature fusion and dynamic perception field technology, it realizes the dynamic perception of the importance of user features. As Figure 2 shown, the four types of features after LLaMA encoding are processed and fused through a weighting module.
[0039] 1.2.1. Multi - dimensional feature fusion The present invention first conducts preliminary fusion of multi - dimensional features, and then constructs a dynamic perception field to realize the dynamic perception of feature importance.
[0040] In the initial feature fusion stage, the user portrait ( ), personal hobbies ( ), historical experiences ( ), and physiological data ( ) of the four types of encoded feature vectors are concatenated to construct a unified feature representation space, realizing the initial fusion of features. Among them, E is the fused feature vector, , with a dimension of n×4, where n is the dimension of a single type of feature vector (n = 768).
[0041] Subsequently, a dynamic perception field is constructed to realize the dynamic perception of user features, and the Weighting Module corresponds to the core component of the dynamic perception field. First, the weights of the four types of features are initialized, and the feature weights within each dimension are evenly distributed. The weight set W is defined as: ; where correspond to the weight coefficients of the four types of features respectively, , .
[0042] It should be noted that the above initial fusion stage is a low-level, structural feature fusion. By using a simple vector concatenation operation, the complete information of the original features remains unchanged, the weights of each feature are not different, and the dimension is extended from N dimensions to N*4. Simply put, it can be understood as the "physical concatenation" of features, similar to simply stacking different information sources together, providing the basic data structure for subsequent dynamic weight allocation and deep fusion.
[0043] The adaptive fusion of the dynamic perception field is a high-level, semantic feature fusion. It uses adaptive weighted fusion based on dynamic weights to dynamically adjust the feature weights according to the results of the initial fusion: continuously optimizing the weight allocation through the gradient descent method.
[0044] 1.2.2. Perception Field Intensity Function The dynamic perception field dynamically perceives the importance of features in each dimension by designing a perception field intensity function . This function includes a direct influence term and an interaction influence term as two core parts, which are used to capture the independent and interactive effects of features respectively: ; Among them, the direct influence term reflects the direct contribution of each feature to the perception field, and the expression formula is: ; This item comprehensively considers the three key elements of feature weight, feature vector and time decay. The importance of each feature is reflected by weighting the feature weight and feature vector, and the time decay factor is introduced to enable the feature impact strength to be dynamically adjusted over time. All weighted features are summed to obtain the overall direct impact strength.
[0045] Interaction Effects The focus is on describing the interaction between features, and the expression formula is: ; In the formula, is the time decay factor; ; For feature correlation, cosine similarity is used to quantify the encoded feature vector and The degree of correlation between them is in the range of [0,1]. The larger the value, the stronger the feature correlation. is the dynamic coupling coefficient. The gradient information is used to characterize the direction of the feature's influence on the perception field. The gradient product is used to reflect the degree of coordination of feature changes. The sigmoid function is used to normalize the value range of the dynamic coupling coefficient to [0,1].
[0046] The interaction term is modeled by two factors: feature correlation and dynamic coupling coefficient. Cosine similarity is used to quantify the degree of correlation between features. The value range is limited to the interval [0,1]. The larger the value, the stronger the feature correlation. In the calculation process, special consideration is given to the directional consistency of the feature vector to eliminate the influence of dimensional differences. Dynamic coupling coefficient The gradient information is used to characterize the direction of the feature's influence on the perception field, the gradient product is used to reflect the degree of coordination of feature changes, and the coefficients are normalized to the [0,1] interval through the sigmoid function. Finally, the interaction effects of all feature pairs are accumulated to capture the complex interaction relationship and combination effect between features.
[0047] 1.2.3 Dynamic Weight Optimization In order to achieve dynamic optimization of feature weights, the present invention adopts an improved momentum gradient descent method to update weights. This method improves the efficiency and stability of weight optimization by introducing momentum terms and second-order gradient information: ; In the formula, and are the feature weights before and after updating respectively; is the learning rate; is the momentum factor; is the second-order gradient; the learning rate Control the step size of weight update, with a default value of 0.01, and the momentum factor Utilize historical gradient information to accelerate convergence, with a default value of 0.9, by introducing second-order gradients , achieving an accurate grasp of the optimization direction and effectively avoiding local optimal solutions. is the first-order gradient.
[0048] Through the above design, the dynamic perception field can capture multi-dimensional features of users and perform adaptive fusion, providing optimized feature representations for the subsequent emotion-music bridge, which is the basic module of the entire music therapy method. Figure 2 The weighted module in
[0049] The dynamic perception field mechanism is a core innovation point and the main protected content of the present invention. This mechanism processes four types of features, namely user portraits, personal hobbies, historical experiences, and physiological data, through the LLaMA model to achieve the vectorized expression of multi-dimensional features. On this basis, through feature fusion and dynamic perception field construction methods, including techniques such as feature splicing, weight initialization, and field strength calculation, the unified expression of features is realized. At the same time, a dynamic evaluation mechanism for feature importance is introduced, considering both direct and indirect impacts, and the gradient descent method with momentum is used for adaptive weight update, achieving the accurate calculation and dynamic adjustment of feature importance.
[0050] 2. Emotional-Music Bridging Adapter The emotional-music bridging adapter is the second innovation point of the present invention, and its main function is to establish a mapping relationship between facial emotional features and the music generation space. As Figure 2 shown, this module connects the facial emotional feature extraction and music generation modules, thereby realizing the deep fusion of facial emotional feature extraction and music generation.
[0051] 2.1 Facial Emotional Feature Extraction 2.1.1 ViT Model Architecture and Parameter Configuration Facial emotion is the most direct and rich carrier of human emotion expression. The present invention uses a Vision Transformer (ViT) model to process the input video frame sequence. As Figure 2 shown, the ViT encoder receives a sequence of facial images in different emotional states as input and outputs the encoded emotional feature vectors. The feature extraction process can be formally expressed as: ; In the formula, is the facial emotional feature vector; represents the encoding process of the ViT encoder; is the facial image sequence; is the th feature component in the facial emotion feature vector. Each component represents a specific aspect in the emotion feature space and jointly constitutes a complete representation of four dimensions: expression coding, emotion dimension, emotion intensity, and temporal sequence features. , m is the total number of feature components.
[0052] The model adopts the following parameter configurations: Infrastructure: ViT-B / 16 (12-layer Transformer, 16×16 image patches); Input resolution: 224×224 pixels, 32-bit color image; Temporal window: 15 consecutive frames (about 0.5 seconds, sampling rate 30fps); Output dimension: 768-dimensional emotion feature vector; Number of attention heads: 12, 64 dimensions for each head; Hidden layer dimension: 3072.
[0053] 2.1.2, Multi-dimensional Analysis of Facial Features To comprehensively capture the user's emotional state, the present invention analyzes facial features from the following four dimensions: Expression coding: Capturing micro-expression features such as facial muscle changes and eye changes, which is the basis for expression recognition; Emotion dimension: Constructing a two-dimensional emotion space through pleasure and arousal to achieve precise emotion localization; Emotion intensity: Quantifying the degree and depth of emotional expression, which determines the influence weight of emotion on music generation; Temporal sequence features: Characterizing the time dynamic characteristics of emotional changes to ensure that the generated music can smoothly transition following emotional changes.
[0054] Each feature dimension is represented by an independent sub-vector and merged into a unified feature vector through linear projection, thereby retaining multi-faceted information.
[0055] 2.2, Bridging Adapter Design The emotion-music bridging adapter is a key component connecting the facial emotion features and the music generation space. The core task of the emotion-music bridging adapter is to solve the heterogeneity problem between the emotion feature space and the music feature space. As Figure 2 shown, this adapter adopts a dual-channel design, including an upward adaptation and a downward adaptation processing path.
[0056] (1) Upward adaptation channel ( ) Lifts and transforms the facial emotion features to the music feature space through a projection layer, ReLU activation, and layer normalization: ; In the formula, is the feature representation output by the upward adaptation channel; Denote layer normalization processing; Denote ReLU activation processing; is the projection weight matrix of the uplink adaptation channel, , is the dimension of the music feature space, is the dimension of the facial emotion feature; is the bias parameter of the uplink adaptation channel; is the set of real numbers. In this embodiment , .
[0057] (2) Downlink adaptation channel ( ) Compresses the music feature dimension through layer normalization, projection layer, and gated unit to extract the core features related to emotions.
[0058] ; In the formula, is the feature representation output by the downlink adaptation channel; is the gated unit processing for selectively retaining key features; is the representation of the music feature space; is the projection weight matrix of the downlink adaptation channel, ; is the bias parameter of the downlink adaptation channel.
[0059] The two channels work together and are integrated through residual connection: ; In the formula, is the bridging feature, which is a bidirectional fusion bridging feature. It should be noted that: is the result of dimensionality increasing and converting the facial emotion feature to the music feature space, is to extract the core features related to emotions from the music feature space, The purpose of which is to achieve bidirectional mapping and information fusion between the emotion feature space and the music feature space; is the residual connection processing for retaining feature information and optimizing gradient propagation.
[0060] 2.3. Optimization and Constraint Mechanism The present invention simultaneously considers two optimization objectives of computational efficiency and generation quality: (1) Computational efficiency objective: Minimize the latency introduced by the emotion-music bridging adapter ; In the formula, the subscript is the set of trainable parameters of the emotion-music bridging adapter, including the uplink channel parameters , downlink channel parameters and the internal parameters of the gating unit; represents the expected value operator, used to calculate the average performance of a random variable; is the delay time introduced by the emotion-music bridging adapter.
[0061] (2) Emotion mapping accuracy objective: Maximize the consistency between emotion expression and music features ; In the formula, represents the facial emotion feature vector and the semantic consistency measure of the music parameter features . It should be noted that is the output of the music generation module encoder in the overall architecture. The training objective of the model is to make the recognized facial emotion features as similar as possible to the corresponding music parameter features. Here, can be understood as that in the following text.
[0062] (3) Music quality objective: Maximize music quality under delay constraints ; In the formula, is the generated music quality measure; is the upper limit threshold of the delay time. The full name of
[0063] in mathematics is "subject to", which means "constrained by". The music quality measure ; In the formula, is the weight coefficient, : Music structure integrity score (0 - 1), measuring the logical form of music; : Melody fluency score (0 - 1), evaluating the naturalness and memorability of the melody; : Harmony coordination score (0 - 1), evaluating the balance and richness of the harmony progression; : Emotion expressiveness score (0 - 1), measuring the matching degree between the music and the target emotion.
[0064] The said emotion-music bridging adapter adopts the following multi-objective balance strategy: (1), Hierarchical priority strategy, dynamically adjusting the weights of optimization objectives according to user scenario requirements: ; In the formula, is the overall optimization objective function; represents the a single optimization objective; is a parameter set of the emotion-music bridging adapter; is a context-related weight coefficient, ; is the current interaction scenario type; When in a real-time treatment scenario, > > , the computational efficiency is prioritized; When in a quality-priority scenario, > > , the music quality is prioritized; When in an emotion response field, > > , the accuracy of emotion mapping is prioritized; (2) Constraint relaxation strategy, allowing partial softening of constraint conditions in extreme scenarios, and the expression formula is: ; In the formula, is the soft constraint penalty term; is the constraint relaxation coefficient; is the th constraint condition, , corresponding to the three constraint conditions of computational efficiency, emotion mapping accuracy, and music quality mentioned above respectively; When the system load exceeds the preset load threshold , decreases, relaxing non-critical constraints; it should be noted that the system load index is calculated by weighted averaging of CPU usage rate, memory occupancy rate, and response latency time. When the system load exceeds , the parameter weights of non-critical constraints such as music quality decay, and system responsiveness is prioritized.
[0065] When the emotion change amplitude exceeds the threshold , increases, strengthening the emotion mapping constraint, that is, ensuring that the music responds to emotion changes in a timely manner.
[0066] Through the emotion-music bridging adapter, the present invention realizes the mapping from the user's facial expression to music generation, enabling the generated music to accurately reflect the user's emotional state and providing a solid technical foundation for personalized music therapy.
[0067] The emotion-music bridging adapter is another key protected point of the present invention. This module first extracts and encodes facial emotion features through the ViT model to obtain core information such as expression encoding, emotion dimensions, emotion intensity, and temporal features. On this basis, a multi-layer feature fusion mechanism is designed. Through feature alignment, attention calculation, and fusion processes, the progressive injection of emotion features is achieved, and through deep integration with the music generation module, an accurate mapping from emotion features to music parameters is ensured.
[0068] 3. Music Generation Module The music generation module is the core execution unit of the present invention, responsible for converting the previously obtained user features and emotional states into personalized therapeutic music. As Figure 2 shown in the music generation module of
[0069] 3.1 Feature Fusion Encoder The feature fusion encoder is the first stage of music generation, responsible for integrating two-way feature inputs and constructing the basic representation for music generation.
[0070] (1) Feature Input The encoder receives two-way feature inputs: The user historical feature information from the dynamic perception field, including the fused representation of user portraits, personal hobbies, historical experiences, and physiological data; The real-time facial emotion features from the emotion-music bridging adapter, which are progressively injected layer by layer during the encoding process through the adapter.
[0071] (2) Encoder Structure As Figure 2 shown on the left side of the music generation module of The encoder is composed of N identical computational blocks stacked on top of each other. Each computational block contains: Add&Norm layer (residual connection and layer normalization operation): to achieve feature normalization and residual connection; Feed-Forward layer (feed-forward neural network layer): to improve the feature representation ability through non-linear transformation;
[0072] (3) Feature Fusion By injecting facial emotion features between the layers of the feature fusion encoder, the progressive fusion of emotion features is achieved. This process is mainly completed through the attention mechanism, specifically including the calculation of Q, K, V and the fusion of attention weights.
[0073] The attention calculation mechanism obtains the query vector (Q), key vector (K), and value vector (V) through three independent linear transformations. The query vector is generated from the hidden state of the current encoding layer while the key vector and value vector are derived from the aligned sentiment features .
[0074] ; ; ; wherein, is the hidden state of the L-th layer, are the weight matrices of the query, key, and value respectively, is the facial sentiment feature representation obtained by processing through a linear layer, a normalization layer, and an activation function: ; In the formula, and are the weight matrix and bias vector of the alignment transformation respectively, is the normalization layer, is the activation function.
[0075] The standard scaled dot-product attention mechanism is adopted to achieve selective fusion of features by calculating the attention weight matrix.
[0076] ; In the formula, is the hidden state of the l +1 encoding layer; represents layer normalization processing; represents a feed-forward neural network for non-linear mapping and dimension transformation of features; is the activation function; d is the scaling factor of the scaled dot-product attention mechanism.
[0077] 3.2. Music Structure Planning Based on the feature representation Z output by the encoder, the present invention first conducts music structure planning to design the overall framework of the music. This process is similar to the composer's conceptualization stage before creation, determining the overall skeleton of the music: ; where is the structure generation function, is the style constraint condition, is the music theory rule set.
[0078] The structure planning adopts multi-dimensional design to form a complete music organizational structure: ; Among them: : The formal structure of the music, including paragraph designs such as theme presentation, development, and recapitulation; : The framework of harmonic progression, including the tonic center, chord progression, and mode change; : The rhythmic organizational structure, including beat type, rhythmic density, and prosody pattern.
[0079] 3.3. Hierarchical decoding generation The music decoding adopts a three - level cascaded decoding architecture, gradually refining the music content from macro to micro: ; ; ; (1) Structure - layer decoding( ) Input: Encoded feature Z and structure framework ; Function: Determine the basic skeleton of the music, including paragraph length, theme position, and tonality planning; Output: Music structure representation, similar to a "sketch" of the music.
[0080] (2) Content - layer decoding( ) Input: Output of the structure layer ; Function: Generate specific music materials, including melody lines, harmonic progressions, and rhythmic patterns; Emotional mapping: Different emotional states are mapped to specific music parameters: Pleasure → Pitch distribution, harmony type (e.g., high pleasure corresponds to ascending melody, bright chords); Arousal → Rhythmic density, note duration (e.g., high arousal corresponds to compact rhythm, short notes).
[0081] (3) Detail - layer decoding( ) Input: Output of the content layer ; Function: Improve the expressive details of the music, including dynamics changes, performance techniques, and timbre adjustment.
[0082] To balance predictability and creativity, the present invention introduces a temperature control mechanism. The temperature parameter T dynamically adjusts the randomness of the decoder. When stable emotions are needed, the temperature is reduced to enhance predictability. When emotional exploration is needed, the temperature is increased to increase variability. The temperature parameter T at time t is expressed as : ; Among them, is the base temperature parameter, is the creative adjustment.
[0083] 3.4 Music generation Input the output features of the detail layer into the pre-trained EnCodec decoder to generate high-fidelity audio waveforms W’ , and its process is expressed as: ; In the formula, Quantize is the vector quantization module that maps continuous features to discrete audio tokens; Φ is the decoder parameter that restores the audio waveform through deconvolution and upsampling; is the pre-trained EnCodec decoder responsible for restoring the quantized audio tokens to high-fidelity waveforms W’ .
[0084] 3.5. Optimization and constraint mechanism To ensure the therapeutic effect and artistic quality of the generated music, the present invention sets multiple constraint conditions: (1) Music coherence constraint ; Among them, is the music coherence metric at time t, is the minimum coherence threshold, set to 0.7.
[0085] (2) Emotional expression constraint This constraint ensures the gradualness of emotional expression, avoids sudden emotional changes during the treatment process, and conforms to the "psychological safety zone" theory.
[0086] ; In the formula, is the L2 norm; is t the expression change vector at time ; the maximum allowable change amplitude is set to 0.3.
[0087] 4. Dynamic feedback mechanism The dynamic feedback mechanism is the core of the closed-loop control of the present invention, enabling it to respond to changes in user emotions in real time and adjust the music generation strategy. As shown by the dynamic feedback mechanism module in Figure 2 , this mechanism connects the emotional feature extraction and music generation modules, forms a complete "perception-generation-feedback" loop, and realizes precise emotional regulation and personalized music therapy experience.
[0088] 4.1. Emotional Feature Monitoring The present invention adopts a unified facial emotional feature extraction architecture to ensure the consistency of feature representation. Although the feedback mechanism and the bridging adapter use the same ViT encoder to extract basic features, there are significant differences between the two in terms of feature processing and application. The bridging adapter focuses on feature mapping transformation, converting facial emotional features into a representation space suitable for music generation, and pays attention to the transformation and adaptation of the feature space. Different from the emotion-music bridging adapter, the dynamic feedback mechanism pays more attention to the temporal changes and comparative analysis of facial expressions rather than single-frame features.
[0089] (1) Expression change vector: Measures the amplitude of expression change between consecutive frames and reflects the degree of emotional fluctuation ; In the formula, and are the consecutive frame features extracted by the encoder of the ViT model, namely at time t and time t-1.
[0090] The system uses the consecutive frame features extracted by the same ViT encoder as in Section 2.1 of this embodiment and to calculate the expression change amount. Designing this index hopes to capture subtle facial expression changes, such as raised eyebrows, upturned corners of the mouth, or widened eyes, etc., and provide instant feedback on emotional changes.
[0091] (2) Emotional state vector: Construct a two-dimensional emotional space to accurately locate the emotional state at the current time t : ; In the present invention, from the 768-dimensional feature vector output by the ViT encoder, the valence and arousal values are extracted through a special mapping layer to construct a two-dimensional emotional space representation that is updated in real time. Among them, valence(t) is the valence value at time t, ranging from [-1,1], with positive values indicating positive emotions and negative values indicating negative emotions; arousal(t) is the arousal value at time t, ranging from [-1,1], with positive values indicating high-energy states and negative values indicating low-energy states.
[0092] (3) Attention focus: Evaluate the user's degree of focus on music: ; The present invention reuses the facial features extracted by ViT and, through an additional Attention component, specifically analyzes eye movements, head postures, and micro-expression changes to evaluate the user's degree of focus on music. Among them, is the attention score at time t, ranging from [0, 1]. The attention level is evaluated by analyzing eye movement patterns, head pose stability, and facial micro-expression changes. A high attention score indicates the user's active engagement with the current music, while a low score may imply distraction or waning interest. is t the set of facial key features at time
[0093] 4.2. Music Feature Analysis Synchronized with emotion monitoring is the feature analysis of the generated music, which enables understanding the characteristics of the self-output music and evaluating its effects to ensure that the music output matches the emotional needs: ; Among them, is the music feature vector at time t, is the audio signal generated at time t. It represents the processing of the feature analyzer.
[0094] The music analyzer extracts and analyzes the following key music parameters in real time: (1) Rhythm dynamic features Rhythm density: The number of notes per minute, reflecting the activity of the music; Accent distribution: The position and intensity of strong beats, shaping the rhythm sense of the music; Tempo change: The stability or gradual change characteristic of the tempo, affecting the emotional stability.
[0095] (2) Harmony complexity index Harmonic texture: Chord density and the number of voices; Tonal stability: The degree of certainty or ambiguity of the tonal center; Dissonance: The usage frequency and intensity of dissonant intervals.
[0096] (3) Timbre expressiveness features Timbre brightness: The proportion of high-frequency energy, affecting the emotional color; Pitch range: The pitch range and the main activity area; Dynamic contour: The volume change curve, shaping the music expressiveness.
[0097] Through this multi-dimensional music analysis, the characteristics of the currently generated music are precisely understood and compared with the user's emotional reactions to evaluate the therapeutic effect of the music. When it is found that specific music features trigger positive emotional reactions, these features will be strengthened; when certain music elements cause negative reactions, these elements will be weakened or adjusted.
[0098] 4.3. Adaptive Control Mechanism Adaptive control is the core link of dynamic feedback. By designing a two-way gating mechanism, precise adjustment of emotional characteristics and music parameters can be achieved.
[0099] 4.3.1 Bidirectional Gating Mechanism The present invention designs two control units, an input gate and an output gate, such as Figure 2 The gating module is shown in Figure 1.
[0100] (1) Input gate: Controls the influence of emotional features on music generation ; (2) Output gate: controls the adjustment intensity of music parameters ; In the formula, and They are t The feature input control gate and parameter output modulation gate at each moment, the former controls the flow of feature information into the control system, and the latter controls the degree of influence of the control signal on the music parameters; is the input weight matrix, is the output weight matrix, is the input bias term; is the output bias term; The sigmoid activation function is used to normalize the gate value to the [0,1] interval.
[0101] The practical application of the gate value is reflected in the regulation of the feature flow: ; ; in, is the original extracted emotional feature vector, is the sentiment feature vector after Gate adjustment, is the original music feature vector generated, is the music feature vector after Gate adjustment.
[0102] ⊙ represents element-wise multiplication, which enables dimension-wise adjustment of features. This design can dynamically control the flow of information. For example: When users show a positive reaction: As the value increases, the influence of emotional features increases; Value adjustment to optimize current music parameters; When the user shows unstable emotions: The value decreases, reducing the impact of emotional fluctuations; As the value increases, the stability of the music characteristics is strengthened; When user attention drops: Adjust the value and introduce novel music elements to attract attention again.
[0103] 4.3.2, PID Controller To achieve more precise regulation of emotional states, the present invention introduces a PID (Proportional-Integral-Derivative) controller, which is a classical feedback control mechanism.
[0104] ; In the formula, is the output control signal of the PID controller; is the error between the target emotional state and the current state; are the proportional term, integral term, and derivative term respectively; where, when exceeds the preset threshold, the control intensity is increased to accelerate emotional adjustment, and vice versa, the control intensity is decreased to maintain a stable state.
[0105] Corresponding relationship between the functions of each part of the PID controller and the treatment scenario: (1) Proportional term ( ): Provide immediate adjustment according to the gap between the current emotional state and the target state; For example: When detecting an anxious mood, immediately reduce the music speed and complexity.
[0106] (2) Integral term ( ): Solve long-term emotional deviations and handle persistent emotional states; For example: For a long-term low mood, gradually introduce bright timbres and ascending melodies.
[0107] (3) Derivative term ( ): Predict the trend of emotional changes, make adjustments in advance, and prevent emotional overshoot; For example: When detecting a rapid emotional change, smooth the change of music characteristics to avoid excessive stimulation.
[0108] The present invention dynamically adjusts the PID parameters according to the magnitude of the emotional deviation: When |e(t)| > threshold: Increase the control intensity to accelerate emotional adjustment; When |e(t)| ≤ threshold: Decrease the control intensity to maintain a stable state.
[0109] threshold is the preset threshold. This design mimics the working method of professional music therapists: closely observing the patient's reactions, timely adjusting the intervention strategy, and balancing short-term responses and long-term effects.
[0110] 4.4, Evaluation of Treatment Effects The present invention comprehensively evaluates the music therapy effect through multi-dimensional indicators to form a quantified treatment effect function : ; In the formula, represents t the music therapy effect at time , and are the weight coefficients of expression change, emotional state and attention level respectively, , and the weights are adjusted according to different treatment stages and goals; among them, is increased during the emotional stability period, is increased during the emotional expression period, is increased during the focused training period.
[0111] Treatment effect evaluation drives the parameter update process: ; wherein, is the music generation parameter, is the learning rate, is the treatment effect gradient. Through this parameter update mechanism, the present invention continuously optimizes the music generation strategy to gradually make the treatment effect approach the ideal goal.
[0112] To quantitatively evaluate the treatment progress, the present invention designs the following specific indicators: (1) Emotional Improvement Index (EMI) ; where V is the pleasure value, a positive value indicates an improvement in mood, and a negative value indicates a deterioration in mood. is the final pleasure value, is the initial pleasure value.
[0113] (2) Emotional Stability Score (ESS) ; The value range of is [0,1], and the higher the value, the more stable the mood. is the standard deviation function,
[0114] (3) Treatment Engagement Index (TEI) ; Comprehensively consider the average attention level and the degree of attention fluctuation. is the average value function.
[0115] 4.5. Optimization Constraints To ensure the effectiveness and safety of the treatment, the present invention sets multiple optimization constraint conditions: (1) Minimize treatment deviation: Ensure that the actual treatment effect is close to the target level ; wherein, is the target treatment effect.
[0116] (2) Music coherence constraint: Ensure the smoothness of music generation ; Among them, is the music coherence metric at time t, is the minimum coherence threshold.
[0117] (3) Emotional smoothness: Avoid drastic fluctuations in emotional states and ensure gradual adjustment ; Among them, is the change in emotional state, is the maximum allowable change amplitude.
[0118] (4) Response latency: Ensure timely response ; Among them, is the response time, is the maximum allowable response delay.
[0119] Through the above dynamic feedback mechanism, the present invention can achieve precise monitoring, timely adjustment, and smooth transition of emotional states, providing users with a personalized music therapy experience. Each module works in coordination to ensure the naturalness and coherence of music generation while guaranteeing the treatment effect.
[0120] The dynamic feedback control mechanism constructs a complete feedback control system, including real-time monitoring schemes for emotional and music features, which can accurately capture key indicators such as facial expression changes, emotional states, and attention levels. By designing an adaptive control mechanism based on bidirectional Gates and combining the precise adjustment ability of a PID controller, real-time optimization of treatment parameters is achieved. At the same time, under multiple constraint conditions, including requirements such as music coherence, emotional smoothness, and response latency, the stability and reliability of the feedback control are ensured.
[0121] Embodiment 2 This embodiment provides a computer terminal, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor.
[0122] The computer terminal can be a smart phone, a tablet computer, a laptop computer, etc. that can execute programs. The processor can be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data. When the processor executes the program, it implements the steps of the emotional music adjustment method based on dynamic perception and multimodal fusion in Embodiment 1.
[0123] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. An emotional music adjustment method based on dynamic perception and multimodal fusion, characterized in that: The following steps are involved: S1. Obtain multi-dimensional feature data of users, including user portraits, personal hobbies, historical experiences and physiological data; S2. Dynamically assigning weights to the multi-dimensional feature data through a dynamic perception field to generate fusion features; S3. Collect user facial image data in real time, extract facial emotion features through the ViT model, and map the facial emotion feature space to the music feature space through the emotion-music bridge adapter; the facial emotion features include expression coding, emotion dimension, emotion intensity and timing features; S4. Input the fusion features and the facial emotion features mapped to the music feature space into the music generation module, and generate personalized therapeutic music through the feature fusion coding, music structure planning and hierarchical decoding architecture in the module; S5. Monitor the user's facial emotional features and generate music features in real time, adjust the music generation parameters through a dynamic feedback mechanism, and form a closed-loop control.
2. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 1 is characterized in that: In step S1, four LlaMA encoders are used to process the text descriptions of four dimensions, namely, user portrait, personal hobbies, historical experiences and physiological data, respectively, and the text descriptions are converted into standardized encoded feature vectors, thereby forming the multi-dimensional feature data.
3. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 2 is characterized in that: Step S2 includes the following specific steps: S21. Concatenate the four types of coded feature vectors, namely user portrait, personal hobbies, historical experience and physiological data, to construct a unified feature representation space and achieve preliminary fusion of features. The expression formula is: ; In the formula, E is the fused feature vector, with a dimension of n×4, where n is the dimension of a single-class feature vector; , , and They are the encoded feature vectors of user portrait, personal hobbies, historical experiences and physiological data respectively; S22. Construct a dynamic perception field to achieve dynamic perception of feature importance; Among them, the weighted module in the dynamic perception field is used to initialize the weights of the four types of features, evenly distribute the feature weights in each dimension, and the weight set W Defined as: ; In the formula, The weight coefficients corresponding to the four types of features are: , , ; The expression formula of the perception field intensity function of the dynamic perception field is as follows: ; In the formula, is the perceived field strength function; t For time; and are direct influence terms and interactive influence terms respectively, and the expression formula is: ; ; In the formula, is the time decay factor; ; For feature correlation, cosine similarity is used to quantify the encoded feature vector and The degree of correlation between them is in the range of [0,1]. The larger the value, the stronger the feature correlation. is the dynamic coupling coefficient, which uses gradient information to characterize the direction of the feature's influence on the perception field, uses gradient product to reflect the degree of coordination of feature changes, and uses the sigmoid function to normalize the value range of the dynamic coupling coefficient to [0,1]; S23. The improved momentum gradient descent method is used to update the feature weights, and the updated feature weights are used to adaptively fuse the multi-dimensional features to generate fused features; wherein the expression formula for weight update is: ; In the formula, and are the feature weights before and after updating respectively; is the learning rate; is the momentum factor; is the second-order gradient, is a first-order gradient.
4. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 1 is characterized in that: In step S3, the encoder of the ViT model receives the facial image sequence as input and outputs the encoded facial emotion feature vector. The expression formula of the feature extraction process is: ; In the formula, is the facial emotion feature vector; Represents the encoding process of the ViT encoder; is a sequence of facial images; is the first feature components, each of which represents a specific aspect in the emotional feature space, and together they constitute a complete representation of the four dimensions of expression coding, emotional dimension, emotional intensity, and temporal characteristics. , m is the total number of characteristic components.
5. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 4 is characterized in that: In step S3, the emotion-music bridge adapter adopts a dual-channel design, including two processing channels of uplink adaptation and downlink adaptation; Among them, the upstream adaptation channel is used to convert the facial emotion features into the music feature space through the projection layer, ReLU activation, and layer normalization. The expression formula is: ; In the formula, Characteristic representation of uplink adaptation channel output; Representation layer normalization processing; ReLU activation processing; is the uplink adaptation channel projection weight matrix, , is the music feature space dimension, It is the dimension of facial emotion characteristics; Adapt channel bias parameters for uplink; is the set of real numbers; The downstream adaptation channel is used to compress the music feature dimensions through layer normalization, projection layer, and gating unit, and extract the core features related to emotions. The expression formula is: ; In the formula, Characteristic representation of the output of the downlink adaptation channel; It is processed by gated units to selectively retain key features; is the representation of the music feature space; is the downlink adaptation channel projection weight matrix, ; Adapt channel bias parameters for downlink; The uplink adaptation channel and the downlink adaptation channel work together and are integrated through residual connections. The expression formula is: ; In the formula, is a bridging feature; It is a residual connection process used to preserve feature information and optimize gradient propagation.
6. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 5 is characterized in that: In step S3, the emotion-music bridge adapter adopts the following optimization and constraint mechanism: (1) Computational efficiency goal: minimize the delay introduced by the emotion-music bridge adapter, that is: ; In the formula, the subscript A set of trainable parameters for the emotion-music bridge adapter, including the upstream channel parameters , downlink channel parameters and internal parameters of the gating unit; represents the expected value operator, which is used to calculate the average performance of a random variable; Delay time introduced for the emotion-music bridge adapter; (2) Emotional mapping accuracy goal: maximize the consistency between emotional expression and music features, that is: ; In the formula, Represents facial emotion feature vector Music parameter characteristics Semantic consistency measure of ; (3) Music quality objective: maximize the music quality under the delay constraint, that is: ; In the formula, To generate music quality metrics; is the upper threshold of the delay time; Among them, music quality measurement It is defined as a multi-factor comprehensive score, and the calculation formula is: ; In the formula, is the weight coefficient; Scoring the musical structural integrity, a measure of the logic of musical form; Melodic fluency was scored to assess melodic naturalness and memorability; Scoring harmony to evaluate the balance and richness of the harmonic progression; score emotional expressiveness, measuring how well the music matches the target emotion; The emotion-music bridge adapter adopts the following multi-objective balancing strategy: (1) Hierarchical priority strategy, dynamically adjust the optimization target weight according to user scenario requirements: ; In the formula, To optimize the objective function overall; Indicates A single optimization target goal; A set of parameters for the emotion-music bridge adapter; is the context-dependent weight coefficient, ; The current interaction scene type; When in a real-time treatment scenario, > > , giving priority to ensuring computational efficiency; When quality is the priority, > > , giving priority to music quality; When in the emotional response field, > > , giving priority to ensuring the accuracy of emotion mapping; (2) Constraint relaxation strategy allows some constraints to be softened in extreme scenarios. The expression formula is: ; In the formula, is the soft constraint penalty term; is the constraint relaxation coefficient; For the Constraints, , corresponding to the three constraints of computational efficiency, emotion mapping accuracy, and music quality; When the system load exceeds the preset load threshold hour, Decrease, relax non-critical constraints; When the emotional change exceeds the threshold hour, Incrementally, strengthen the emotion mapping constraints.
7. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 6 is characterized in that: In step S4, the method for generating personalized therapeutic music includes the following specific steps: S41. Using a feature fusion encoder, integrating the fusion feature and the facial emotion feature, and constructing a basic feature representation for music generation; The feature fusion encoder is composed of N identical computing blocks stacked together, each of which includes: residual connection and layer normalization operations, feedforward neural network layers and multi-head attention layers. The residual connection and layer normalization operations are used to achieve feature normalization and residual connection; the feedforward neural network layer is used to improve feature representation capabilities through nonlinear transformation; the multi-head attention layer is used to capture long-distance dependencies between features; The facial emotion features are injected between the encoding layers of the feature fusion encoder through the attention mechanism to achieve the progressive fusion of emotion features. Among them, the attention mechanism obtains the query vector Q, key vector K and value vector V through three independent linear transformations. The calculation formula is: ; ; ; In the formula, For the l The hidden state of the encoding layer, are the weight matrices for query, key, and value, respectively. for The facial emotion feature representation obtained after processing by the linear layer, normalization layer and activation function: ; In the formula, and are the weight matrix and bias vector of the alignment transformation, is the normalization layer, is the activation function; S42. Based on the basic feature representation output by the feature fusion encoder, music structure planning is performed to form a structural framework of the music; The structural framework of the music The expression formula is: ; In the formula, Generate functions for structures; It is the basic feature representation output by the feature fusion encoder; is the style constraint; for music theory rule sets; The structural planning adopts a multi-dimensional design, which is expressed as: ; In the formula, For the musical form and structure; to frame the harmonic progression; Organize the structure for the rhythm; S43. A three-layer cascade decoding architecture is used to represent the basic feature Z and the structural framework To decode, the expression formula is: ; ; ; In the formula, , and They represent the structure layer decoding process, the content layer decoding process and the detail layer decoding process respectively. , and They are the corresponding outputs of each layer respectively; S44. Output features of detail layer Input the pre-trained EnCodec decoder to generate high-fidelity audio waveform W’ , the process is expressed as: ; In the formula, Quantize is a vector quantization module that converts continuous features is mapped to discrete audio tags; Φ is the decoder parameter, and the audio waveform is restored through deconvolution and upsampling; A pre-trained EnCodec decoder responsible for restoring the quantized audio markers to high-fidelity waveforms W’ ; The music generation process in step S4 adopts the following optimization and constraint mechanism: (1) Musical continuity constraints, namely: ; In the formula, is the music coherence measure at time t, is the minimum coherence threshold; (2) Emotional expression constraints, namely: ; In the formula, is the L2 norm; for t The expression changes at each moment. is the maximum allowable variation.
8. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 7 is characterized in that: Step S5 includes the following specific steps: S51. Using the same ViT model as in step S3 to monitor the user's facial emotional features in real time, thereby calculating the expression change vector, emotional state vector and attention concentration; Among them, the calculation formula of the expression change vector is: ; In the formula, and Continuous frame features extracted by the encoder of the ViT model; The calculation formula of the emotional state vector is: ; In the formula, for t The emotional state vector at the moment; for t The pleasure value of the moment, ranging from [-1,1], where positive values indicate positive emotions and negative values indicate negative emotions; for t The awakening value at the moment, ranging from [-1,1], positive values indicate high energy state, and negative values indicate low energy state; The calculation formula for attention concentration is: ; In the formula, for t The attention score at the moment, ranging from [0,1]; For the attention component, the user's concentration on the music is evaluated by analyzing eye movements, head posture and micro-expression changes; for t The set of key facial features at the moment; S52. Perform feature analysis on the generated music in real time, and the expression formula is: ; In the formula, for t The music feature vector at the moment, for t The audio signal generated at each moment; A music analyzer is used to analyze the rhythm dynamics, harmonic complexity and timbre characteristics of the generated music; S53. The intensity of the input emotional features and the output music parameters is adjusted through a bidirectional gating mechanism, and the calculation formula is: ; ; In the formula, and They are t The characteristic input control gate and parameter output modulation gate at each moment; is the input weight matrix, is the output weight matrix, is the input bias term; is the output bias term; is the sigmoid activation function, which is used to normalize the gate value to the [0,1] interval; S54. Use PID controller to dynamically adjust music parameters according to emotional deviation, and the calculation formula is: ; In the formula, is the output control signal of the PID controller; is the error between the target emotional state and the current state; are proportional term, integral term and differential term respectively; among them, when When the preset threshold is exceeded, the control intensity is increased to speed up the emotional adjustment, and vice versa, the control intensity is reduced to maintain a stable state; S55. Comprehensively evaluate the effect of music therapy through multi-dimensional indicators to form a quantitative treatment effect function: ; In the formula, express t The therapeutic effect of music at all times; , and are the weight coefficients of expression change, emotional state and attention level, , adjust the weight according to different treatment stages and goals; among them, increase , which increases during the emotional expression period , increased during the focused training period ; S56. The update of the music generation parameters is driven according to the treatment effect. The expression formula of the update process is: ; In the formula, and They are and The music generation parameters at each moment, is the learning rate, for The gradient of treatment effect at each moment.
9. The emotional music adjustment method based on dynamic perception and multimodal fusion according to claim 8 is characterized in that: The following optimization and constraint mechanisms are used in the closed-loop control of step S5: (1) Minimize treatment bias, that is: ; In the formula, To target treatment effect; (2) Musical continuity constraints, namely: ; In the formula, for t A measure of musical coherence at a moment in time; is the minimum coherence threshold; (3) Emotional smoothness, that is: ; In the formula, is the change in emotional state; is the maximum allowable variation; (4) Response delay, that is: ; In the formula, is the response time; is the maximum allowed response delay.
10. A computer terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the emotional music adjustment method based on dynamic perception and multimodal fusion as described in any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Smart home equipment music recommendation method based on continuous emotion of user
CN118245629A
Multi-mode emotion collection and acousto-optic comprehensive emotion treatment device
CN118717121A
A method and system for constructing user portraits and delivering precise advertisements in digital marketing
CN119741067A
Method for analysing comprehensive state of a subject
US20160350801A1
Methods, apparatuses and computer program products for gaze-driven adaptive content generation
US20250130636A1
Cited By
Intelligent music regulation and control system and method based on electroencephalogram signal multi-modal feature recognition
CN120437460A
Wearable device video live broadcast method and system based on intelligent AI large model driving
CN120916014A
Wearable device video live streaming method and system based on intelligent AI large model driving
CN120916014B
Virtual image model construction method and system based on image cloning
CN121349311A
Data processing method and device for virtual scene
CN121371614A