Intelligent music generation algorithm and system dynamically adapting to environmental emotion

By embedding music control parameters in the Lie group trajectory space and introducing user feedback-driven dynamic fine-tuning, the problems of trajectory discontinuity and static emotion mapping in existing music generation systems are solved, achieving personalized and adaptive melody generation effects.

CN121096293APending Publication Date: 2025-12-09PINGDINGSHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511229144.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing music generation systems suffer from discontinuous control trajectories, static and unadjustable emotion mapping, lack of user feedback loops, and weak adaptive capabilities, resulting in abrupt changes in rhythm control signals, inconsistent emotion outputs, and a simplistic trajectory structure during melody generation.

Method used

A music control parameter modeling method based on Lie group trajectory space is adopted, which combines multimodal emotional feature information, constructs a continuous music control vector sequence through emotional state function, embeds the trajectory in Lie group trajectory space, introduces a user feedback-driven dynamic fine-tuning mechanism, and forms a closed-loop adaptive generation mechanism.

Benefits of technology

It achieves stability of trajectory structure and natural state transition in the music generation process, improves the personalized adaptation of melody generation to user emotions, ensures that the generation result is highly consistent with the emotional state, and solves the structural alignment problem between control information and melody generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096293A_ABST
    Figure CN121096293A_ABST
Patent Text Reader

Abstract

The invention relates to the field of music artificial intelligence, and discloses an intelligent music generation algorithm dynamically adapting to environment emotions, which comprises the following steps: acquiring emotional state information of a user in a current environment, and constructing an emotional state function changing along with time for describing an emotional change trend of the user in continuous time; inputting the emotional state function into a pre-trained emotional mapping model to obtain a corresponding music control vector sequence, the control vector sequence being defined in a Lie algebra space of music control parameters; the invention also discloses an intelligent music generation system dynamically adapting to the environmental emotion. The system comprises an emotion recognition module; an emotion mapping module; a track embedding module; a music generation module; and a feedback control module. According to the method, the Lie group trajectory modeling and user emotion feedback fine tuning mechanism is introduced, so that the structural continuity of melody generation and dynamic self-adaption of emotion response are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of music artificial intelligence, specifically to an intelligent music generation algorithm and system that dynamically adapts to environmental emotions. Background Technology

[0002] In daily life, music has long since transcended passive background content, becoming an active interactive medium highly correlated with users' emotions and behavioral states. As people's demand for personalized music experiences increases, traditional fixed-style playback modes can no longer satisfy the dynamic preferences of different users in various emotional scenarios. Especially in scenarios such as psychological adjustment, game accompaniment, and virtual voice synthesis, there is a greater need for systems to respond in real-time to the user's emotional state, thereby outputting melodic fragments with emotional expressiveness and strong structural coherence.

[0003] Existing techniques typically employ coarse-grained control of generated melodies using static emotion labels. Some methods leverage deep neural networks to process emotion and melody information within a unified encoding space, thereby achieving functions such as style-driven processing, rhythm recognition, and melodic fluency. Other research introduces attention mechanisms to achieve overall consistency between melody rhythm and control vectors without relying on explicit labels. These techniques have shown promising results in enhancing melodic diversity and ensuring rhythmic regularity.

[0004] However, existing technologies still have some shortcomings. First, most existing control vectors are in Euclidean space, and the control trajectory lacks geometric constraints, which makes the rhythm control signal prone to abrupt changes during generation, making smooth evolution impossible. Second, the emotion mapping model remains static after training, lacking the ability to perceive real-time user feedback. Once the model misjudges the user's emotion, the subsequent melody output will continuously deviate from the target state. In addition, existing regulation mechanisms are mostly based on label-level control, with a single path and lack of trajectory structure or dynamic continuity, resulting in severe drift and lack of temporal consistency in the regulation results. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an intelligent music generation algorithm and system that dynamically adapts to environmental emotions, solving the problems of discontinuous control trajectories, static and unadjustable emotion mapping, lack of user feedback loops, weak adaptive capabilities, and insufficient fusion of multi-dimensional control information in existing technologies.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent music generation algorithm that dynamically adapts to environmental emotions, comprising the following steps:

[0007] Obtain the user's emotional state information in the current environment and construct an emotional state function that changes over time to describe the user's emotional change trend over continuous time;

[0008] The emotional state function is input into a pre-trained emotion mapping model to obtain the corresponding music control vector sequence, which is defined in the Lie algebra space of the music control parameters.

[0009] Perform an exponential mapping operation on the control vector sequence to obtain the music control parameter trajectory embedded in the Lie group trajectory space;

[0010] The trajectory of the music control parameters is passed as a conditional input to the music generation model to guide the model to generate a melody fragment that matches the emotional state.

[0011] Based on user emotional feedback information during the playback of the melody segment, the emotion mapping model is dynamically fine-tuned and updated, thereby forming a closed-loop adaptive music generation mechanism.

[0012] Preferably, the emotional state information includes:

[0013] Emotional feature information from multimodal data sources, including user facial images, audio recordings, and natural language text;

[0014] The facial images are used to extract the user's facial expressions, the audio is used to analyze tone, speech rate and intensity, and the text information is used to identify the emotional tendencies in the language expression.

[0015] The multimodal emotional features, after being fused, constitute the original input of the emotional state.

[0016] Preferably, the emotional state function includes:

[0017] A continuous function structure based on time series is used to represent the user's emotional evolution process over continuous time.

[0018] This function contains at least two components representing pleasure and arousal, which respectively reflect the positiveness and activation state of the user's subjective feelings, and has the characteristic of changing smoothly over time.

[0019] Preferably, the music control vector sequence includes:

[0020] A multidimensional set of parameters used to control the music generation model, with each control vector corresponding to the music generation state at a point in time;

[0021] The vector includes rhythm speed parameters, melody center pitch parameters, pitch fluctuation amplitude parameters, and mode type parameters, which are used to adjust the rhythm style, pitch range distribution, and emotional expression of the generated melody, respectively.

[0022] Preferably, the Lie algebra space includes:

[0023] The vector space defined in the real number field constitutes the basic structure for the evolution of the trajectory of music control parameters;

[0024] This space supports addition and scalar multiplication operations and has continuous differentiability.

[0025] Preferably, the exponential mapping operation includes:

[0026] The process of mapping a sequence of music control vectors defined in a Lie algebra space to a Lie group trajectory space;

[0027] The mapping maintains the continuity and differentiability of the control parameters over time.

[0028] Preferably, the music control parameter trajectory includes:

[0029] A set of time-varying control parameters that evolve in the Lie group trajectory space, which provide guiding signals during melody generation;

[0030] The trajectory is controllable and continuous, and is used to adjust the rhythm, mode, melodic direction and note density of the melody, so that the generated music is more in line with the target emotional intention.

[0031] Preferably, the melody fragment includes:

[0032] The music generation model outputs a sequence of notes under the guidance of a given control parameter trajectory, and each melodic fragment can be used as a component of a complete musical structure;

[0033] The musical note sequence is organized and arranged according to the rhythm and emotional characteristics encoded by the control parameters, and can express the musical style and emotional color that matches the input emotional state.

[0034] Preferably, the dynamic fine-tuning update includes:

[0035] Based on the current user's emotional feedback, the parameters of some neural network weights in the emotional mapping model are fine-tuned.

[0036] The fine-tuning process employs an incremental learning strategy to maintain the stability of the original model structure;

[0037] Simultaneously, the response sensitivity of the mapping model is dynamically adjusted based on the feedback of emotional bias.

[0038] This invention also includes an intelligent music generation system that dynamically adapts to environmental emotions, comprising:

[0039] The emotion recognition module is used to collect multimodal data of the user's current environment, including facial images, audio, and text input, and to identify the user's current emotional state based on the data, thereby constructing a continuous emotion state function that changes over time.

[0040] The emotion mapping module is used to receive the emotion state function and convert it into a sequence of music control vectors defined in the Lie algebra space through a pre-trained mapping model, so as to express the high-level control intention of music generation.

[0041] The trajectory embedding module is used to perform an exponential mapping operation on the control vector sequence, embedding it into the Lie group trajectory space to obtain a music control parameter trajectory with time continuity and differentiability.

[0042] The music generation module is used to take the music control parameter trajectory as a guiding condition input, combine it with the historical note sequence, and generate a melody fragment that matches the user's current emotional state through a neural network model.

[0043] The feedback control module is used to monitor the user's real-time emotional feedback during the playback of the melody segment, and to fine-tune and update the parameters of the emotion mapping module based on the feedback.

[0044] This invention provides an intelligent music generation algorithm and system that dynamically adapts to environmental emotions. It has the following beneficial effects:

[0045] 1. This invention employs a music control parameter modeling method based on Lie group trajectory space, enabling the control vector sequence to possess continuity and differentiability in space, achieving a stable trajectory structure and natural state transitions. Compared to existing technologies that directly model control sequences based on linear Euclidean space, this overcomes the problems of obvious trajectory jumps and strong control discontinuities.

[0046] 2. This invention introduces a dynamic fine-tuning mechanism for an emotion mapping model driven by user emotion feedback, achieving an adaptive closed loop between the music generation process and the user's subjective feelings. It achieves melody style updates based on genuine emotion perception, avoiding the problem of static mapping models in traditional methods failing to continuously match the user's current feelings, and significantly improving the personalization and adaptability of interactive responses.

[0047] 3. This invention employs a technical approach that uses Lie group trajectories as conditional input to guide melody generation, successfully achieving structural alignment between control information and melody generation behavior, resulting in a high degree of consistency between the generated outcome and the emotional state. Compared to existing methods that only regulate the generation direction through labels or latent vector embedding, this invention effectively solves the problems of weak control constraints and easy deviation within the generation process.

[0048] 4. This invention employs a trajectory modeling approach that integrates multi-dimensional control information such as rhythm, structure, and emotion. The music control parameter trajectory constructed by this invention possesses global interpretability and time alignment capabilities, resulting in more coherent rhythmic levels and emotional fluctuations in the generated music. Compared to previous methods that relied on manual features or simple feature splicing, this invention overcomes the bottlenecks of fragmented control dimensions and low information fusion efficiency. Attached Figure Description

[0049] Figure 1 This is a schematic diagram of the algorithm flow of the present invention;

[0050] Figure 2 This is a schematic diagram of the system architecture of the present invention. Detailed Implementation

[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] Please see the appendix Figure 1 This invention provides an intelligent music generation algorithm that dynamically adapts to environmental emotions, comprising the following steps:

[0053] S1. Obtain the user's emotional state information in the current environment and construct an emotional state function that changes over time to describe the user's emotional change trend over continuous time.

[0054] To achieve emotion-driven music generation, it is first necessary to establish an input model closely related to the user's actual emotional state. The acquisition and modeling of emotional states, as the system's perceptual input, directly determines the relevance and adaptability of music generation through its accuracy and expressiveness. Therefore, it is essential to first effectively collect the user's current emotional characteristics from the environment and model them as a time-varying emotional state function, laying the foundation for subsequent derivation of control vectors and trajectory generation.

[0055] In general, a user's emotional state is difficult to accurately express using only a single modality. Therefore, this invention employs a multimodal fusion strategy to obtain emotional feature information. As a possible implementation, the sources of emotional data acquisition include, but are not limited to, facial images, speech audio signals, and natural language text.

[0056] In this embodiment, facial image information can be captured in real time by an integrated camera device to capture images of the user's facial area. Based on a facial expression recognition model (such as a CNN-ResNet structure), emotional feature points are extracted, and the emotional category and intensity parameters corresponding to the expression are output.

[0057] In one possible implementation, speech audio data is collected in real time by microphones in the environment and processed by an audio sentiment analysis module. This module extracts low-level features of the speech signal, such as pitch, energy, and rate, and inputs them into a trained acoustic sentiment classification model, outputting a sentiment label distribution or sentiment score.

[0058] As an alternative, natural language data can also be collected via text input, which is suitable for scenarios where users express emotions through text. After the text data is processed by sentiment analysis algorithms (such as the BERT model based on Transformer), it can output sentiment polarity scores related to emotions.

[0059] After the multimodal emotion feature collection is completed, it is processed uniformly through a feature fusion module. Generally, weighted averaging, attention mechanisms, or multilayer perceptrons are used to fuse the feature vectors, resulting in a fused vector representing the user's current emotional state. This fused vector is denoted as:

[0060] ;

[0061] in, This represents the fused emotional state vector. For the first Modal class, This represents the transpose operation of a vector.

[0062] In a preferred embodiment, the fused emotion state vector is further mapped to a two-dimensional continuous function, specifically constructed as follows:

[0063] ;

[0064] in, Indicates time Emotional pleasure Indicates time wake-up rate Indicates time A two-dimensional continuous vector function of emotion. This represents the transpose operation of a vector.

[0065] Specifically, in order to capture the dynamic trend of emotion changes, the serialized input emotion vector can be smoothed by using a sliding window average or a weighted filter, thereby suppressing short-term noise and improving the stability of modeling.

[0066] In one implementation, the emotion state function can be constructed based on state estimation methods, such as Kalman filtering or exponential moving average algorithms, so that it has temporal continuity while maintaining response sensitivity.

[0067] The smoothing function can be estimated in the following form:

[0068] ;

[0069] ;

[0070] in, These are the normalized Gaussian weight coefficients. Parameters for controlling the smoothness Current time The output of the emotional pleasure smoothing function, For at a certain point in time The original emotional pleasure level observation value at the location, The natural exponential function is used to construct the Gaussian weight decay kernel. These are the normalized Gaussian weight coefficients. Indicates time Emotional pleasure Indicates time wake-up rate This represents the total number of historical observation points within the sliding window. For the time step index within the sliding window, For at a certain point in time The original emotional arousal observation value, This refers to the current time.

[0071] In some embodiments, the emotion state function can be directly called as input in the subsequent neural network model as a time function object, or it can be discretized into a vector sequence of fixed time intervals and input into the mapping network.

[0072] To achieve real-time music adaptation, the system needs to discretely sample the above function at a fixed sampling period. For example, sampling once per second will construct 10 emotional state points within 10 seconds for subsequent mapping.

[0073] As a supplementary form, the emotional state function can also be associated with contextual events, such as the dynamic migration of the position of specific semantic tags (anger, sadness) in the emotional spectrum, and can also be used to adjust the musical structure or rhythmic style.

[0074] S2. Input the emotion state function into the pre-trained emotion mapping model to obtain the corresponding music control vector sequence, which is defined in the Lie algebra space of the music control parameters.

[0075] After constructing the emotion state function, the system needs to use this function as input and further map it to the music control parameter space to drive the melody generation module for structured control. This step is a key part of the technical solution of this invention, connecting the semantic bridge between emotion perception input and music behavior output.

[0076] Generally, there is no direct linear relationship between emotional information and musical parameters, requiring nonlinear modeling methods for mapping and fitting. Therefore, this invention constructs and pre-trains a neural network-based mapping model to achieve functional approximation from the emotional state function to the control vector.

[0077] In this embodiment, the emotion mapping model is a deep neural network. Its input is the current emotion state vector, and its output is a music control vector defined in the Lie algebra space. This network model has end-to-end mapping capabilities and can capture complex relationships between nonlinear, multi-scale, and multi-modal components.

[0078] Specifically, the model employs a feedforward neural network with at least three layers, including an input layer, hidden layers, and an output layer. The input layer accepts a 2-dimensional emotion state vector; the hidden layers use the ReLU activation function to enhance feature representation; and the output layer uses linear activation to output a continuous control parameter vector.

[0079] In some embodiments, to enhance the generalization ability of the model, normalization layers, residual connections, or attention mechanism modules can be introduced into the network to deal with complex emotional expression structures.

[0080] The output is a sequence of vectors defined in the Lie algebra space, denoted as:

[0081] ;

[0082] in, This represents the transpose operation of a vector. Indicates time The control vector at any given time contains four dimensions of music control parameters, the meanings of which are as follows:

[0083] A factor used to represent the tempo control of music;

[0084] This indicates the amount of adjustment in the center pitch of the melody;

[0085] It reflects the fluctuation range of pitch and is used to control the dynamic rise and fall of melody;

[0086] A continuous encoding representation of musical mode types.

[0087] In one possible implementation, the control vector is generated every 200 milliseconds, forming a time-updated vector sequence to support real-time control requirements.

[0088] As an alternative, the model training phase uses the mapping relationship between the emotion state function and the control parameters of manually labeled music segments as training data, and optimizes it by minimizing the loss function.

[0089] In this embodiment, the loss function is the sum of two parts: one part measures the mean square error between the predicted control vector and the actual control parameters, and the other part penalizes drastic fluctuations in the control vector over time. Specific definitions and meanings are explained above.

[0090] In addition, to ensure that the output vector has the structural properties of a Lie algebra, a structural constraint term is introduced during the training process so that the output results satisfy the requirements of linearity, closure and additivity.

[0091] In some specific scenarios, the system also supports targeted fine-tuning of the model structure. For example, during long-term user use, the model parameters can be updated through a feedback mechanism to better match individual emotional styles and preferences.

[0092] It is important to note that the music control vector sequence forms the input basis for the subsequent generation of Lie group trajectories. Therefore, its mathematical structure must possess the basic properties of Lie algebras, including: linear space structure, mapping relationship with Lie group exponents, and first-order differentiability.

[0093] S3. Perform an exponential mapping operation on the control vector sequence to obtain the music control parameter trajectory embedded in the Lie group trajectory space;

[0094] In the semantically driven music control parameter generation method proposed in this invention, the aforementioned content-aware module, semantic guidance module, and rhythm structure modeling unit work together to extract and construct a control vector sequence expressing temporal dependence and local semantic continuity. This control vector sequence not only encodes information such as rhythmic period, structural hierarchy, melodic dynamics, and dynamic control, but also enhances the semantic consistency between temporal sequences through a structural attention mechanism.

[0095] Generally, the aforementioned control vector sequence resides in a Euclidean vector space. Modeling the control trajectory in this space can easily introduce discontinuous state transitions, which is detrimental to modeling the rhythmic smoothness, structural transitions, and parameter differentiability requirements in music generation. Therefore, during the trajectory generation stage, the control vector sequence needs to be embedded in a trajectory space with a Lie group structure. By using exponential mapping, the trajectory can be continuously evolved in terms of spatial structure, thus maintaining the interpretability of trajectory generation at both the mathematical and physical levels.

[0096] As an option, the system will control vector Mapping to the Lie algebra space yields the intermediate representation vector. The mapping method is a linear transformation, specifically defined as:

[0097] ;

[0098] in, This represents the mapping matrix used for dimension reconstruction and orientation recalibration. This is the offset parameter.

[0099] Specifically, in one possible implementation, Lie groups are used to construct composite structures of two-dimensional translation and rotation, suitable for modeling control sequences with rhythmic directionality and paragraph dynamic behavior. In this structure, This includes a composite control state that incorporates rotational and translational components.

[0100] right Applying an exponential mapping operation yields the control state points in the Lie group trajectory space, which constitute a trajectory sequence. This sequence is the trajectory of the music control parameters embedded in the Lie group trajectory space, possessing continuity, integrability, and consistency of group structure.

[0101] In some embodiments, to improve the adaptability of trajectories in the generator, the system constructs a trajectory interpolation module to perform Bézier spline interpolation or inter-sample exponential weighted smoothing on the trajectory sequences. This module performs difference calculations based on the right invariants of the Lie group between trajectories, ensuring that the trajectory fitting behavior conforms to the intra-group computational rules.

[0102] As a technological extension, the trajectory control submodule can also introduce a residual term based on velocity constraints to regularize the rate of change of the trajectory state. For example, the system can calculate:

[0103] ;

[0104] in, Indicates time The residual term of the rate of change of trajectory state at time t. Denotes the logarithmic mapping from Lie groups to Lie algebras. This represents the control state in the trajectory space of the Lie group.

[0105] A smoothing loss is introduced to reduce discontinuous jumps. The above logarithmic mapping operation is a left-invariant difference mapping that backtracks from a Lie group to a Lie algebra. This can be understood as a time step. The rate of state evolution.

[0106] Furthermore, the system constructs a trajectory error feedback mechanism during the training phase, comparing the differences between the predicted trajectory and the target trajectory in the Lie group space, thereby updating the gradient of the transformation parameters for the control vector-to-trajectory mapping. This feedback process helps enhance the target alignment capability of trajectory generation and the accuracy of control response.

[0107] S4. Pass the music control parameter trajectory as a conditional input to the music generation model to guide the model to generate melody fragments that match the emotional state.

[0108] In the provided semantic-driven music generation method, the aforementioned steps have extracted the multidimensional emotional state of the current target segment through the emotion recognition module, and based on the emotional state and contextual semantic information, combined with the content-aware feature generation module, rhythm control module and trajectory modeling mechanism, a music control parameter trajectory embedded in the Lie group trajectory space is constructed.

[0109] Generally, the control parameter trajectory encodes composite control signals such as rhythmic change trends, structural evolution direction, and emotional dynamic intensity in the form of a time series. This trajectory is continuous and differentiable in the Lie group space and possesses a stable dynamic evolution law. To achieve an effective match between the generated final melody content and the emotional state, the aforementioned music control parameter trajectory needs to be passed as a conditional input to the music generation model to complete the melody output control based on emotion perception.

[0110] Alternatively, the system uses this trajectory sequence as conditional input features and feeds it into the melody generation model. This model can employ a decoder framework based on the Transformer structure, or a sequence generator combined with a gated recurrent network. The trajectory information is input via structured embedding, specifically represented as follows:

[0111] ;

[0112] in, For the corresponding time step Trajectory embedding representation This is a trajectory embedding function used to map Lie group elements to the continuous vector space required by the generator. This represents the control state in the trajectory space of the Lie group.

[0113] In one possible implementation, the trajectory embedding module performs [the following] by […]. Applying a Lie group logarithmic mapping and then using a linear transformation layer to flatten the features, which then serves as the fusion condition for the position alignment control vector and the input of the melody decoder.

[0114] Specifically, the melody generation model relies on the historical note sequence at each time step. Trajectory input vector With global emotional state encoding Together, they constitute the decoding input conditions:

[0115] ;

[0116] in, This represents a high-dimensional emotion state vector obtained through the emotion extraction module, defined in a fixed emotion category or continuous emotion axis space. For the corresponding time step Trajectory embedding representation For the current time step Output the note or melody symbol to be generated. In time step Previous historical note sequences.

[0117] In some embodiments, to achieve temporal consistency between trajectory input and melody generation, the system designs an alignment control mechanism. This mechanism utilizes beat position encoding to resample the trajectory sequence, ensuring it aligns with the melody generation cycle at the beat granularity. This process uses the rhythm segment division boundaries provided by the rhythm structure module for beat alignment.

[0118] As a technological extension, trajectory conditions can also be fused with other input modalities (such as lyric emotion tags and melody start seeds) through multiple channels. Fusion strategies include channel attention mechanisms, gating weighting mechanisms, or bidirectional cross-attention modules to enhance the responsiveness of the generated melody to emotional commands.

[0119] Furthermore, during training, the system introduces a trajectory-guided generation loss function, combining melody generation loss and control trajectory similarity loss, to ensure that the trajectory input substantially constrains the generation behavior:

[0120] ;

[0121] in, For standard melody cross-entropy loss, This represents the consistency constraint term between the Lie group trajectory input and the target generated rhythm sequence. For loss weight hyperparameters, This represents the total loss function of the melody generation system.

[0122] S5. Based on user emotional feedback information during the playback of melody segments, the emotion mapping model is dynamically fine-tuned and updated to form a closed-loop adaptive music generation mechanism.

[0123] In the semantically driven music generation system, the aforementioned steps have extracted the target emotional state through an emotion recognition module and, combined with a content-aware control mechanism, rhythm modeling strategy, and Lie group trajectory structure, generated a melody fragment that satisfies the target emotional intent. Simultaneously, the melody generation process incorporates music control parameter trajectories as conditional control information, ensuring that the generated result maintains consistency with the expected emotional state in terms of rhythmic form, dynamic curve, and pitch trend.

[0124] Generally, modeling emotional states is a static process, meaning the emotion mapping model has already converged its parameters during the training phase and remains unchanged during the inference phase. However, in real-world scenarios, users' emotional perception of melodic segments exhibits significant individual differences and dynamic fluctuations, making it difficult for static emotion mapping strategies to accurately capture the current subjective emotional experience of users. Therefore, it is necessary to introduce an information update mechanism based on user emotional feedback during playback to dynamically fine-tune the emotion mapping model, thereby achieving closed-loop, adaptive music generation control.

[0125] In this embodiment, the system collects the user's emotional feedback information in real time during the playback of the melody segment. This feedback information can be obtained through multimodal signals, such as, but not limited to, facial expression recognition, changes in physiological parameters (such as heart rate and skin conductance), changes in voice tone, or active evaluation signals.

[0126] Specifically, the system compares the emotional outputs predicted by the original emotion mapping model. Emotional tags collected from user feedback Construct the emotion mapping error signal:

[0127] ;

[0128] in, For emotional output, For emotion tags, Indicates time The emotional error term.

[0129] The aforementioned error terms serve as optimization driving signals and are fed back to the emotion mapping module for parameter fine-tuning.

[0130] As an alternative, the system employs an online fine-tuning strategy, updating only a subset of parameters in the emotion mapping model with low amplitude to avoid structural shifts in the model due to short-term feedback fluctuations.

[0131] In some embodiments, to stabilize the feedback training process and improve individual adaptability, the system introduces a historical feedback accumulation mechanism, which aggregates the emotion error sequences within multiple playback cycles into a feedback window, calculates the cumulative loss, and uses it for gradient updates.

[0132] ;

[0133] in, For the feedback window size, This is the start time of the current feedback cycle. For emotional output, For emotion tags, Represents the feedback loss function. At the current time point,

[0134] During the feedback optimization process, the system employs a weighted update strategy, allocating weights based on the confidence level or perceptual consistency of the feedback information as the gradient. For example, when the facial expression recognition result has high consistency with physiological parameters, the weight coefficient is increased, and vice versa, to avoid noise signals affecting the optimization result.

[0135] To avoid overfitting or model degradation during the feedback process, the system introduces a freezing strategy and regularization constraints, allowing only specific intermediate layer or low-frequency activation channel parameters in the sentiment mapping model to participate in the update. The specific constraints are expressed as follows:

[0136] ;

[0137] in, Indicates the currently updated parameter. Represents the initial parameters of the original model. For the strength of the regularization term, This represents the regularization loss term.

[0138] Furthermore, in this embodiment, the updated emotion mapping model participates in the next round of melody generation, forming a closed-loop structure. Based on the updated emotion representation, the system reconstructs the music control parameter trajectory and sequentially passes it to the trajectory embedding module and the melody generation decoder module, achieving melody control based on dynamic optimization through individual feedback.

[0139] The intelligent music generation system that dynamically adapts to environmental emotions described below and the intelligent music generation algorithm that dynamically adapts to environmental emotions described above can be referred to as corresponding to each other.

[0140] Please see the appendix Figure 2 The present invention also provides an intelligent music generation system that dynamically adapts to environmental emotions, comprising:

[0141] The emotion recognition module is used to collect multimodal data from the user's current environment, including facial images, audio, and text input. Based on the data, it identifies the user's current emotional state and then constructs a continuous emotion state function that changes over time.

[0142] The emotion mapping module receives the emotion state function and converts it into a sequence of music control vectors defined in the Lie algebra space through a pre-trained mapping model, in order to express the high-level control intention of music generation.

[0143] The trajectory embedding module is used to perform an exponential mapping operation on the control vector sequence, embedding it into the Lie group trajectory space to obtain a music control parameter trajectory with time continuity and differentiability.

[0144] The music generation module is used to take the trajectory of music control parameters as the guiding condition input, combine it with the historical note sequence, and generate a melody fragment that matches the user's current emotional state through a neural network model.

[0145] The feedback control module is used to monitor the user's real-time emotional feedback during the playback of melody segments, and to fine-tune and update the parameters of the emotion mapping module based on this feedback.

[0146] The system in this embodiment can be used to execute the above algorithm embodiments, and its principle and technical effect are similar, so they will not be described again here.

[0147] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A smart music generation algorithm that dynamically adapts to environmental emotions, characterized in that, Includes the following steps: Obtain the user's emotional state information in the current environment and construct an emotional state function that changes over time to describe the user's emotional change trend over continuous time; The emotional state function is input into a pre-trained emotion mapping model to obtain the corresponding music control vector sequence, which is defined in the Lie algebra space of the music control parameters. Perform an exponential mapping operation on the control vector sequence to obtain the music control parameter trajectory embedded in the Lie group trajectory space; The trajectory of the music control parameters is passed as a conditional input to the music generation model to guide the model to generate a melody fragment that matches the emotional state. Based on user emotional feedback information during the playback of the melody segment, the emotion mapping model is dynamically fine-tuned and updated, thereby forming a closed-loop adaptive music generation mechanism.

2. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The emotional state information includes: Emotional feature information from multimodal data sources, including user facial images, audio recordings, and natural language text; The facial images are used to extract the user's facial expressions, the audio is used to analyze tone, speech rate and intensity, and the text information is used to identify the emotional tendencies in the language expression. The multimodal emotional features, after being fused, constitute the original input of the emotional state.

3. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The emotional state function includes: A continuous function structure based on time series is used to represent the user's emotional evolution process over continuous time. This function contains at least two components representing pleasure and arousal, which respectively reflect the positiveness and activation state of the user's subjective feelings, and has the characteristic of changing smoothly over time.

4. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The music control vector sequence includes: A multidimensional set of parameters used to control the music generation model, with each control vector corresponding to the music generation state at a point in time; The vector includes rhythm speed parameters, melody center pitch parameters, pitch fluctuation amplitude parameters, and mode type parameters, which are used to adjust the rhythm style, pitch range distribution, and emotional expression of the generated melody, respectively.

5. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The Lie algebra space includes: The vector space defined in the real number field constitutes the basic structure for the evolution of the trajectory of music control parameters; This space supports addition and scalar multiplication operations and has continuous differentiability.

6. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The exponential mapping operation includes: The process of mapping a sequence of music control vectors defined in a Lie algebra space to a Lie group trajectory space; The mapping maintains the continuity and differentiability of the control parameters over time.

7. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The music control parameter trajectory includes: A set of time-varying control parameters that evolve in the Lie group trajectory space, which provide guiding signals during melody generation; The trajectory is controllable and continuous, and is used to adjust the rhythm, mode, melodic direction and note density of the melody, so that the generated music is more in line with the target emotional intention.

8. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The melody fragment includes: The music generation model outputs a sequence of notes under the guidance of a given control parameter trajectory, and each melodic fragment can be used as a component of a complete musical structure; The musical note sequence is organized and arranged according to the rhythm and emotional characteristics encoded by the control parameters, and can express the musical style and emotional color that matches the input emotional state.

9. The intelligent music generation algorithm for dynamically adapting to environmental emotions according to claim 1, characterized in that, The dynamic fine-tuning update includes: Based on the current user's emotional feedback, the parameters of some neural network weights in the emotional mapping model are fine-tuned. The fine-tuning process employs an incremental learning strategy to maintain the stability of the original model structure; Simultaneously, the response sensitivity of the mapping model is dynamically adjusted based on the feedback of emotional bias.

10. A dynamically adaptable intelligent music generation system, comprising an intelligent music generation algorithm for dynamically adapting to environmental emotions according to any one of claims 1-9, characterized in that, include: The emotion recognition module is used to collect multimodal data of the user's current environment, including facial images, audio, and text input, and to identify the user's current emotional state based on the data, thereby constructing a continuous emotion state function that changes over time. The emotion mapping module is used to receive the emotion state function and convert it into a sequence of music control vectors defined in the Lie algebra space through a pre-trained mapping model, so as to express the high-level control intention of music generation. The trajectory embedding module is used to perform an exponential mapping operation on the control vector sequence, embedding it into the Lie group trajectory space to obtain a music control parameter trajectory with time continuity and differentiability. The music generation module is used to take the music control parameter trajectory as a guiding condition input, combine it with the historical note sequence, and generate a melody fragment that matches the user's current emotional state through a neural network model. The feedback control module is used to monitor the user's real-time emotional feedback during the playback of the melody segment, and to fine-tune and update the parameters of the emotion mapping module based on the feedback.