Music adaptive generation method and system based on multi-modal emotion recognition

By collecting facial visual and voice physiological data, fusing multi-dimensional emotional feature vectors, constructing an emotional gradient intensity sequence and direction consistency index, and using sliding fitting to process the emotional evolution trend, adaptive music parameters are generated. This solves the shortcomings of existing technologies in music generation and user emotion adaptation, and achieves dynamic and accurate matching and smooth experience between music and user emotions.

CN122369409APending Publication Date: 2026-07-10SHENZHEN KUAIGE INTELLIGENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610265067.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-05
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies cannot accurately capture users' continuously evolving discrete emotional tags, making it difficult to achieve music generation and user self-adaptation, and lacking in-depth mining and modeling of the temporal continuity and evolution trend of emotions.

Method used

By collecting facial visual features and speech physiological data through sensors, fusing multi-dimensional emotional feature vectors, extracting emotional dependence intensity, constructing emotional gradient intensity sequence and direction consistency index, performing sliding fitting processing, analyzing emotional evolution trend vector, generating initial music parameter sequence, and optimizing music parameters through filtering and smoothing processing to achieve adaptive music generation.

Benefits of technology

It accurately identifies the unidirectional gradual change in user emotions, realizes the continuous and progressive planning of music parameters, ensures the dynamic and accurate matching between music and user emotions, and improves the smoothness and immersion of the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369409A_ABST
    Figure CN122369409A_ABST
Patent Text Reader

Abstract

This invention relates to the field of artificial intelligence technology and discloses a method and system for adaptive music generation based on multimodal emotion recognition. The method includes obtaining a multidimensional emotion feature vector through sensors; determining an emotion gradient intensity sequence and direction consistency index based on the multidimensional emotion feature vector; if the emotion gradient intensity sequence and direction consistency index meet certain conditions, obtaining an emotion evolution trend vector through sliding fitting; obtaining an initial music parameter sequence based on the emotion evolution trend vector; obtaining a current feedback vector based on the initial music parameter sequence; comparing the angle between the current feedback vector and the emotion evolution trend vector; if the angle is greater than a preset angle threshold, re-planning to obtain an optimized music parameter sequence; and synthesizing an adaptive music stream audio based on the optimized music parameter sequence. This method can solve the problem of existing technologies failing to achieve natural and smooth synchronization between music and user emotions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a music adaptive generation method and system based on multimodal emotion recognition. Background Technology

[0002] Currently, with the deepening application of artificial intelligence technology in the field of music creation, how to use multimodal emotion recognition to drive music generation technology has become a research hotspot in the field of artificial intelligence technology.

[0003] Current technologies often collect user emotional information through single or limited modalities such as images and physiological data. After classifying and labeling the collected discrete emotional features, a single-moment emotional tag is obtained and directly mapped to a preset music parameter library. Pre-made music segments are then selected or switched for playback. Some systems attempt to combine temporal emotional data, but only perform simple temporal splicing of discrete emotional tags, without in-depth mining and modeling of the continuous evolution of emotions. The spliced ​​emotional tag sequence is directly matched one by one with corresponding music styles, rhythms, and harmonic elements. However, such solutions fail to fully utilize the technological advantages of big data analysis, lacking analysis of the temporal continuity and evolutionary trends of discrete emotional tags. The rigid mapping of discrete emotional tags to music parameters fails to capture the gradual and fluctuating characteristics of user emotions, and also struggles to maintain the structural consistency of core elements such as music rhythm and harmony.

[0004] In summary, existing technologies cannot accurately capture the continuously evolving discrete emotional tags of users, making it difficult to achieve music generation and user adaptive adaptation. Summary of the Invention

[0005] This invention provides a music adaptive generation method and system based on multimodal emotion recognition to solve the problem that existing technologies cannot achieve natural and smooth synchronization between music and user emotions.

[0006] In a first aspect, to address the aforementioned technical problems, this invention provides a music adaptive generation method based on multimodal emotion recognition, comprising: By collecting temporal data of facial visual features and physiological data of speech through sensors, and fusing the temporal data of facial visual features and the physiological data of speech, a multidimensional emotion feature vector is obtained. Based on the multidimensional emotional feature vector, the emotional dependence intensity is extracted, and the emotional dependence intensity is arranged in time sequence to obtain the emotional gradient intensity sequence. A displacement vector is constructed on the emotional gradient intensity sequence, and the cosine of the angle between the displacement vectors is calculated to obtain the direction consistency index. If the emotional intensity gradient sequence changes monotonically and the directional consistency index exceeds a preset directional threshold, then the recent temporal sequence of the multidimensional emotional feature vector is subjected to sliding fitting to obtain an emotional evolution trend vector. The emotional evolution trend vector is analyzed to obtain rhythmic and harmonic feature components. The rhythmic and harmonic feature components are then preprocessed to obtain an initial music parameter sequence. The preset audio synthesis engine is driven by the initial music parameter sequence to obtain the actual emotional feature vector. The difference operation is then performed on the expected emotional baseline features obtained by mapping the initial music parameter sequence to obtain the current feedback vector. Compare the angle between the current feedback vector and the emotion evolution trend vector. If the angle is greater than a preset angle threshold, then re-plan based on the actual emotion feature vector, and obtain an optimized music parameter sequence after filtering and smoothing. Based on the optimized music parameter sequence, the timbre texture density vector and pitch micro-variation trajectory are deconstructed. The two are combined to generate the initial timbre texture and perform phase alignment and pitch shifting to obtain a smooth timbre texture layer. The smooth timbre texture layer is then modulated to obtain an adaptive music stream audio.

[0007] Secondly, the present invention provides a music adaptive generation system based on multimodal emotion recognition, comprising: The feature acquisition module collects temporal data of facial visual features and physiological data of speech through sensors, and fuses the temporal data of facial visual features and the physiological data of speech to obtain a multidimensional emotion feature vector. The emotional feature fusion module extracts the emotional dependence intensity based on the multi-dimensional emotional feature vector, arranges the emotional dependence intensity in chronological order to obtain an emotional gradient intensity sequence, constructs a displacement vector for the emotional gradient intensity sequence, and calculates the cosine of the angle between the displacement vectors to obtain a direction consistency index. The sentiment trend analysis module performs sliding fitting on the recent time series of the multidimensional sentiment feature vector if the sentiment intensity sequence changes monotonically and the direction consistency index exceeds a preset direction threshold, in order to obtain the sentiment evolution trend vector. The music parameter generation module parses the emotional evolution trend vector to obtain rhythmic feature components and harmonic feature components, and preprocesses the rhythmic feature components and harmonic feature components to obtain an initial music parameter sequence. The emotion feedback acquisition module drives a preset audio synthesis engine based on the initial music parameter sequence to obtain the actual emotion feature vector, and performs a difference operation on the expected emotion benchmark feature obtained by mapping the initial music parameter sequence to obtain the current feedback vector; The music parameter optimization module compares the angle between the current feedback vector and the emotion evolution trend vector. If the angle is greater than a preset angle threshold, the module re-plans the parameters based on the actual emotion feature vector and obtains an optimized music parameter sequence after filtering and smoothing. The music stream synthesis module deconstructs the timbre texture density vector and pitch micro-variation trajectory based on the optimized music parameter sequence, combines the two to perform initial timbre texture generation and phase alignment pitch shifting processing to obtain a smooth timbre texture layer, and modulates the smooth timbre texture layer to obtain adaptive music stream audio.

[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention extracts the dependence strength and determines the gradual change law of multidimensional emotional feature vectors, and captures the continuous evolution trend of emotions by combining sliding fitting processing. It can accurately identify the unidirectional gradual change state of user emotions, effectively solving the defects of existing technologies that only identify discrete emotional labels and do not explore the temporal dependence of emotions, resulting in the inability to predict the direction of emotional evolution and the lack of foresight in music generation. It provides the core basis for the smooth planning of music parameters that fits the user's real state.

[0009] (2) This invention maps the emotional evolution trend vector to the potential spatial features of music parameters, plans the rhythm BPM linear interpolation path and the dark gradient curve, and generates the initial parameter sequence through filtering and smoothing. This realizes the continuous and progressive planning of the core music parameters, solves the problem of the rigid mapping of discrete emotions and music parameters and the abrupt break in the transition in the prior art, and enables the music rhythm and harmonic tone to flow naturally with the emotions, greatly improving the smoothness and immersion of the listening experience.

[0010] (3) This invention determines the matching degree by calculating the angle between the feedback emotional deviation vector and the emotional evolution trend vector, and re-plans the music parameter path and curve by combining the angle size, forming a closed-loop adjustment mechanism of "trend prediction - parameter generation - feedback optimization", which ensures the dynamic and accurate matching between adaptive music and the user's real-time emotions. It solves the problems of existing technologies such as one-way music generation, no feedback correction mechanism, and easy emotional drift, and significantly improves the flexibility and adaptability of adaptive music generation. Attached Figure Description

[0011] Figure 1 This is a schematic diagram of the music adaptive generation method based on multimodal emotion recognition provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the music adaptive generation system based on multimodal emotion recognition provided in the second embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] Reference Figure 1 The first embodiment of the present invention provides a music adaptive generation method based on multimodal emotion recognition, including the following steps: S11, collect facial visual feature time-series data and speech physiological data through sensors, and fuse the facial visual feature time-series data and the speech physiological data to obtain a multi-dimensional emotion feature vector; S12, extract the emotional dependence intensity based on the multidimensional emotional feature vector, arrange the emotional dependence intensity in time sequence to obtain the emotional gradient intensity sequence, construct the displacement vector for the emotional gradient intensity sequence, and calculate the cosine of the angle between the displacement vectors to obtain the direction consistency index. S13, if the emotional intensity sequence changes monotonically and the direction consistency index exceeds the preset direction threshold, then the recent time series of the multidimensional emotional feature vector is subjected to sliding fitting processing to obtain the emotional evolution trend vector. S14, parse the emotional evolution trend vector to obtain rhythmic feature components and harmonic feature components, preprocess the rhythmic feature components and harmonic feature components to obtain the initial music parameter sequence; S15, drive the preset audio synthesis engine according to the initial music parameter sequence to obtain the actual emotional feature vector, and perform a difference operation on the expected emotional benchmark feature obtained by mapping the initial music parameter sequence to obtain the current feedback vector; S16, compare the angle between the current feedback vector and the emotion evolution trend vector. If the angle is greater than a preset angle threshold, then re-plan based on the actual emotion feature vector and obtain an optimized music parameter sequence after filtering and smoothing. S17, based on the optimized music parameter sequence, the timbre texture density vector and pitch micro-variation trajectory are deconstructed, and the two are combined to perform initial timbre texture generation and phase alignment pitch shifting processing to obtain a smooth timbre texture layer. The smooth timbre texture layer is then modulated to obtain adaptive music stream audio.

[0014] In step S11, temporal data of facial visual features and physiological data of speech are collected by sensors. The temporal data of facial visual features and the physiological data of speech are fused to obtain a multidimensional emotion feature vector, including: The facial visual feature time-series data and speech physiological data are collected by sensors. If the timestamp deviation of the facial visual feature time-series data and the speech physiological data is less than a preset deviation threshold, the facial visual feature time-series data and the speech physiological data are subjected to time-series alignment and standardization processing to construct a multimodal joint feature matrix. The multimodal joint feature matrix is ​​encoded using a pre-trained multimodal temporal encoder to obtain the hidden state code. By mapping the hidden layer state encoding through a preset emotion mapping network, a multidimensional emotion feature vector is obtained.

[0015] It should be noted that the sensor is a multimodal sensor group composed of a visual acquisition sensor and a physiological voice acquisition sensor. The visual acquisition sensor is a high-definition industrial camera with a frame rate of no less than 30 frames per second, used to capture dynamic images of the user's face. The physiological voice acquisition sensor is a sound pickup device integrating a heart rate acquisition module, with a sound sampling rate of 16kHz and a heart rate acquisition frequency of 1 beat per second, simultaneously acquiring both voice signals and physiological heart rate signals. The visual acquisition sensor captures the user's facial image in real time, locates 68 feature points in areas such as eyebrows, eyes, and mouth using a facial key point detection algorithm, extracts dynamic information such as the position and displacement rate of these feature points, and continuously records them along a timeline to form temporal data of facial visual features. The physiological voice acquisition sensor acquires the user's voice sound wave signal through the sound pickup module, extracts temporal features such as tone and speech rate, and simultaneously records continuous heart rate variability data through the heart rate acquisition module, integrating these data along a timeline to form physiological voice data.

[0016] The timestamp deviation is calculated by taking the visual feature time-series data and the speech physiological data and taking the data point acquisition time at the same time dimension as the absolute difference. The preset deviation threshold is set by combining the acquisition frequency of facial visual features and speech physiological features and the experimental data on the synchronicity of human emotional expression. No less than 500 sets of human emotional expression synchronicity control experiments were conducted. 100 test subjects of different ages and genders were selected to complete 5 sets of natural expression experiments of different emotions. Multimodal data under different timestamp deviation gradients were collected synchronously and artificially labeled with emotional matching degree. The experimental results showed that when the timestamp deviation between visual and speech physiological data exceeded 100 milliseconds, the emotional matching degree accuracy of the 500 sets of samples was less than 60%, and could not effectively correspond to the same emotional state. When the deviation was ≤80ms, the sample matching degree accuracy was higher than 92%. Therefore, the preset deviation threshold in this invention is set to 80 milliseconds.

[0017] The temporal alignment of the facial visual feature temporal data and speech physiological data is performed using a linear interpolation completion method. The time axis of the facial visual feature temporal data with higher acquisition frequency is used as the reference axis. The speech physiological data is mapped to this reference axis according to its acquisition timestamp. For time points without corresponding speech physiological data on the reference axis, linear interpolation is performed based on the feature values ​​of the two adjacent speech physiological data points. Specifically, feature value change weights are assigned according to the proportion of the time interval between the point to be completed and the points before and after, and the feature value of the point to be completed is calculated to complete the speech physiological feature value for that time point. After temporal alignment, the facial visual feature temporal data and speech physiological data are standardized. For numerical features such as coordinates and displacement in facial visual features, the feature values ​​are mapped to the [-1,1] interval, and the absolute value difference of the feature values ​​is eliminated through linear transformation. For physiological acoustic features such as heart rate and fundamental frequency in speech physiological features, mean-variance standardization is used. The multimodal joint feature matrix is ​​a two-dimensional matrix. Its row dimension consists of continuous data points sorted along the time axis. Each row vector corresponds to all feature data at the same time point. The number of row dimensions is consistent with the total number of data points after time alignment. Its column dimension consists of all feature dimensions of facial visual features and speech physiological features after standardization.

[0018] The constructed multimodal joint feature matrix is ​​input into a pre-trained multimodal temporal encoder. This encoder employs a Transformer architecture, consisting of multiple stacked encoder layers. Each layer contains a multi-head self-attention sublayer and a feedforward fully connected sublayer, with layers connected via residual connections and layer normalization. The multimodal joint feature matrix is ​​processed sequentially, with each row treated as an input feature vector for a given time step. This vector first undergoes a linear transformation through an input embedding layer, mapping the original feature dimensions to a unified hidden layer dimension, and then superimposed with a learnable positional encoding to preserve temporal positional information. After contextual modeling through multiple Transformer blocks, the output vector corresponding to the first special marker [CLS] in the sequence is taken as the aggregated representation of the entire input sequence, yielding the hidden state encoding. This encoding integrates the emotional dependence of facial visual features and speech physiological features in the temporal dimension, with a preset dimension of 256.

[0019] A 256-dimensional hidden state encoding is input into a pre-defined sentiment mapping network. This network employs a two-layer fully connected structure. The first layer maps the 256 dimensions to 128 dimensions using the ReLU activation function. The second layer maps the 128 dimensions to a pre-defined number of sentiment dimensions, such as 8, corresponding to core sentiment components such as pleasure, arousal, tension, and sadness. The output layer uses the Tanh activation function to constrain the values ​​within the range [-1, 1], ultimately yielding a multi-dimensional sentiment feature vector. This sentiment mapping network is pre-trained jointly with a multimodal temporal encoder. The training dataset consists of a large amount of multimodal sentiment temporal data labeled with continuous sentiment dimension values. The loss function is mean squared error, and the optimization objective is to minimize the deviation between the predicted sentiment features and the manually labeled sentiment features, ensuring that the mapped multi-dimensional sentiment feature vector accurately represents the user's real-time emotional state.

[0020] In step S12, emotional dependence intensity is extracted based on the multidimensional emotional feature vector. The emotional dependence intensity is arranged chronologically to obtain an emotional gradient intensity sequence. A displacement vector is constructed from the emotional gradient intensity sequence, and the cosine of the angle between the displacement vectors is calculated to obtain a directional consistency index, including: The multidimensional emotion feature vector is subjected to sliding segmentation to obtain a continuous time window feature subset; Temporal feature extraction is performed on the feature subset of the continuous time window to obtain the hidden layer state vector; Calculate the Euclidean distance between adjacent hidden state vectors to obtain the emotional dependence intensity, and arrange the emotional dependence intensity in time sequence to obtain the emotional gradient intensity sequence; If the emotional gradient intensity sequence does not exceed the preset gradient threshold, then the emotional gradient intensity sequence corresponds to an emotional stable state, and the hidden state vector corresponding to the emotional stable state is the reference vector. If the emotional gradient intensity sequence exceeds a preset gradient threshold, a displacement vector is obtained by performing a difference operation between the hidden state vector and the reference vector. Calculate the cosine of the angle between the displacement vectors, and use the mean of the cosine of the angle as an index of directional consistency.

[0021] It is worth noting that the sliding segmentation of the multidimensional sentiment feature vector is performed using a fixed-length window and an overlapping sliding method. The window length is set to 3 seconds, and the sliding step size is set to 1.5 seconds, meaning that after each sliding, the new window overlaps with the previous window by 50% of the feature data. Following this rule, the multidimensional sentiment feature vector is sequentially segmented along the time axis. The feature vectors within each window form an independent time window feature subset. After continuous segmentation, a continuous time window feature subset arranged in chronological order is obtained.

[0022] In this process, temporal feature extraction is performed on the feature subsets of the continuous time window to obtain the hidden state vector. The Long Short-Term Memory (LSTM) network used is a dedicated network structure for emotional temporal feature extraction. The LSM network employs a single-layer LSM network for temporal feature extraction, with 256 neurons in the hidden layer. The input sequence is a one-dimensional sequence of emotional feature vectors within a single time window feature subset, arranged chronologically. The length of the input sequence is the same as the number of feature vectors within the time window. The input, forget, and output gates all use the Sigmoid activation function, while the cell state activation function uses the Tanh activation function. After the LSM network completes the temporal operation on the feature sequence of the entire time window, it extracts only the 256-dimensional hidden state corresponding to the last time step of the sequence as the feature condensation result for that window. This result is the hidden state vector of the corresponding time window feature subset.

[0023] The training dataset for the Long Short-Term Memory Network is a multimodal emotional temporal dataset, which consists of temporal data of facial visual features, speech physiological data, and corresponding emotional annotations that are consistent with the data collection dimensions of this invention. The dataset contains more than 10,000 valid emotional temporal samples, covering typical emotions such as joy, calmness, tension, sadness, and excitement, as well as the gradual transition states between these emotions. Each sample contains 68 facial key point temporal features, speech fundamental frequency jitter / speech rate unevenness temporal features, and heart rate variability RR interval sequences, and is accompanied by emotional gradient trend and intensity labels annotated by professional annotators. The sample time length is 5s-30s, which matches the emotional collection time in actual applications. At the same time, the dataset is preprocessed with timestamp alignment and standardization, which is completely consistent with the feature processing rules in step S11.

[0024] The network initialization uses Xavier normal initialization to initialize the weight matrix and bias terms of the Long Short-Term Memory (LSTM) network layers. The initial values ​​of the weight matrix follow a normal distribution with a mean of 0 and a variance of 2 / (input dimension + output dimension). The initial value of the bias terms is uniformly set to 0.01. The hidden layer states and cell states of the network are initialized to all zeros at the beginning of each batch of training samples. The training objective is temporal fitting of sentiment features. The loss function is mean squared error loss, which measures the deviation between the hidden layer state vector output by the network and the labeled sentiment feature vector. The Adam optimizer is selected for training, with a learning rate of 0.001. A step decay strategy is used for the learning rate, multiplying by 0.5 every 100 epochs. The batch size is set to 32, and the total number of training epochs is set to 300. During training, the dataset is divided into training, validation, and test sets in a 7:2:1 ratio. The model performance is validated on the validation set after each epoch of training. Training is stopped when the loss on the validation set no longer decreases after 20 consecutive epochs to avoid overfitting. At the same time, generalization validation is performed on the test set, requiring the accuracy of the sentiment feature fitting on the test set to be no less than 92%.

[0025] After the model completes training and passes validation and testing, the weight matrices and bias terms of all layers of the network are fixed, and training-related modules such as optimizers and loss functions used during training are removed, retaining only the forward propagation computation structure of the network. During temporal feature extraction, this pre-trained LSTM network only performs forward propagation operations.

[0026] Specifically, for the hidden state vectors arranged in chronological order, the Euclidean distance between adjacent hidden state vectors is calculated sequentially. This distance value represents the emotional dependence strength. A larger distance indicates a more significant difference in emotional representation between adjacent time windows and a more drastic change in emotion. A smaller distance indicates that the emotional state remains relatively stable and the emotional dependence between time windows is higher. The emotional dependence strengths calculated from all adjacent hidden state vectors are arranged sequentially in chronological order to form a one-dimensional numerical sequence. This sequence is the emotional gradient intensity sequence, where each value corresponds to the degree of emotional change between two adjacent time windows.

[0027] The preset gradation threshold is set by combining emotion recognition experimental data with the sensitivity requirements of music parameter adjustments. At least 600 sets of human emotion gradation experiments were conducted, selecting 120 test subjects of different ages and genders. Each subject completed 5 sets of different types of emotion gradation expressions, such as from calm to pleasure, from calm to tension, etc. Emotion gradation intensity sequences were simultaneously collected and generated. Statistical analysis of the obtained emotion gradation intensity sequences showed that when the emotion dependence intensity was below 0.35, it was mostly minor emotional fluctuations or noise interference, while above 0.45 it represented significant real emotional changes. Therefore, the preset gradation threshold in this invention is 0.4.

[0028] If the emotional gradient intensity sequence does not exceed the preset gradient threshold, the emotional gradient intensity sequence corresponds to an emotional stable state. When the emotional gradient intensity sequence exceeds the preset gradient threshold, the hidden state vector corresponding to the previous emotional stable state is used as the reference vector. The hidden state vector of the current time window is subtracted from the reference vector, and the resulting difference vector is the displacement vector. The direction of this vector points to the main trend of the user's emotional change, and the magnitude of the vector reflects the amplitude of the emotional change. The displacement vectors corresponding to five consecutive time windows are constructed sequentially, and the cosine value of the angle between two adjacent displacement vectors is calculated. The mean of multiple consecutive cosine values ​​of the angle is used as the direction consistency index, and the index value range is [-1, 1].

[0029] For example, the frame rate of the multidimensional emotion feature vector is 30 frames / second. A fixed window of 3 seconds and a sliding step of 1.5 seconds are used for segmentation, resulting in continuous temporal window feature subsets such as Window 1, Window 2, and Window 3 arranged in chronological order. Each window contains 90 frames of feature vectors. Temporal feature extraction is performed on each window to obtain 256-dimensional hidden state vectors V1, V2, and V3. The Euclidean distance between adjacent vectors is calculated sequentially, yielding emotion dependence intensities of 0.32 and 0.42, arranged in the order of [0.32, 0.42], representing a gradual change in emotion intensity sequence. 0.42 exceeds the preset gradual change threshold of 0.4. Using V1 as the base vector, displacement vectors V2-V1 and V3-V1 are constructed. After normalization of the two displacement vectors, the cosine of the angle between them is calculated to be 0.83, which is taken as the directional consistency index.

[0030] In step S13, if the emotional intensity gradient sequence changes monotonically and the directional consistency index exceeds a preset directional threshold, then the recent temporal sequence of the multidimensional emotional feature vector is subjected to sliding fitting processing to obtain an emotional evolution trend vector, including: If the emotional intensity gradient sequence shows a monotonous change that is continuously increasing or continuously decreasing, and the directional consistency index exceeds a preset directional threshold, then a multimodal temporal feature matrix is ​​constructed based on the multidimensional emotional feature vector. Based on the multimodal temporal feature matrix, a multinomial fitting process is performed to calculate the change slope value and the fitting residual value. The change slope value and the fitting residual value are then integrated sequentially according to the feature dimension to obtain the multimodal fitting parameter set. The parameters in the multimodal fitting parameter set are arranged sequentially according to their dimensions to construct a high-dimensional feature vector, which is then standardized to obtain the emotion evolution trend vector.

[0031] It should be noted that the core criterion for determining whether an emotional gradation intensity sequence is monotonically changing is the consistency of the trend of continuous values ​​in the sequence. Three to five consecutive emotional dependence intensities exceeding a preset gradation threshold are selected as judgment samples. If the sample values ​​show a continuous increasing or decreasing trend, and there is no reverse fluctuation between adjacent values, then the emotional gradation intensity sequence is determined to be monotonically changing. The preset directional threshold is set comprehensively based on the directional stability experiment of continuous evolution of human emotions and the trend prediction sensitivity requirements of music adaptation. At least 800 sets of human emotional unidirectional gradation experiments were conducted, selecting 160 test subjects of different ages and genders. Each subject completed five sets of typical unidirectional emotional gradation expressions, such as from calm to pleasure, from pleasure to tension, etc., and directional consistency indices for each experiment were collected and generated simultaneously.

[0032] Statistical analysis of the experimental data showed that when the directional consistency index was below 0.75, emotional changes were prone to repetitive directional shifts, rather than genuine unidirectional gradual changes. When the index was above 0.80, over 95% of the experimental samples exhibited stable unidirectional emotional evolution. Therefore, the preset directional threshold in this invention was set to 0.78. The multimodal temporal feature matrix was constructed with the time axis as the row dimension. Five time windows were arranged sequentially, with each row corresponding to the feature data of one time window. The column dimension was the feature dimension of the real-time multidimensional emotional feature vector, which was consistent with the feature dimension output in step S11, exemplarily set to 8 dimensions, including emotional components such as pleasure, arousal, tension, and sadness. The real-time multidimensional emotional feature vectors of the five time windows were filled into a two-dimensional matrix according to the rule of "time window corresponding to row, feature dimension corresponding to column," forming the multimodal temporal feature matrix.

[0033] The multimodal time series feature matrix is ​​fitted using a quadratic polynomial sliding window method. The sliding window length is consistent with the row dimension of the multimodal time series feature matrix, which covers all feature data of the most recent five time windows without window overlap. The fitting process uses the time window as the independent variable and the values ​​of each feature dimension as the dependent variable. The coefficients of the quadratic polynomials for the five feature data points in each column are solved by the least squares method to obtain the quadratic polynomial fitting curve corresponding to each feature dimension.

[0034] The slope value is the first derivative of the fitted curve for each feature dimension. The last time window number is used as the calculation node. A positive slope indicates that the emotion dimension is increasing, while a negative slope indicates that it is decreasing. The larger the absolute value of the slope, the more drastic the emotion change.

[0035] The fitting residual values ​​are the average absolute values ​​of the deviations between the actual values ​​of each feature dimension and the predicted values ​​on the corresponding fitting curves. The smaller the residual value, the better the fitting effect and the stronger the regularity of sentiment evolution. The slope values ​​corresponding to the changes of all feature dimensions in the multimodal time series feature matrix are integrated with the fitting residual values ​​in the order of feature dimensions to form a one-dimensional parameter set containing a slope subset and a residual subset. This set is the multimodal fitting parameter set.

[0036] A high-dimensional feature vector is constructed based on the multimodal fitting parameter set. All values ​​in the parameter set are arranged in a fixed order, with the slope value first and the residual value last, forming a one-dimensional high-dimensional feature vector. The dimension of the vector is twice the column dimension of the multimodal time-series feature matrix. The constructed high-dimensional feature vector is subjected to interval mapping standardization, which is a linear extremum mapping. The maximum and minimum values ​​of all values ​​in the high-dimensional feature vector are used as the mapping reference. Each value in the vector is mapped to the interval [-1, 1] through a linear transformation, while maintaining the relative magnitude relationship between the values ​​during the transformation.

[0037] For example, in this embodiment, four consecutive emotional dependence intensities exceeding 0.4 [0.42, 0.45, 0.47, 0.50] in the emotional gradient intensity sequence are selected, showing a continuous increasing trend, which is judged as monotonic change, and its directional consistency index is 0.81, which exceeds the preset directional threshold of 0.78. Eight-dimensional real-time multi-dimensional sentiment feature vectors from the five most recent time windows were selected, and a 5x8 multimodal temporal feature matrix was constructed in chronological order. A quadratic polynomial sliding fit was performed on the feature data in each column of the matrix to obtain the slope values ​​[0.08, 0.09, -0.02, -0.03, 0.07, 0.01, -0.01, 0.06] and the fitting residual values ​​[0.05, 0.04, 0.06, 0.07, 0.04, 0.08, 0.07, 0.05] corresponding to the eight feature dimensions. These were then integrated to form a 16-dimensional multimodal fitting parameter set. This parameter set was used to construct a 16-dimensional high-dimensional feature vector in sequence. After interval mapping standardization, the sentiment evolution trend vector of each dimension value in the interval [-1, 1] was obtained. This vector accurately represents the unidirectional gradual trend of the user from calm to pleasure.

[0038] In step S14, the emotional evolution trend vector is parsed to obtain rhythmic feature components and harmonic feature components. The rhythmic feature components and harmonic feature components are preprocessed to obtain an initial music parameter sequence.

[0039] Specifically, the emotional evolution trend vector is analyzed to obtain rhythmic and harmonic feature components. Preprocessing of the rhythmic and harmonic feature components includes: The emotional evolution trend vector is processed by nonlinear mapping and dimensional transformation to obtain the music parameter feature vector; The dimension of the music parameter feature vector is divided into two parts according to a preset ratio. The first half is the rhythm feature component, and the second half is the harmony feature component. The rhythmic feature components are numerically integrated and range-mapped to obtain the target BPM value, and a BPM change path is constructed by combining linear interpolation. The harmonic feature components are numerically integrated and range-mapped to obtain target and declared darkness values, and then combined with Sigmoid nonlinear interpolation to generate harmonic color change curves.

[0040] It should be noted that the music parameter feature vector obtained by mapping the emotional evolution trend vector is achieved through a generator of a conditional generative adversarial network. The network structure of this generator is a network structure with 3 fully connected layers, batch normalization, and ReLU activation. The input layer dimension is consistent with the emotional evolution trend vector dimension, which is 256 dimensions for example. The hidden layer dimensions are set to 512 dimensions and 256 dimensions respectively. The output layer dimension is the music parameter feature vector dimension, which is 64 dimensions. A batch normalization layer is added between each layer to eliminate the influence of dimensions. The activation function uses ReLU to ensure the non-linear expression of the feature mapping. The output layer uses the Sigmoid function to map the output value to the [0,1] interval. The training dataset consists of emotional evolution trend vector samples and corresponding labeled music parameter feature vector samples, containing no less than 5,000 valid samples, covering typical emotions such as joy, calmness, tension, and sadness, as well as gradual transition states. The music parameter feature vector samples are labeled by professional music producers according to emotional trends, containing core potential features of rhythm, harmony, and timbre. All samples have undergone interval mapping standardization, linearly mapping the feature values ​​of all samples to the [0,1] interval, which is consistent with the range of values ​​of the generated emotional evolution trend vector.

[0041] The generator and the discriminator of the conditional generative adversarial network are jointly trained. The discriminator adopts a network structure with three fully connected layers and LeakyReLU activation. The input layer is a concatenated vector of the emotion evolution trend vector and the music parameter feature vector output by the generator, with a dimension of 256 + 64 = 320. The hidden layers have dimensions of 256 and 128 respectively. The output layer is a 1-dimensional scalar used to distinguish whether the input latent space feature vector is a real labeled sample or a fake sample from the generator. The discriminator and the generator are trained synchronously. The activation function is LeakyReLU with a negative slope of 0.2 to avoid gradient vanishing. The Adam optimizer is used with a learning rate of 0.0002. It shares the training dataset and strategy with the generator, using the sentiment evolution trend vector as the conditional input. The loss function employs Wasserstein loss combined with gradient penalty. The Adam optimizer is selected with a learning rate of 0.0002, a batch size of 64, and a total of 500 training epochs. During training, the training and validation sets are divided in an 8:2 ratio. Training stops when the validation set loss no longer decreases for 30 consecutive epochs. After training, the generator weights are fixed, and only forward propagation is performed. The standardized sentiment evolution trend vector is input into the pre-trained generator. After nonlinear mapping and batch normalization through three fully connected layers, a 64-dimensional high-dimensional dense feature vector is output, which is the music parameter feature vector.

[0042] The analysis of the music parameter feature vector adopts a dimensional block mapping method. First, the 64-dimensional music parameter feature vector is divided into rhythmic feature components and harmonic feature components in a 1:1 ratio. For example, the first 32 dimensions are divided into rhythmic feature components and the last 32 dimensions are divided into harmonic feature components. The rhythmic feature components encode the core information such as rhythmic pulsation and speed changes corresponding to the emotional trend, while the harmonic feature components encode the core information such as harmony darkness, color tendency, and tonality characteristics corresponding to the emotional trend. During the analysis process, the values ​​of each dimension remain unchanged, and only dimensional block extraction is performed. The extracted rhythmic and harmonic feature components are both 32-dimensional, and the values ​​are still in the range of [0,1].

[0043] The BPM change path is constructed based on the rhythmic feature components. First, the 32-dimensional rhythmic feature components are mapped to a single value through a one-dimensional fully connected layer. The one-dimensional fully connected layer is a single hidden layer fully connected structure with a 32-dimensional input dimension and a 1-dimensional output dimension. There are no additional activation functions within the layer. The value mapped by the one-dimensional fully connected layer is mapped to a preset BPM value range [60, 180] to obtain the target BPM value. The preset BPM value range is determined based on the industry-standard range of matching emotions with music rhythm and human physiological perception experiments. 60 BPM is a soothing slow rhythm, matching calm and sad emotions, and 180 BPM is a fast rhythm, matching excited and tense emotions. This range covers all typical music rhythm intervals corresponding to human emotions. It has been verified by no less than 500 sets of human auditory experiments. The BPM change within this range is linearly positively correlated with human emotional arousal and there is no physiological discomfort. It is the optimal rhythm interval for emotion-adaptive music generation. The current BPM value is the system's real-time baseline BPM value, set to 120. The transition duration is determined based on the rate of emotional evolution; the more drastic the emotional change, the shorter the transition duration. For example, the transition duration is taken as 5 seconds. With time as the independent variable and BPM value as the dependent variable, linear interpolation is performed between the current BPM value and the target BPM value. The instantaneous BPM value at each time point is calculated with a time step of 0.3 seconds. All instantaneous BPM values ​​are arranged in chronological order to form the BPM change path. The process involves generating a harmonic hue change curve based on the harmonic feature components. First, the 32-dimensional harmonic feature components are mapped to a single numerical value through a one-dimensional fully connected layer. This one-dimensional fully connected layer is a single-hidden-layer fully connected structure with a 32-dimensional input and a 1-dimensional output, and no additional activation function. The one-dimensional fully connected layer maps the 32-dimensional harmonic feature components to an intermediate value, which is then linearly mapped to the range [0,1] to obtain the target and declared darkness values. 0 represents the darkest harmonic hue, and 1 represents the brightest harmonic hue. The current and declared darkness values ​​are set to the system's real-time baseline value of 0.4. The transition duration is consistent with the rhythm BPM transition duration, which is 5 seconds. A sigmoid non-linear interpolation method is used to generate the gradient curve, with time as the independent variable and the declared darkness value as the dependent variable. Sigmoid interpolation is performed between the current value and the target value, calculating the declared darkness value at each time point with a time step of 0.3 seconds. All values ​​are arranged in chronological order to form the harmonic hue change curve.

[0044] Specifically, in one implementation, obtaining the initial music parameter sequence includes: The standard frequency of the main tone of the music is preset as the reference frequency. The trigger interval corresponding to the reference frequency is calculated according to each instantaneous BPM value in the BPM change path. The rhythm frequency value is calculated according to the trigger interval. The rhythm frequency value is arranged in chronological order to obtain the initial rhythm frequency sequence. The brightness values ​​in the harmonic color change curve are calculated according to a preset numerical mapping rule to obtain the corresponding four-dimensional harmonic parameter values. The four-dimensional harmonic parameter values ​​are then arranged in chronological order to obtain the harmonic parameter sequence. The initial rhythm frequency sequence and the harmony parameter sequence are merged and subjected to moving average filtering to obtain the initial music parameter sequence.

[0045] It is worth noting that the reference frequency is selected from the standard frequency of the preset musical tonic, for example, A4=440Hz, which is the basic trigger frequency of the core rhythm instruments, such as percussion and low-frequency pulse instruments. The modulation relationship between the instantaneous BPM value of the rhythm and the reference frequency is as follows: the BPM value determines the trigger interval of the reference frequency. The larger the BPM value, the shorter the trigger interval and the higher the rhythm frequency. According to each instantaneous BPM value in the BPM change path, the trigger interval of the corresponding reference frequency is calculated, where the trigger interval is 60 divided by BPM. Then, according to the rule that the rhythm frequency is equal to 1 divided by the trigger interval, the trigger interval is converted into a continuous rhythm frequency value in Hz. The rhythm frequency values ​​of all time points are arranged in chronological order to form an initial rhythm frequency sequence. During the modulation process, the reference frequency remains unchanged, and frequency modulation is achieved by adjusting the trigger interval only through the BPM value. The time step of the sequence is consistent with the BPM transition path.

[0046] The mapping between harmonic timbre brightness values ​​and harmonic parameters employs a pre-defined one-to-one correspondence rule. This rule is based on professional music theory and validated through 800 sets of human experiments verifying emotion-harmonic matching. According to music theory, minor triads, low cutoff frequencies, and low overtone gains correspond to dark harmonic timbre, while major seventh chords, high cutoff frequencies, and high overtone gains correspond to bright harmonic timbre. Experiments allow test subjects to subjectively rate the harmonic parameters matched to different emotional trends. The parameter combination with the highest score is selected as the correspondence rule for each brightness value, ensuring the auditory matching degree between harmonic timbre and emotional trend. Harmonic parameters include four core dimensions: root note position, chord type, filter cutoff frequency, and overtone gain. These are all key parameters in music synthesis, linearly mapping the harmonic timbre brightness value range [0,1] to the pre-defined value range of each harmonic parameter.

[0047] For example, the root note position is mapped to [C3, G5], the chord type is gradually mapped from a dark minor triad to a bright major seventh chord, the filter cutoff frequency is mapped to [500Hz, 5000Hz], and the overtone gain is mapped to [0, 0.8]. Based on the value of each time point in the harmonic color change curve, the corresponding harmonic parameter values ​​in the four dimensions are calculated according to the mapping rules. The four-dimensional harmonic parameter values ​​of all time points are arranged in chronological order to form a harmonic parameter sequence. The time step of the sequence is consistent with the chord color darkness gradient curve, and all parameter values ​​are standardized. The standardization process adopts multi-dimensional interval mapping standardization, which linearly maps the root note position, filter cutoff frequency, and overtone gain to the [0, 1] interval according to their preset value ranges. The chord type is assigned values ​​of 0.0, 0.33, 0.66, and 1.0 according to the brightness gradient of "minor triad, major triad, seventh chord, major seventh chord".

[0048] The initial rhythm frequency sequence is a one-dimensional numerical sequence, and the harmony parameter sequence is a four-dimensional numerical sequence. Both have identical time steps and total durations. A column-merging method aligned along the time axis is used to integrate the rhythm frequency values ​​and four-dimensional harmony parameter values ​​at the same time point into a single five-dimensional feature vector. These five-dimensional feature vectors are then arranged sequentially according to time, forming the merged five-dimensional music parameter sequence. A fixed-length, non-overlapping moving average filter is applied to the merged five-dimensional music parameter sequence. The filter window length is set to three time steps (0.9 seconds). The filter slides along the sequence in chronological order, calculating the mean of each of the three five-dimensional feature vectors within each window. This mean is used as the filtered value at the center time point of the window. Edge padding is applied to the beginning and end of the sequence during the filtering process to ensure that the duration of the filtered sequence matches that of the original sequence. The five-dimensional music parameter sequence after moving average filtering is the initial music parameter sequence.

[0049] For example, a 256-dimensional emotional evolution trend vector representing the gradual transition from calm to pleasure is input into a pre-trained conditional generative adversarial network generator, which outputs a 64-dimensional music parameter feature vector. The first 32 dimensions and the last 32 dimensions are analyzed to obtain rhythmic and harmonic feature components. The rhythmic feature component is mapped to a target BPM value of 140. Using the current BPM of 120 as a baseline, a 5-second transition duration and a 0.3-second step size are used for linear interpolation to generate a BPM change path. The harmonic feature component is mapped to a target and a statement darkness of 0.8. Using the current darkness of 0.4 as a baseline, a 5-second transition duration is used for linear interpolation. The duration of the passage is interpolated by Sigmoid to generate a harmonic color change curve. Using 440Hz as the reference frequency, the trigger interval is modulated according to the BPM path to generate an initial rhythm frequency sequence. The harmonic darkness curve is mapped according to rules to a four-dimensional harmonic parameter sequence including root note position and chord type. The one-dimensional rhythm frequency sequence and the four-dimensional harmonic parameter sequence are aligned and merged along the time axis to form a five-dimensional sequence. After a 3-step moving average filter, a smooth initial music parameter sequence is obtained. This sequence perfectly matches the user's emotional gradual change from calm to pleasant, with no parameter jumps.

[0050] In step S15, a preset audio synthesis engine is driven according to the initial music parameter sequence to obtain the actual emotional feature vector. This vector is then combined with the expected emotional baseline features mapped from the initial music parameter sequence and subjected to a difference operation to obtain the current feedback vector, including: The initial music parameter sequence drives a preset audio synthesis engine to synthesize short music segments, and at the same time, feature mapping is performed based on the initial music parameter sequence to generate corresponding expected emotional benchmark features. Collect time-series data of user facial visual features and voice physiological data during the short music clip playback period, and fuse them to obtain an actual emotional feature vector. The difference between the actual emotional feature vector and the expected emotional baseline feature is calculated dimension by dimension, and the difference is arranged according to the feature dimension to obtain the current feedback vector.

[0051] It should be noted that the expected sentiment baseline feature vector is constructed using a music parameter-sentiment feature mapping network. This network employs a deep fully connected network structure with three fully connected layers. The input layer dimension is consistent with the feature dimension of the initial smoothed music parameter sequence, the hidden layer dimensions are set to 256 and 128 dimensions respectively, and the output layer dimension is 32 dimensions. Each fully connected layer is followed by a batch normalization layer to eliminate the influence of dimensions. The ReLU activation function is used uniformly, and the Tanh activation function is used in the output layer to map the output values ​​to the [-1, 1] interval. The pre-training dataset contains no fewer than 6000 pairs of initial music parameter feature vector-sentiment feature vector pairs. The samples are labeled by professional music producers with corresponding emotional states, covering typical emotions such as joy, calmness, tension, and sadness, as well as gradual transition states. The dataset preprocessing uses interval mapping normalization to linearly map the music parameter feature vector values ​​to the [-1, 1] interval. The network pre-training uses the mean squared error loss function, with the Adam optimizer selected. The learning rate is set to 0.0005, the batch size to 32, and the total number of training epochs to 300. The training and validation sets are divided in an 8:2 ratio. Training stops when the validation set loss stops decreasing for 20 consecutive epochs. After training, the network weights are fixed, and only forward propagation is performed. The global music parameter feature vector is obtained by averaging the smoothed sequence of initial music parameters along the time axis. This vector is input into the pre-trained mapping network, and after nonlinear mapping, it outputs a 32-dimensional standardized feature vector, which is the expected sentiment baseline feature vector.

[0052] The emotional components of the 32-dimensional expected emotional baseline feature vector are arranged in a fixed order, specifically divided as follows: pleasure in dimensions 1-8, arousal in dimensions 9-16, tension in dimensions 17-24, and sadness in dimensions 25-32. All dimensions are standardized and fall within the range of [-1, 1], where positive values ​​indicate high intensity, negative values ​​indicate low intensity, and zero values ​​indicate no significant emotional tendency. When driving a preset digital audio synthesis engine to generate short music clips, the initial smoothed sequence of music parameters is input into the professional digital audio synthesis engine at a time granularity of 0.05 seconds / frame. The preset digital audio synthesis engine can be based on FM synthesis or physical modeling synthesis. The rhythm frequency dimension controls the trigger interval of percussion and low-frequency pulses, while the harmony parameter dimension controls chord type, filter cutoff frequency, overtone gain, and root note position. Short music clips of 5-8 seconds are synthesized, with the clip duration matching the time window for emotional feature acquisition.

[0053] The multi-dimensional emotional feature vector contained in the short music segment is mapped using a multimodal emotional feature mapping network. The network structure employs a one-dimensional convolutional layer, cross-modal attention fusion, and two fully connected layers. The input layer is a multi-dimensional emotional feature vector within the short music playback period, exemplified by an 8-dimensional vector containing facial micro-expressions, heart rate variability, and speech features. The sequence length is 100 frames, corresponding to a 5-second playback duration. First, two one-dimensional convolutional layers with a kernel size of 3, a stride of 1, and 64 output channels are used to extract single-modal temporal features. Then, cross-modal attention fusion is performed... The attention mechanism weights and fuses the importance of each feature dimension. The core logic is as follows: First, the single-modal temporal features extracted by two one-dimensional convolutional layers are mapped to a 64-dimensional feature space. Then, facial micro-expression features are used as query vectors, and heart rate variability features and speech features are used as key-value pairs. The attention weights between the query vector and each key-value pair are calculated. These weights represent the importance of different modal features to the current emotional state. Finally, all the weighted modal features are summed and fused to obtain a fused multimodal temporal feature vector. Finally, two fully connected layers with dimensions of 128 and 32 are used to map the vector to a fixed-dimensional feature vector. The output layer uses the Tanh activation function to map the values ​​to the [-1,1] interval to obtain a 32-dimensional actual emotional feature vector.

[0054] The training dataset contains no fewer than 8000 sets of multimodal emotion feature vector sequences-emotion feature vector pairing samples, covering typical emotions and gradual states such as joy, calmness, tension, and sadness. The multimodal emotion feature vector sequences have the same feature dimensions as those collected in step S11. The emotion feature vectors are labeled by professional emotion annotators according to the corresponding emotion states of the samples. All samples are standardized by interval mapping, with a uniform numerical range of [-1, 1]. The mean squared error loss function is used in network training to measure the deviation between the network's output emotion feature vectors and the manually labeled vectors. The Adam optimizer is used, with a learning rate of 0.001, a batch size of 64, and a total of 400 training epochs. The training, validation, and test sets are divided in a 7:2:1 ratio. Training stops when the validation set loss no longer decreases for 25 consecutive epochs. The emotion feature mapping accuracy of the test set is no less than 93%. After training, the network weights are fixed, and only forward propagation is performed.

[0055] The mapping process involves extracting multi-dimensional emotional feature vectors from a short music clip playback period, standardizing them, and then inputting them into a pre-trained multimodal emotional feature mapping network. After convolutional feature extraction, cross-modal fusion, and mapping through fully connected layers, a 32-dimensional standardized feature vector is directly output, which is the actual emotional feature vector.

[0056] The current feedback vector is obtained by calculating the difference between the actual emotional feature vector and the expected emotional baseline feature vector. The emotional deviation vector is obtained by calculating the difference across dimensions, that is, by successively subtracting the expected baseline value from the actual emotional feature vector, and arranging the differences across all dimensions in a fixed order to form a 32-dimensional current feedback vector. In the deviation vector, a positive difference in a certain dimension indicates that the actual response of that emotional component is stronger than expected, while a negative difference indicates that the actual response is weaker than expected. The absolute value of the difference represents the degree of deviation in that dimension.

[0057] For example, a smoothed sequence of 5-dimensional initial music parameters (i.e., the average of 1-dimensional rhythm frequency and 4-dimensional harmony parameters) is input into a music parameter-emotion feature mapping network, outputting a 32-dimensional expected emotion baseline feature vector. Its core emotion components are: pleasure 0.75, arousal 0.70, and tension -0.40. This smoothed sequence is then input into a digital audio synthesis engine at 0.05 seconds / frame to generate a 6-second short music clip. An 8-dimensional multi-dimensional emotion feature vector is extracted from this 6-second clip and input into a pre-trained multimodal emotion feature mapping network to obtain a 32-dimensional actual emotion feature vector. Its core emotion components are: pleasure 0.62, arousal 0.72, and tension -0.25. Differences are calculated dimension-by-dimensionally between the two vectors to obtain the core components of the current feedback vector: pleasure -0.13, arousal +0.02, and tension +0.15. This deviation vector indicates that the user's actual pleasure is lower than expected, and tension is higher than expected.

[0058] In step S16, the angle between the current feedback vector and the emotion evolution trend vector is compared. If the angle is greater than a preset angle threshold, the actual emotion feature vector is re-planned, and an optimized music parameter sequence is obtained after filtering and smoothing.

[0059] Specifically, if the included angle is greater than a preset angle threshold, then the music parameter sequence is re-planned based on the actual emotional feature vector, and after filtering and smoothing, an optimized music parameter sequence is obtained, including: If the included angle is greater than a preset angle threshold, the loss weight ratio is calculated based on the deviation ratio between the included angle and the preset angle threshold. The target BPM value is obtained by numerically integrating and range mapping the rhythm feature components. After adjusting the target BPM value according to the loss weight ratio, a quadratic planning BPM change path is obtained by combining linear interpolation. The harmonic feature components are numerically integrated and range-mapped to obtain target and declared darkness values. After adjusting the target and declared darkness values ​​according to the loss weight ratio, the secondary harmonic color change curve is obtained by combining Sigmoid nonlinear interpolation. The BPM variation path and the secondary harmonic color variation curve of the quadratic programming are filtered and smoothed to obtain an optimized music parameter sequence.

[0060] It should be noted that when comparing the angle between the current feedback vector and the emotional evolution trend vector, the core basis for calculating the angle is the product of the dot product and the product of the magnitudes of the two vectors. The angle is inferred from the cosine value. The closer the cosine value is to 1, the smaller the angle, indicating that the feedback deviation is consistent with the direction of the emotional evolution trend; the closer the cosine value is to -1, the larger the angle, indicating that the feedback deviation deviates from the trend direction. The determination of the preset angle threshold is combined with human emotional auditory adaptation experiments and the efficiency requirements of music parameter optimization. At least 700 sets of emotional deviation auditory tests were conducted, and 140 test subjects listened to music clips with different deviation angles and scored them. The results showed that when the angle was greater than 60 degrees, more than 90% of the test subjects perceived a disharmony between the music and the emotion. Therefore, this invention sets the preset angle threshold to 60 degrees.

[0061] The loss weight ratio is used to quantify the correction intensity of music parameters. Based on the preset angle threshold of 60 degrees, the deviation ratio between the actual angle and the threshold is calculated. If the angle is 60 degrees, the deviation ratio is 0 and the weight ratio is 1. If the angle is 180 degrees, the deviation ratio is 1 and the weight ratio is 3. The loss weight ratio is calculated by subtracting 60 from the actual angle, dividing by 120, multiplying by 2, and adding 1.

[0062] The method for obtaining the secondary planning BPM change path is entirely consistent with the core logic of "constructing the BPM change path based on the rhythmic feature components" in step S14. The only difference is that after obtaining the target BPM value through a one-dimensional fully connected layer mapping, the difference between the target BPM value and the current BPM value is multiplied by the loss weight ratio, and then linear interpolation is performed to achieve a greater range of BPM adjustments. Similarly, the method for obtaining the secondary harmonic color change curve is entirely consistent with the core logic of "generating a harmonic color change curve based on the harmonic feature components" in step S14. The only difference is that after obtaining the target and declared darkness values ​​through a one-dimensional fully connected layer mapping, the difference between the target and declared darkness values ​​and the current values ​​is multiplied by the loss weight ratio, and then Sigmoid nonlinear interpolation is performed to achieve harmonic color adjustments that better match emotional trends. During the replanning process, core parameters such as time step and transition duration remain unchanged; only the adjustment range of the target parameters is amplified by adjusting the loss weight ratio.

[0063] The method of filtering and smoothing the quadratic planning BPM change path and the quadratic harmonic color change curve to obtain an optimized music parameter sequence is completely consistent with the process in step S14 of "filtering and smoothing the instantaneous BPM linear interpolation transition path and the harmonic color change curve to obtain an initial music parameter sequence".

[0064] For example, the angle between the 32-dimensional current feedback vector and the 32-dimensional emotion evolution trend vector is 90 degrees, exceeding the preset angle threshold of 60 degrees. Therefore, the calculated loss weight ratio is 1.5. This weight is applied to the replanning process. The original target BPM value of 140 is multiplied by the weight and adjusted to 150. The original target and declared darkness value of 0.8 is multiplied by the weight and adjusted to 1.0. The BPM path and declared darkness curve are replanned according to the logic of step S14. After reference frequency modulation, harmony parameter mapping, sequence merging, and moving average filtering, an optimized music parameter sequence is obtained. This sequence effectively corrects the deviation between the original music and the user's emotion, achieving precise adaptation.

[0065] In step S17, the timbre texture density vector and pitch micro-variation trajectory are deconstructed based on the optimized music parameter sequence. These are then combined to perform initial timbre texture generation and phase-aligned pitch shifting to obtain a smooth timbre texture layer. The smooth timbre texture layer is then modulated to obtain adaptive music stream audio, including: The optimized music parameter sequence is extracted by dividing it into blocks according to the feature dimension to obtain the timbre texture density vector and the pitch micro-variation trajectory; An initial timbre texture data stream is generated based on the timbre texture density vector, and the initial timbre texture data stream is phase-aligned and pitch-shifted in combination with the pitch micro-variation trajectory to obtain a smooth timbre texture layer. The instantaneous loudness value is calculated frame by frame for the smooth timbre texture layer to obtain the instantaneous loudness envelope, and the preset low-frequency pulsating oscillator is reset according to the peak-valley change of the instantaneous loudness envelope; The smooth timbre texture layer is subjected to amplitude modulation processing based on the low-frequency pulsating oscillator, and then synthesized into adaptive music stream audio after global peak normalization.

[0066] It should be noted that the deconstruction employs a combination of fixed-dimensional block division and temporal unpacking. First, the five-dimensional optimized music parameter sequence is expanded along the time axis, extracting the overtone gain dimension value and root note position dimension value for each time step. The values ​​of other dimensions are not included in the deconstruction at this stage. The overtone gain values ​​of consecutive time steps are arranged sequentially to construct a one-dimensional timbre texture density vector. This vector dimension is consistent with the time step size of the music flow, directly representing the temporal change of particle density with the original values. Using the basic pitch corresponding to the root note position of the extracted consecutive time steps as a reference, a random perturbation signal with an amplitude controlled by a pitch perturbation coefficient of 0.02 is superimposed to generate continuous curve data, which is the pitch micro-variation trajectory. The preset pitch perturbation coefficient of 0.02 was determined through 800 sets of human auditory comfort experiments. The experiments show that a coefficient of 0.02 can ensure that the melody is natural, smooth, and without any sense of incongruity.

[0067] The initial timbre texture data stream is generated using a particle synthesis algorithm. This algorithm employs additive synthesis, using the timbre texture density vector as the sole input. A larger vector value results in a greater number and higher density of generated audio particles, while a smaller value leads to sparser particles. The generated data stream has a fixed sampling rate of 44.1kHz and a bit depth of 16bit, consistent with standard audio. Pitch shifting is performed using a real-time phase-aligned pitch shifting algorithm. This algorithm employs a phase vocoder-based strategy to force the alignment of phase continuity between adjacent frames, ensuring uninterrupted phase continuity during pitch shifting. Using the pitch micro-change trajectory as the pitch shifting reference, each audio frame of the initial timbre texture data stream is pitch-shifted point-by-point, ensuring uninterrupted phase continuity during pitch shifting. After pitch shifting, a 5-millisecond linear crossfade-in / fade-out strategy is used to stitch adjacent audio frames together, eliminating inter-frame abrupt changes caused by pitch shifting, ultimately forming a seamless, smooth timbre texture layer.

[0068] The extraction of the instantaneous loudness envelope employs a 10-millisecond sliding time window. The instantaneous loudness value is calculated frame-by-frame for the audio data of the smooth timbre texture layer. This instantaneous loudness value is calculated using an A-weighted comprehensive loudness algorithm, which calculates short-time energy, performs A-weighted filtering, and converts it to decibels. These values ​​are then arranged chronologically to form a one-dimensional loudness envelope curve, representing the real-time trend of audio loudness changes. The preset low-frequency pulsating oscillator is a basic electronic oscillator commonly used in audio synthesis, specifically a sinusoidal low-frequency oscillator. All parameters have been calibrated using 800 sets of human auditory comfort experiments and emotion-rhythm matching experiments, and are fixed preset values. These include an oscillation frequency of 2Hz, an initial phase of 0rad, a sampling rate of 44100Hz, and a bit depth of 16bit. The preset reset rule for the low-frequency pulsating oscillator is as follows: using the peak point of the loudness envelope curve as the trigger signal, the phase of the oscillator is reset to 0, and the frequency of the oscillator is set to 2Hz. This frequency matches the basic rhythm of human emotions, enhancing emotional immersion. The amplitude is determined by the mean of the loudness envelope. The modulation process uses amplitude modulation, using the output of the reset low-frequency pulsating oscillator as the modulation signal to perform point-by-point amplitude modulation on the audio data of the smooth timbre texture layer, so that the pulsating rhythm of the music is synchronized with the fluctuation trend of the user's emotions. After modulation, the audio data is normalized by using global peak normalization to map to the [-1,1] interval, eliminating the volume overload caused by modulation. The final output continuous audio stream is the adaptive music stream audio.

[0069] For example, the five-dimensional optimized music parameter sequence is deconstructed to obtain a timbre texture density vector representing the improvement in pleasure, with values ​​increasing from 0.3 to 0.7 and a slight pitch change trajectory, the root note slowly curving from C4 to D4. An initial timbre texture data stream is generated using a particle synthesis algorithm, and real-time phase alignment and pitch shifting are performed by combining perturbation path data. After 5 milliseconds of crossfade-in and fade-out, a smooth timbre texture layer is obtained. The instantaneous loudness envelope of this layer is extracted, triggering a phase reset of a 2Hz low-frequency pulsating oscillator. The music pulsation is bound to the loudness change through amplitude modulation. After normalization, a smooth, adaptive music stream audio that is highly adapted to the trend from calm to pleasure is output.

[0070] In summary, this invention solves the problem of existing technologies failing to achieve natural and smooth synchronization between music and user emotions by employing emotion-adaptive generation and a multimodal feedback closed-loop mechanism.

[0071] Reference Figure 2 The second embodiment of the present invention provides a music adaptive generation system based on multimodal emotion recognition, comprising: The feature acquisition module collects temporal data of facial visual features and physiological data of speech through sensors, and fuses the temporal data of facial visual features and the physiological data of speech to obtain a multidimensional emotion feature vector. The emotional feature fusion module extracts the emotional dependence intensity based on the multi-dimensional emotional feature vector, arranges the emotional dependence intensity in chronological order to obtain an emotional gradient intensity sequence, constructs a displacement vector for the emotional gradient intensity sequence, and calculates the cosine of the angle between the displacement vectors to obtain a direction consistency index. The sentiment trend analysis module performs sliding fitting on the recent time series of the multidimensional sentiment feature vector if the sentiment intensity sequence changes monotonically and the direction consistency index exceeds a preset direction threshold, in order to obtain the sentiment evolution trend vector. The music parameter generation module parses the emotional evolution trend vector to obtain rhythmic feature components and harmonic feature components, and preprocesses the rhythmic feature components and harmonic feature components to obtain an initial music parameter sequence. The emotion feedback acquisition module drives a preset audio synthesis engine based on the initial music parameter sequence to obtain the actual emotion feature vector, and performs a difference operation on the expected emotion benchmark feature obtained by mapping the initial music parameter sequence to obtain the current feedback vector; The music parameter optimization module compares the angle between the current feedback vector and the emotion evolution trend vector. If the angle is greater than a preset angle threshold, the module re-plans the parameters based on the actual emotion feature vector and obtains an optimized music parameter sequence after filtering and smoothing. The music stream synthesis module deconstructs the timbre texture density vector and pitch micro-variation trajectory based on the optimized music parameter sequence, combines the two to perform initial timbre texture generation and phase alignment pitch shifting processing to obtain a smooth timbre texture layer, and modulates the smooth timbre texture layer to obtain adaptive music stream audio.

[0072] It should be noted that the music adaptive generation system based on multimodal emotion recognition provided in this embodiment of the invention is used to execute all the process steps of the music adaptive generation method based on multimodal emotion recognition in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0073] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0074] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A music adaptive generation method based on multimodal emotion recognition, characterized in that, include: By collecting temporal data of facial visual features and physiological data of speech through sensors, and fusing the temporal data of facial visual features and the physiological data of speech, a multidimensional emotion feature vector is obtained. Based on the multidimensional emotional feature vector, the emotional dependence intensity is extracted, and the emotional dependence intensity is arranged in time sequence to obtain the emotional gradient intensity sequence. A displacement vector is constructed on the emotional gradient intensity sequence, and the cosine of the angle between the displacement vectors is calculated to obtain the direction consistency index. If the emotional intensity gradient sequence changes monotonically and the directional consistency index exceeds a preset directional threshold, then the recent temporal sequence of the multidimensional emotional feature vector is subjected to sliding fitting to obtain an emotional evolution trend vector. The emotional evolution trend vector is analyzed to obtain rhythmic and harmonic feature components. The rhythmic and harmonic feature components are then preprocessed to obtain an initial music parameter sequence. The preset audio synthesis engine is driven by the initial music parameter sequence to obtain the actual emotional feature vector. The difference operation is then performed on the expected emotional baseline features obtained by mapping the initial music parameter sequence to obtain the current feedback vector. Compare the angle between the current feedback vector and the emotion evolution trend vector. If the angle is greater than a preset angle threshold, then re-plan based on the actual emotion feature vector, and obtain an optimized music parameter sequence after filtering and smoothing. Based on the optimized music parameter sequence, the timbre texture density vector and pitch micro-variation trajectory are deconstructed. The two are combined to generate the initial timbre texture and perform phase alignment and pitch shifting to obtain a smooth timbre texture layer. The smooth timbre texture layer is then modulated to obtain an adaptive music stream audio.

2. The music adaptive generation method based on multimodal emotion recognition according to claim 1, characterized in that, By collecting temporal data of facial visual features and speech physiological data through sensors, and fusing the temporal data of facial visual features and the speech physiological data, a multidimensional emotion feature vector is obtained, including: The facial visual feature time-series data and speech physiological data are collected by sensors. If the timestamp deviation of the facial visual feature time-series data and the speech physiological data is less than a preset deviation threshold, the facial visual feature time-series data and the speech physiological data are subjected to time-series alignment and standardization processing to construct a multimodal joint feature matrix. The multimodal joint feature matrix is ​​encoded using a pre-trained multimodal temporal encoder to obtain the hidden state code. By mapping the hidden layer state encoding through a preset emotion mapping network, a multidimensional emotion feature vector is obtained.

3. The music adaptive generation method based on multimodal emotion recognition according to claim 1, characterized in that, Emotional dependence intensity is extracted based on the multidimensional emotional feature vector. The emotional dependence intensity is then arranged chronologically to obtain an emotional gradient intensity sequence. A displacement vector is constructed from the emotional gradient intensity sequence, and the cosine of the angle between the displacement vectors is calculated to obtain a directional consistency index, including: The multidimensional emotion feature vector is subjected to sliding segmentation to obtain a continuous time window feature subset; Temporal feature extraction is performed on the feature subset of the continuous time window to obtain the hidden layer state vector; Calculate the Euclidean distance between adjacent hidden state vectors to obtain the emotional dependence intensity, and arrange the emotional dependence intensity in time sequence to obtain the emotional gradient intensity sequence; If the emotional gradient intensity sequence does not exceed the preset gradient threshold, then the emotional gradient intensity sequence corresponds to an emotional stable state, and the hidden state vector corresponding to the emotional stable state is the reference vector. If the emotional gradient intensity sequence exceeds a preset gradient threshold, a displacement vector is obtained by performing a difference operation between the hidden state vector and the reference vector. Calculate the cosine of the angle between the displacement vectors, and use the mean of the cosine of the angle as an index of directional consistency.

4. The music adaptive generation method based on multimodal emotion recognition according to claim 1, characterized in that, If the emotional intensity gradient sequence changes monotonically and the directional consistency index exceeds a preset directional threshold, then the recent temporal sequence of the multidimensional emotional feature vector is subjected to sliding fitting to obtain an emotional evolution trend vector, including: If the emotional intensity gradient sequence shows a monotonous change that is continuously increasing or continuously decreasing, and the directional consistency index exceeds a preset directional threshold, then a multimodal temporal feature matrix is ​​constructed based on the multidimensional emotional feature vector. Based on the multimodal temporal feature matrix, a multinomial fitting process is performed to calculate the change slope value and the fitting residual value. The change slope value and the fitting residual value are then integrated sequentially according to the feature dimension to obtain the multimodal fitting parameter set. The parameters in the multimodal fitting parameter set are arranged sequentially according to their dimensions to construct a high-dimensional feature vector, which is then standardized to obtain the emotion evolution trend vector.

5. The music adaptive generation method based on multimodal emotion recognition according to claim 1, characterized in that, The emotional evolution trend vector is analyzed to obtain rhythmic and harmonic feature components. Preprocessing of the rhythmic and harmonic feature components includes: The emotional evolution trend vector is processed by nonlinear mapping and dimensional transformation to obtain the music parameter feature vector; The dimension of the music parameter feature vector is divided into two parts according to a preset ratio. The first half is the rhythm feature component, and the second half is the harmony feature component. The rhythmic feature components are numerically integrated and range-mapped to obtain the target BPM value, and a BPM change path is constructed by combining linear interpolation. The harmonic feature components are numerically integrated and range-mapped to obtain target and declared darkness values, and then combined with Sigmoid nonlinear interpolation to generate harmonic color change curves.

6. The music adaptive generation method based on multimodal emotion recognition according to claim 5, characterized in that, The process of obtaining the initial music parameter sequence includes: The standard frequency of the main tone of the music is preset as the reference frequency. The trigger interval corresponding to the reference frequency is calculated according to each instantaneous BPM value in the BPM change path. The rhythm frequency value is calculated according to the trigger interval. The rhythm frequency value is arranged in chronological order to obtain the initial rhythm frequency sequence. The brightness values ​​in the harmonic color change curve are calculated according to a preset numerical mapping rule to obtain the corresponding four-dimensional harmonic parameter values. The four-dimensional harmonic parameter values ​​are then arranged in chronological order to obtain the harmonic parameter sequence. The initial rhythm frequency sequence and the harmony parameter sequence are merged and subjected to moving average filtering to obtain the initial music parameter sequence.

7. The music adaptive generation method based on multimodal emotion recognition according to claim 1, characterized in that, The preset audio synthesis engine is driven by the initial music parameter sequence to obtain the actual emotional feature vector. This vector is then combined with the expected emotional baseline features mapped from the initial music parameter sequence and subjected to a difference operation to obtain the current feedback vector, which includes: The initial music parameter sequence drives a preset audio synthesis engine to synthesize short music segments, and at the same time, feature mapping is performed based on the initial music parameter sequence to generate corresponding expected emotional benchmark features. Collect time-series data of user facial visual features and voice physiological data during the short music clip playback period, and fuse them to obtain an actual emotional feature vector. The difference between the actual emotional feature vector and the expected emotional baseline feature is calculated dimension by dimension, and the difference is arranged according to the feature dimension to obtain the current feedback vector.

8. The music adaptive generation method based on multimodal emotion recognition according to claim 1, characterized in that, If the included angle is greater than a preset angle threshold, then the actual emotional feature vector is re-planned, and after filtering and smoothing, an optimized music parameter sequence is obtained, including: If the included angle is greater than a preset angle threshold, the loss weight ratio is calculated based on the deviation ratio between the included angle and the preset angle threshold. The target BPM value is obtained by numerically integrating and range mapping the rhythm feature components. After adjusting the target BPM value according to the loss weight ratio, a quadratic planning BPM change path is obtained by combining linear interpolation. The harmonic feature components are numerically integrated and range-mapped to obtain target and declared darkness values. After adjusting the target and declared darkness values ​​according to the loss weight ratio, the secondary harmonic color change curve is obtained by combining Sigmoid nonlinear interpolation. The BPM variation path and the secondary harmonic color variation curve of the quadratic programming are filtered and smoothed to obtain an optimized music parameter sequence.

9. The music adaptive generation method based on multimodal emotion recognition according to claim 1, characterized in that, Based on the optimized music parameter sequence, the timbre texture density vector and pitch micro-variation trajectory are deconstructed. Combining these two elements, initial timbre texture generation and phase-aligned pitch shifting are performed to obtain a smooth timbre texture layer. The smooth timbre texture layer is then modulated to obtain adaptive music stream audio, including: The optimized music parameter sequence is extracted by dividing it into blocks according to the feature dimension to obtain the timbre texture density vector and the pitch micro-variation trajectory; An initial timbre texture data stream is generated based on the timbre texture density vector, and the initial timbre texture data stream is phase-aligned and pitch-shifted in combination with the pitch micro-variation trajectory to obtain a smooth timbre texture layer. The instantaneous loudness value is calculated frame by frame for the smooth timbre texture layer to obtain the instantaneous loudness envelope, and the preset low-frequency pulsating oscillator is reset according to the peak-valley change of the instantaneous loudness envelope; The smooth timbre texture layer is subjected to amplitude modulation processing based on the low-frequency pulsating oscillator, and then synthesized into adaptive music stream audio after global peak normalization.

10. A music adaptive generation system based on multimodal emotion recognition, characterized in that, include: The feature acquisition module collects temporal data of facial visual features and physiological data of speech through sensors, and fuses the temporal data of facial visual features and the physiological data of speech to obtain a multidimensional emotion feature vector. The emotional feature fusion module extracts the emotional dependence intensity based on the multi-dimensional emotional feature vector, arranges the emotional dependence intensity in chronological order to obtain an emotional gradient intensity sequence, constructs a displacement vector for the emotional gradient intensity sequence, and calculates the cosine of the angle between the displacement vectors to obtain a direction consistency index. The sentiment trend analysis module performs sliding fitting on the recent time series of the multidimensional sentiment feature vector if the sentiment intensity sequence changes monotonically and the direction consistency index exceeds a preset direction threshold, in order to obtain the sentiment evolution trend vector. The music parameter generation module parses the emotional evolution trend vector to obtain rhythmic feature components and harmonic feature components, and preprocesses the rhythmic feature components and harmonic feature components to obtain an initial music parameter sequence. The emotion feedback acquisition module drives a preset audio synthesis engine based on the initial music parameter sequence to obtain the actual emotion feature vector, and performs a difference operation on the expected emotion benchmark feature obtained by mapping the initial music parameter sequence to obtain the current feedback vector; The music parameter optimization module compares the angle between the current feedback vector and the emotion evolution trend vector. If the angle is greater than a preset angle threshold, the module re-plans the parameters based on the actual emotion feature vector and obtains an optimized music parameter sequence after filtering and smoothing. The music stream synthesis module deconstructs the timbre texture density vector and pitch micro-variation trajectory based on the optimized music parameter sequence, combines the two to perform initial timbre texture generation and phase alignment pitch shifting processing to obtain a smooth timbre texture layer, and modulates the smooth timbre texture layer to obtain adaptive music stream audio.