A personalized music automatic generation system and method fusing multi-modal emotion perception

By integrating multimodal perception data and historical preference data, an emotional state vector and a music preference vector are constructed to generate a personalized music recommendation table. This solves the problem of difficulty in capturing users' real-time emotional state in existing technologies, and improves the accuracy and timeliness of personalized music recommendations.

CN122153113APending Publication Date: 2026-06-05SHENZHEN XUANTONG INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN XUANTONG INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-03-04
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing personalized music recommendation systems struggle to capture and respond to users' real-time changing emotional states, resulting in a discrepancy between the recommended results and the user's current emotional needs, and failing to achieve real-time, accurate matching driven by emotions.

Method used

By collecting multimodal perception data (physiological signals, facial signals, and interaction signals), cross-modal emotion extraction is performed. Combined with users' historical music preference data, emotional state vectors and music preference vectors are constructed, and a personalized music recommendation table is generated using music matching thresholds.

Benefits of technology

It achieves real-time dynamic adaptation driven by emotions, improving the accuracy and timeliness of personalized music recommendations and ensuring that the recommended content is highly consistent with the user's real-time emotional state and long-term preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153113A_ABST
    Figure CN122153113A_ABST
Patent Text Reader

Abstract

The application provides a personalized music automatic generation system and method fusing multi-modal emotion perception, relates to the technical field of personalized music automatic generation, collects multi-modal perception data of a target user, performs emotion cross-modal extraction on the multi-modal perception data to obtain an emotion state vector; determines a music preference vector based on historical music preference data, performs expected tendency on music preference of the target user according to the emotion state vector and the music preference vector to obtain each music expectation value of the target user under a current emotion state; predicts a current expected music type of the target user according to each music expectation value and a music matching threshold to obtain a music response degree of the target user for each music type; and generates a personalized music recommendation table according to each music response degree. The application can fuse the current emotion state of the user and the historical preference data, realize real-time dynamic adaptation driven by emotion, and improve the accuracy and timeliness of personalized music recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of personalized music automatic generation technology, and more specifically, to a personalized music automatic generation system and method that integrates multimodal emotion perception. Background Technology

[0002] Currently, mainstream music recommendation systems primarily rely on collaborative filtering, content analysis, and hybrid recommendation algorithms. These systems typically build preference models based on users' historical playback records, collection behavior, and explicit rating data, and then generate recommendation lists by analyzing metadata of music content (such as genre, rhythm, and sentiment tags) or mining similarities between user groups. Some advanced systems attempt to introduce basic emotion recognition technology, such as adjusting recommendation strategies by analyzing users' listening periods or simple facial expression data. In addition, deep learning-based music generation technology has also made progress, capable of automatically generating music clips based on constraints such as style and rhythm. Existing technologies have laid the foundation for data modeling and content generation to achieve personalized music services.

[0003] Current personalized music recommendation technologies primarily rely on static analysis of users' historical playback records, collection behavior, and similar user preferences. While this can build long-term interest profiles, it struggles to capture and respond to users' real-time changing emotional states. Some systems attempting to incorporate emotion perception often use single-modal data or outdated manual labels, lacking effective fusion and dynamic analysis of multi-dimensional real-time emotional information such as users' physiological signals, facial expressions, voice tone, and interactive behaviors. This leads to discrepancies between recommendation results and users' current emotional needs, failing to achieve emotion-driven, real-time, and accurate adaptation. This impacts the immediacy and effectiveness of music recommendations in scenarios such as emotional support and stress management. Therefore, how to integrate users' current emotional states and historical preference data to achieve emotion-driven, real-time, dynamic adaptation and improve the accuracy and timeliness of personalized music recommendations is a challenge facing the industry. Summary of the Invention

[0004] This application provides a personalized music automatic generation system and method that integrates multimodal emotion perception. It can integrate the user's current emotional state and historical preference data to achieve real-time dynamic adaptation driven by emotion, thereby improving the accuracy and timeliness of personalized music recommendations.

[0005] In a first aspect, this application provides a personalized music automatic generation method that integrates multimodal emotion perception, the control method comprising the following steps: Collect multimodal perception data of the target user, and perform cross-modal emotion extraction on the multimodal perception data to obtain the target user's emotion state vector; The system acquires the target user's historical music preference data, determines a music preference vector based on the historical music preference data, and then determines the target user's music preference expectation tendency based on the emotional state vector and the music preference vector, thereby obtaining the target user's various music expectation values ​​in the current emotional state. Initialize the music matching threshold, and predict the current music type desired by the target user based on each music expectation value and the music matching threshold to obtain the music response degree of the target user for each music type; A personalized music recommendation list is generated based on the target user's current emotional state, according to each music responsiveness level.

[0006] In this embodiment, the multimodal sensing data includes physiological signals, facial signals, and interaction signals.

[0007] In this embodiment, the process of extracting cross-modal emotions from the multimodal perception data to obtain the target user's emotional state vector specifically includes: Physiological features, facial features, and interaction features are extracted from the multimodal perception data. The physiological features, facial features, and interaction features are fused across modalities to obtain the target user's emotional state vector.

[0008] In this embodiment, determining the music preference vector based on the historical music preference data specifically includes: Based on the historical music preference data, feature vectors are generated to obtain historical music feature vectors; Determine the intensity of each music preference in the historical music feature vector; The music preference vector is determined based on the intensity of each music preference.

[0009] In this embodiment, the expected music preferences of the target user are determined based on the emotional state vector and the music preference vector, thereby obtaining the expected music values ​​of the target user in the current emotional state. Specifically, this includes: The weights of each preference of the target user are determined based on the music preference vector. Based on the emotional state vector, the target user's current emotional state is personalized by converting the music type, and the target user's emotional bias in the current emotional state is obtained. For each emotional bias, the expected tendency is determined by the emotional bias and the corresponding preference weight, and the music expectation value for the corresponding emotional bias is obtained, thereby obtaining the music expectation value for the target user in the current emotional state.

[0010] In this embodiment, the music expectation value is used to quantify the user's acceptance of a specific type of music in the current emotional state.

[0011] In this embodiment, the music type currently desired by the target user is predicted based on each music expectation value and the music matching threshold, and the music response level of the target user for each music type is obtained, specifically including: Initialize the music genre feature library; Type matching is performed based on various music expectation values ​​and music type feature databases to obtain the matching degree of each music type; Matching intensity is filtered based on the matching degree of each music type and the music matching threshold, thereby obtaining each desired music type; Based on the expected music values ​​corresponding to each desired music type, the expected response is calculated to obtain the target user's music response level for each music type.

[0012] In this embodiment, the music responsiveness is used to quantify the intensity of a target user's response to a certain type of music in the current state.

[0013] In this embodiment, generating a personalized music recommendation table for the target user in their current emotional state based on various music responsiveness metrics specifically includes: Determine the music response sequence based on each music responsivity; A personalized music recommendation list is generated based on the music response sequence to reflect the target user's current emotional state.

[0014] Secondly, this application provides a personalized music automatic generation system that integrates multimodal emotion perception, used to execute a personalized music automatic generation method that integrates multimodal emotion perception, the generation system comprising: The state data acquisition module is used to collect multimodal perception data of the target user, and to extract emotions across modalities from the multimodal perception data to obtain the emotional state vector of the target user. The music preference expectation module is used to obtain the target user's historical music preference data, determine the music preference vector based on the historical music preference data, and then determine the expected tendency of the target user's music preference according to the emotional state vector and the music preference vector, thereby obtaining the target user's various music expectation values ​​in the current emotional state. The music type response module is used to initialize the music matching threshold, predict the music type currently desired by the target user based on each music expectation value and the music matching threshold, and obtain the music response degree of the target user for each music type. The personalized music generation module is used to generate a personalized music recommendation list for the target user based on the responsiveness of each music and their current emotional state.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: Multimodal perception data of the target user is collected, and cross-modal emotion extraction is performed on the multimodal perception data to obtain the target user's emotion state vector; historical music preference data of the target user is obtained, and a music preference vector is determined based on the historical music preference data; then, the expected tendency of the target user's music preference is determined according to the emotion state vector and the music preference vector, thereby obtaining the expected values ​​of each music of the target user in the current emotional state; a music matching threshold is initialized, and the music type currently expected by the target user is predicted according to each music expectation value and the music matching threshold to obtain the music response degree of the target user for each music type; a personalized music recommendation table for the target user in the current emotional state is generated according to each music response degree.

[0016] Therefore, this application firstly constructs a precise emotion perception capability by integrating multi-source heterogeneous data of target users, improving the accuracy and reliability of emotion recognition. It integrates multi-source heterogeneous data into an emotion state vector, achieving a quantitative representation of the target user's instantaneous emotions and providing a stable data foundation for subsequent personalized music generation. Secondly, it constructs a long-term preference profile using the target user's historical music preference data, dynamically integrating it with real-time emotional states. Through the interactive calculation of emotion vectors and preference vectors, it can accurately predict the user's music demand tendency under specific emotions, generating music expectation values ​​and realizing an intelligent transformation from liking to needing, providing a core foundation for accurate personalized music recommendations. The decision-making foundation is solidified by music matching thresholds, which enable intelligent filtering of user music expectations. Focusing on user needs, it accurately identifies music types that match the target user's current emotional state and calculates quantified music responsiveness, providing clear and actionable guidance for subsequent personalized music generation. Finally, the quantified music responsiveness is mapped to an executable personalized recommendation strategy, achieving a closed loop from sentiment analysis to music content delivery. This intelligently generates a structured personalized music recommendation table, ensuring that the recommended content highly matches the user's real-time emotional state and long-term preferences, optimizing the efficiency and accuracy of music recommendations, and completing an intelligent service from multimodal emotion perception to personalized music presentation.

[0017] In summary, the technical solution adopted in this application can integrate the user's current emotional state and historical preference data to achieve real-time dynamic adaptation driven by emotion, thereby improving the accuracy and timeliness of personalized music recommendations. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this embodiment of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is an exemplary flowchart of a personalized music automatic generation method that integrates multimodal emotion perception, provided in this application. Figure 2 This is a flowchart illustrating the process of obtaining the expected values ​​of various music tracks for the target user in their current emotional state, as provided in this application. Figure 3 This is a flowchart illustrating the process of obtaining the target user's music response to various music genres, as provided in this application. Figure 4 This is a module structure diagram of a personalized music automatic generation system that integrates multimodal emotion perception, provided in this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0021] This application provides a personalized music automatic generation system and method that integrates multimodal emotion perception. The core of the system involves: collecting multimodal perception data of a target user; extracting cross-modal emotions from the multimodal perception data to obtain the target user's emotion state vector; acquiring the target user's historical music preference data; determining a music preference vector based on the historical music preference data; then, based on the emotion state vector and the music preference vector, assigning expected tendencies to the target user's music preferences to obtain various music expectation values ​​for the target user in the current emotional state; initializing music matching thresholds; predicting the target user's current desired music types based on each music expectation value and the music matching thresholds to obtain the target user's music responsiveness to each music type; and generating a personalized music recommendation table for the target user in the current emotional state based on each music responsiveness.

[0022] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1As shown, this figure is an exemplary flowchart of a personalized music automatic generation method integrating multimodal emotion perception according to this embodiment of the present application. The control method includes the following steps: In step S1, multimodal perception data of the target user is collected, and cross-modal emotion extraction is performed on the multimodal perception data to obtain the emotional state vector of the target user.

[0023] In practical implementation, firstly, multimodal perception data of the target user can be collected. Specifically, the RR interval sequence can be obtained using photoplethysmography and used as a physiological signal. Using a computer vision algorithm and a DlibHOG detector, 68 facial feature points are identified, and the AU intensity of each facial feature point in the Facial Action Coding System (FACS) is calculated. Each AU intensity is used as a facial signal, and the touch position, pressure, and duration of the target user on the music software panel are recorded as interaction signals. Thus, the set of physiological signals, facial signals, and interaction signals is used as the target user's multimodal perception data, which includes physiological signals, facial signals, and interaction signals.

[0024] In this embodiment, the process of extracting cross-modal emotions from the multimodal perception data to obtain the target user's emotional state vector can be achieved through the following steps: Physiological features, facial features, and interaction features are extracted from the multimodal perception data. The physiological features, facial features, and interaction features are fused across modalities to obtain the target user's emotional state vector.

[0025] In practical implementation, physiological features, facial features, and interaction features can be extracted from multimodal sensing data. Specifically, the average heart rate, standard deviation of the RR interval, and root mean square of the difference between adjacent RR intervals can be calculated based on the RR interval sequence. These factors can then be used as physiological features. Core AU intensity values ​​can be extracted as facial features. Interaction signals are analyzed in fixed time windows. Within each time window, the operation frequency (number of touches per unit time), operation force (mean and variance of touch pressure), and operation duration (average duration of a single touch) are calculated. These operation frequency, operation force, and operation duration are then used as interaction features. Finally, physiological features and facial features can be... Cross-modal fusion of interaction features yields the target user's emotional state vector. Specifically, physiological features, facial features, and interaction features are encoded through independent fully connected neural network layers and mapped to the same semantic space to form a set of initial feature representations. Based on a pre-defined multi-head attention fusion network, the feature representations of each modality are used simultaneously as sources of queries, keys, and values. Using attention weights, the features of the three modalities are weighted and summed separately. The results are input into a recurrent neural network layer to model the temporal dynamic changes of emotional states. Finally, through a fully connected output layer, the temporal representation is mapped to an emotional state vector. For example, the emotional state vector includes focus (0 (distracted) to 1 (focused)), energy (0 (fatigued) to 1 (energetic)) and arousal (0 (calm) to 1 (excited)).

[0026] It should be noted that by integrating multi-source heterogeneous data from target users, a precise emotion perception capability was constructed, improving the accuracy and reliability of emotion recognition. By integrating multi-source heterogeneous data into an emotion state vector, a quantitative representation of the instantaneous emotions of target users was achieved, providing a stable data foundation for subsequent personalized music generation.

[0027] In step S2, the historical music preference data of the target user is obtained, a music preference vector is determined based on the historical music preference data, and then the expected tendency of the target user's music preference is determined according to the emotional state vector and the music preference vector, so as to obtain the various music expectation values ​​of the target user in the current emotional state.

[0028] In practical implementation, obtaining the target user's historical music preference data can be achieved by extracting the target user's listening to various music genres, the total duration of each genre, and the total historical listening time from the music software backend. This data is then used as the target user's historical music preference data, which includes the target user's listening to various music genres, the total duration of each genre, and the total historical listening time.

[0029] In this embodiment, determining the music preference vector based on the historical music preference data can be achieved through the following steps: Based on the historical music preference data, feature vectors are generated to obtain historical music feature vectors; Determine the intensity of each music preference in the historical music feature vector; The music preference vector is determined based on the intensity of each music preference.

[0030] In practical implementation, firstly, feature vectorization can be performed based on historical music preference data to obtain historical music feature vectors. That is, the music type and total duration of the type in the historical music preference data can be used as vector elements to form a vector, thus obtaining the historical music feature vector. Then, the intensity of each music preference in the historical music feature vector can be determined. That is, for each music type in the historical music feature vector, the total duration of the corresponding music type is divided by the total historical listening time, and the result is used as the music preference intensity of that music type, thus obtaining the music preference intensity of each music type. Finally, the music preference vector can be determined based on the intensity of each music preference. That is, the intensity of each music preference can be sorted from high to low, and each music preference intensity and the corresponding music type can be used as vector elements to form a vector, thus obtaining the resulting vector as the music preference vector.

[0031] Preferably, in this embodiment, the target user's music preferences are expected based on the emotional state vector and the music preference vector, thereby obtaining the target user's various music expectation values ​​in the current emotional state, with reference to... Figure 2 As shown in the figure, this is a flowchart illustrating the process of obtaining various music expectation values ​​of the target user in the current emotional state in some embodiments of this application. In this embodiment, obtaining various music expectation values ​​of the target user in the current emotional state can be achieved through the following steps: In step S21, the preference weights of the target user are determined based on the music preference vector; In step S22, the target user's current emotional state is personalized by music type based on the emotional state vector to obtain the target user's various emotional biases in the current emotional state. In step S23, for each emotional bias, the expected tendency is determined by the emotional bias and the corresponding preference weight, and the music expectation value for the corresponding emotional bias is obtained, thereby obtaining the music expectation values ​​for the target user in the current emotional state.

[0032] In practical implementation, firstly, the preference weights of the target user can be determined based on the music preference vector. This involves summing the intensities of each music preference in the vector, dividing each intensity by the sum, and using the results as the preference weights for each user. Secondly, the music genre can be personalized based on the target user's current emotional state using the emotional state vector, yielding the user's emotional bias in that state. This can be achieved by pre-setting a standardized mapping matrix that defines the influence of different emotional dimensions on various music genres, such as (classical music, focus coefficient 0.9, energy coefficient 0.6, arousal coefficient 0.4). Here, the focus coefficient refers to the degree to which the music genre supports the user's concentration ability, the energy coefficient refers to the degree of match between the music genre and the user's energy level, and the arousal coefficient is... To determine the degree to which a music genre adapts to user arousal, for each music genre, each element in the emotional state vector is multiplied by the corresponding coefficient for that music genre, and the results are summed. The sum of these multiplications is used as the emotional bias of that music genre. Finally, for each emotional bias, the expected tendency is calculated by multiplying the emotional bias by the corresponding preference weight, resulting in the music expectation value for that emotional bias. This yields the music expectation values ​​for the target user in their current emotional state. In other words, for each emotional bias, the emotional bias is multiplied by the corresponding preference weight, and the result is used as the music expectation value for that emotional bias. It should be noted that the music expectation value is used to quantitatively represent the user's acceptance of a specific type of music in their current emotional state.

[0033] It should be noted that by constructing a long-term preference profile based on the target user's historical music preference data and dynamically integrating it with the real-time emotional state, the system can accurately predict the user's music demand tendency under specific emotions through interactive calculation of emotion vectors and preference vectors. This generates music expectation values ​​and realizes the intelligent transformation from liking to need, providing a core decision-making basis for accurate personalized music recommendations.

[0034] In step S3, the music matching threshold is initialized, and the music type currently desired by the target user is predicted based on each music expectation value and the music matching threshold, so as to obtain the music response degree of the target user for each music type.

[0035] In the specific implementation, the music matching threshold is initialized. The music matching threshold can be preset based on historical data statistics. The music matching threshold is a boundary value used to divide the intensity of music matching.

[0036] Preferably, in this embodiment, the target user's current desired music type is predicted based on each music expectation value and music matching threshold to obtain the target user's music response level for each music type, and then referenced. Figure 3 As shown in the figure, this is a flowchart illustrating the process of obtaining the target user's music response to various music genres in some embodiments of this application. In this embodiment, obtaining the target user's music response to various music genres can be achieved through the following steps: In step S31, the music type feature library is initialized; In step S32, type matching is performed based on each music expectation value and music type feature library to obtain the matching degree of each music type; In step S33, matching intensity is filtered based on the matching degree of each music type and the music matching threshold to obtain each desired music type; In step S34, the expected response is performed based on the music expectation value corresponding to each expected music type to obtain the target user's music response degree for each music type.

[0037] In practical implementation, firstly, a music genre feature library can be initialized. This library can be pre-defined based on experimental data analysis, and includes focus fit, energy fit, and arousal fit for each music genre. Then, genre matching can be performed between the expected values ​​of each music genre and the feature library to obtain the matching degree for each genre. Specifically, the emotional state vector can be multiplied by the corresponding features in the music genre feature library, the results summed, and then multiplied by the expected values ​​of each music genre. The resulting sum is used as the matching degree for each music genre. Furthermore, matching strength can be filtered based on the matching degree and a music matching threshold to obtain the matching degree for each genre. The desired music type is determined by comparing the matching degree of each music type with a music matching threshold. All music types with matching degrees exceeding the threshold are extracted, and the resulting music types are used as the desired music types. Finally, the desired response is calculated based on the expected music value corresponding to each desired music type, yielding the target user's music response level for each music type. This is achieved by summing the expected music values ​​for all desired music types, then dividing each expected music value by the summation, and using the result as the target user's music response level for each music type. It should be noted that music response level quantifies the intensity of the target user's reaction to a particular type of music in the current state.

[0038] It should be noted that by using music matching thresholds, the system can intelligently filter users' music expectations, focusing on user needs. It can accurately identify the music type that matches the target user's current emotional state and calculate a quantitative music response, providing clear and actionable guidance for subsequent personalized music generation.

[0039] In step S4, a personalized music recommendation table is generated based on each music responsiveness to determine the target user's current emotional state.

[0040] In this embodiment, generating a personalized music recommendation list for the target user in their current emotional state based on each music responsiveness can be achieved through the following steps: Determine the music response sequence based on each music responsivity; A personalized music recommendation list is generated based on the music response sequence to reflect the target user's current emotional state.

[0041] In practical implementation, firstly, a music response sequence can be determined based on each music response level. That is, the music response levels can be sorted from largest to smallest, and the resulting sequence is used as the music response sequence. Then, a personalized music recommendation table for the target user in the current emotional state can be generated based on the music response sequence. That is, the music type can be extracted sequentially from the music response sequence, and the music type corresponding to each music response level in the music response sequence can be extracted and merged. For example, if multiple music response levels correspond to the same music type, they can be merged to obtain a music type sequence. Then, each music type in the music type sequence is used as the selection criterion for recommended music. The music names corresponding to the music type sequence are extracted from the backend database, and the extracted music names are used to generate a table. The resulting table serves as the personalized music recommendation table for the target user in the current emotional state. The personalized music recommendation table includes: recommended music name, corresponding music matching degree, and corresponding music type.

[0042] It should be noted that by mapping quantified music responsiveness to an executable personalized recommendation strategy, a closed loop from sentiment analysis to music content delivery is achieved. This intelligently generates a structured personalized music recommendation table, ensuring that the recommended content highly matches the user's real-time emotional state and long-term preferences. This optimizes the efficiency and accuracy of music recommendations, completing an intelligent service from multimodal emotion perception to personalized music presentation.

[0043] Therefore, this application firstly constructs a precise emotion perception capability by integrating multi-source heterogeneous data of target users, improving the accuracy and reliability of emotion recognition. It integrates multi-source heterogeneous data into an emotion state vector, achieving a quantitative representation of the target user's instantaneous emotions and providing a stable data foundation for subsequent personalized music generation. Secondly, it constructs a long-term preference profile using the target user's historical music preference data, dynamically integrating it with real-time emotional states. Through the interactive calculation of emotion vectors and preference vectors, it can accurately predict the user's music demand tendency under specific emotions, generating music expectation values ​​and realizing an intelligent transformation from liking to needing, providing a core foundation for accurate personalized music recommendations. The decision-making foundation is solidified by music matching thresholds, which enable intelligent filtering of user music expectations. Focusing on user needs, it accurately identifies music types that match the target user's current emotional state and calculates quantified music responsiveness, providing clear and actionable guidance for subsequent personalized music generation. Finally, the quantified music responsiveness is mapped to an executable personalized recommendation strategy, achieving a closed loop from sentiment analysis to music content delivery. This intelligently generates a structured personalized music recommendation table, ensuring that the recommended content highly matches the user's real-time emotional state and long-term preferences, optimizing the efficiency and accuracy of music recommendations, and completing an intelligent service from multimodal emotion perception to personalized music presentation.

[0044] In summary, the technical solution adopted in this application can integrate the user's current emotional state and historical preference data to achieve real-time dynamic adaptation driven by emotion, thereby improving the accuracy and timeliness of personalized music recommendations.

[0045] Example 2: This application provides a personalized music automatic generation system that integrates multimodal emotion perception, referencing... Figure 4 As shown in the figure, this is a module structure diagram of a personalized music automatic generation system integrating multimodal emotion perception according to this embodiment of the present application. The generation system includes: The state data acquisition module 100 is used to acquire multimodal perception data of the target user, and to perform cross-modal emotion extraction on the multimodal perception data to obtain the emotional state vector of the target user. The music preference expectation module 200 is used to acquire the target user's historical music preference data, determine the music preference vector based on the historical music preference data, and then determine the target user's music preference tendency according to the emotional state vector and the music preference vector, thereby obtaining the target user's various music expectation values ​​in the current emotional state. The music type response module 300 is used to initialize the music matching threshold, predict the music type currently desired by the target user based on each music expectation value and the music matching threshold, and obtain the music response degree of the target user for each music type. The personalized music generation module 400 is used to generate a personalized music recommendation list for the target user based on the responsiveness of each music, according to the user's current emotional state.

[0046] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0047] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0048] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A personalized music automatic generation method integrating multimodal emotion perception, characterized in that, The control method includes the following steps: Collect multimodal perception data of the target user, and perform cross-modal emotion extraction on the multimodal perception data to obtain the target user's emotion state vector; The system acquires the target user's historical music preference data, determines a music preference vector based on the historical music preference data, and then determines the target user's music preference expectation tendency based on the emotional state vector and the music preference vector, thereby obtaining the target user's various music expectation values ​​in the current emotional state. Initialize the music matching threshold, and predict the current music type desired by the target user based on each music expectation value and the music matching threshold to obtain the music response degree of the target user for each music type; A personalized music recommendation list is generated based on the target user's current emotional state, according to each music responsiveness level.

2. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, The multimodal sensing data includes physiological signals, facial signals, and interaction signals.

3. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, The process of extracting cross-modal emotions from the multimodal perception data to obtain the target user's emotional state vector specifically includes: Physiological features, facial features, and interaction features are extracted from the multimodal perception data. The physiological features, facial features, and interaction features are fused across modalities to obtain the target user's emotional state vector.

4. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, Determining the music preference vector based on the historical music preference data specifically includes: Based on the historical music preference data, feature vectors are generated to obtain historical music feature vectors; Determine the intensity of each music preference in the historical music feature vector; The music preference vector is determined based on the intensity of each music preference.

5. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, Based on the emotional state vector and the music preference vector, the expected music preferences of the target user are determined, thereby obtaining the various music expectation values ​​of the target user in the current emotional state, specifically including: The weights of each preference of the target user are determined based on the music preference vector. Based on the emotional state vector, the target user's current emotional state is personalized by converting the music type, and the target user's emotional bias in the current emotional state is obtained. For each emotional bias, the expected tendency is determined by the emotional bias and the corresponding preference weight, and the music expectation value for the corresponding emotional bias is obtained, thereby obtaining the music expectation value for the target user in the current emotional state.

6. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, The music expectation value is used to quantify the degree to which a user is receptive to a specific type of music in their current emotional state.

7. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, Based on the music expectation values ​​and the music matching threshold, the target user's current desired music type is predicted, and the target user's music response to each music type is obtained, specifically including: Initialize the music genre feature library; Type matching is performed based on various music expectation values ​​and music type feature databases to obtain the matching degree of each music type; Matching intensity is filtered based on the matching degree of each music type and the music matching threshold, thereby obtaining each desired music type; Based on the expected music values ​​corresponding to each desired music type, the expected response is calculated to obtain the target user's music response level for each music type.

8. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, The music responsiveness is used to quantify the intensity of a target user's response to a certain type of music in the current state.

9. The personalized music automatic generation method integrating multimodal emotion perception as described in claim 1, characterized in that, Based on various music responsiveness metrics, a personalized music recommendation list is generated for the target user in their current emotional state, specifically including: Determine the music response sequence based on each music responsivity; A personalized music recommendation list is generated based on the music response sequence to reflect the target user's current emotional state.

10. A personalized music automatic generation system integrating multimodal emotion perception, used to execute the personalized music automatic generation method integrating multimodal emotion perception as described in any one of claims 1 to 9, characterized in that, The generation system includes: The state data acquisition module is used to collect multimodal perception data of the target user, and to extract emotions across modalities from the multimodal perception data to obtain the emotional state vector of the target user. The music preference expectation module is used to obtain the target user's historical music preference data, determine the music preference vector based on the historical music preference data, and then determine the expected tendency of the target user's music preference according to the emotional state vector and the music preference vector, thereby obtaining the target user's various music expectation values ​​in the current emotional state. The music type response module is used to initialize the music matching threshold, predict the music type currently desired by the target user based on each music expectation value and the music matching threshold, and obtain the music response degree of the target user for each music type. The personalized music generation module is used to generate a personalized music recommendation list for the target user based on the responsiveness of each music and their current emotional state.