Intelligent music playing method, system and device adapting to emotion and medium

By combining deep learning models with breathing, pulse and sound data, identifying the user's emotional state and recommending music based on the emotion recognition results, the problem that the existing technology cannot adapt to user's emotions in real time is solved, and high-precision emotion recognition and personalized music recommendation are achieved.

CN120045738APending Publication Date: 2025-05-27FOSHAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510097622.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing smart music playback technology cannot fully consider the user's real-time emotional changes, resulting in the inability to accurately meet the user's current emotional needs, and the emotional recognition methods are relatively single, with low accuracy and reliability.

Method used

By obtaining the user's breathing data, pulse data, sound data and music preference data, performing feature processing, using deep learning models for emotion recognition, and establishing a music track playlist based on the recognition results and music preference data, realizing intelligent music playback that adapts to emotions.

Benefits of technology

It improves the accuracy of user emotions recognition, can more comprehensively capture users' real-time emotional changes, provide a highly personalized music experience, and help users regulate their emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045738A_ABST
    Figure CN120045738A_ABST
Patent Text Reader

Abstract

The invention provides an emotion-adaptive intelligent music playing method, system and device and a medium, and relates to the technical field of intelligent music playing, and the method comprises the steps: obtaining breath data, pulse data, sound data and music preference data of a user; performing feature processing on the respiration data, the pulse data and the sound data to obtain respiration features, pulse features and sound features; according to the breathing features, the pulse features and the sound features, utilizing a deep learning model to obtain an emotion recognition result of the user; establishing a music track playing list according to the emotion recognition result and the music preference data; and playing music according to the music track playlist. According to the method, the multi-modal data and the deep learning model are utilized, the recognition accuracy of the user emotion is improved, and the corresponding music is played according to the recognition result of the user emotion, so that the user is effectively helped to perform emotion adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent music playing, and in particular to an intelligent music playing method, system, device and medium adapted to emotions. Background Art

[0002] Intelligent music playing technology can effectively help users regulate their emotions by identifying the user's emotional state and providing corresponding music. Most of the existing intelligent music playing technologies recommend based on the user's historical play records and music type preferences, and fail to fully consider the user's real-time emotional changes, and cannot accurately meet the user's current emotional needs. In addition, the existing emotional recognition means are relatively single, resulting in low accuracy and reliability of emotional recognition. Summary of the Invention

[0003] This application provides an intelligent music playing method, system, device and medium adapted to emotions to solve one or more technical problems existing in the prior art, and at least provide a beneficial choice or create conditions.

[0004] On the one hand, this application provides an intelligent music playing method adapted to emotions, including the following steps: Obtain the user's breathing data, pulse data, voice data and music preference data; Perform feature processing on the breathing data, the pulse data and the voice data to obtain breathing features, pulse features and voice features; According to the breathing features, the pulse features and the voice features, use a deep learning model to obtain the user's emotional recognition result; Establish a music track playlist according to the emotional recognition result and the music preference data; play music according to the music track playlist.

[0005] Further, the performing feature processing on the breathing data, the pulse data and the voice data to obtain breathing features, pulse features and voice features includes: Perform preprocessing on the breathing data to obtain breathing frequency features, breathing depth features and breathing rhythm features; Perform feature fusion processing on the breathing frequency features, the breathing depth features and the breathing rhythm features to obtain the breathing features; Perform preprocessing on the pulse data to obtain heart rate features and heart rate variability features; Perform feature fusion processing on the heart rate features and the heart rate variability features to obtain the pulse features; Perform voice activity detection processing and endpoint localization processing on the voice data to obtain voice segment data; Preprocess the speech segment data to obtain pitch features and loudness features; Perform feature fusion processing on the pitch features and the loudness features to obtain the voice features.

[0006] Further, the deep learning includes a breathing data model, a pulse data model, and a voice data model; Based on the breathing features, the pulse features, and the voice features, using a deep learning model, obtain the emotion recognition result of the user, including: Based on the breathing features, combined with the breathing data model, obtain a breathing emotion probability vector; Based on the pulse features, combined with the pulse data model, obtain a pulse emotion probability vector; Based on the voice features, combined with the voice data model, obtain a voice emotion probability vector; Perform fusion processing on the breathing emotion probability vector, the pulse emotion probability vector, and the voice emotion probability vector to obtain the emotion recognition result of the user.

[0007] Further, based on the emotion recognition result and the music preference data, establish a music track playlist, including: Based on the music preference data, establish an emotion music library for the user; Based on the emotion recognition result, screen out matching music tracks from the emotion music library and establish the music track playlist.

[0008] Further, based on the music preference data, establish the emotion music library for the user, including: Based on the music preference data, obtain the user's favorite tracks; Perform emotion characteristic classification and annotation on the favorite tracks to obtain emotion labels corresponding to the favorite tracks; Based on the favorite tracks and their corresponding emotion labels, establish the emotion music library for the user.

[0009] Further, based on the emotion recognition result, screen out matching music tracks from the emotion music library and establish the music track playlist, including: Based on the emotion recognition result, screen out matching music tracks from the emotion music library to obtain a set of music tracks suitable for the current emotion; Based on the set of music tracks suitable for the current emotion, combined with a preset sorting rule, establish the music track playlist.

[0010] Furthermore, the respiration data model includes a long short-term memory network model; the pulse data model includes a convolutional neural network model; and the voice data model includes a deep belief network model.

[0011] On the other hand, the present application provides an emotion-adaptive intelligent music playback system, including: a data acquisition module, a feature processing module, an emotion recognition module, and a music playback module; The data acquisition module is configured to acquire the user's respiration data, pulse data, voice data, and music preference data; The feature processing module is configured to perform feature processing on the respiration data, the pulse data, and the voice data to obtain respiration features, pulse features, and voice features; The emotion recognition module is configured to obtain an emotion recognition result of the user by using a deep learning model according to the respiration features, the pulse features, and the voice features; The music playback module is configured to establish a music track playback list according to the emotion recognition result and the music preference data; and play music according to the music track playback list.

[0012] On the other hand, the present application provides an emotion-adaptive intelligent music playback device, including: a processor and a memory; the memory is configured to store a program; when the program is executed by the processor, the processor implements the foregoing emotion-adaptive intelligent music playback method.

[0013] On the other hand, the present application provides a computer-readable storage medium, in which a processor-executable program is stored, and the processor-executable program is used to implement the foregoing emotion-adaptive intelligent music playback method when executed by the processor.

[0014] The beneficial effects of the present application are as follows: The present application provides an emotion-adaptive intelligent music playback method, including: acquiring the user's respiration data, pulse data, voice data, and music preference data; performing feature processing on the respiration data, the pulse data, and the voice data to obtain respiration features, pulse features, and voice features; obtaining an emotion recognition result of the user by using a deep learning model according to the respiration features, the pulse features, and the voice features; establishing a music track playback list according to the emotion recognition result and the music preference data; and playing music according to the music track playback list. The present application improves the accuracy of user emotion recognition by using multi-modal data and a deep learning model, and plays corresponding music according to the emotion recognition result of the user, thereby effectively helping the user to regulate emotions. The present application also provides corresponding systems, devices, and media. The beneficial effects of the systems, devices, and media are similar to those of the method and will not be repeated here.

[0015] Other features and advantages of the present application will be described in the subsequent specification, and in part will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the specification, claims and drawings. Description of the Drawings

[0016] The drawings are used to provide a further understanding of the technical solutions of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solutions of the present invention, and do not constitute a limitation to the technical solutions of the present invention.

[0017] Figure 1 is a flowchart of the intelligent music playing method for adapting to emotions provided by the present application; Figure 2 is a schematic diagram of determining respiratory characteristics, pulse characteristics and voice characteristics provided by the present application; Figure 3 is a schematic diagram of obtaining the emotion recognition result of the user by using a deep learning model according to the respiratory characteristics, pulse characteristics and voice characteristics provided by the present application; Figure 4 is a schematic diagram of establishing a music track playlist according to the emotion recognition result and music preference data provided by the present application; Figure 5 is a structural diagram of the intelligent music playing system for adapting to emotions provided by the present application; Figure 6 is a structural diagram of the intelligent music playing device for adapting to emotions provided by the present application. Detailed Embodiments

[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0019] The present application will be further described below with reference to the drawings in the specification and specific embodiments. The described embodiments should not be regarded as a limitation to the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0020] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0022] Most existing intelligent music playback technologies recommend songs based on users' historical play records and music genre preferences, ignoring users' real-time mood changes and unable to provide a truly mood-fitting music experience. In terms of emotion recognition, although existing technologies have tried to use physiological data for emotion monitoring, most rely on single-modal data sources and are difficult to comprehensively and accurately reflect complex emotional states. For example, relying solely on a certain type of physiological signal may not be able to capture the subtle differences in emotions, resulting in inaccurate recognition results.

[0023] Therefore, integrating multi-modal data for emotion recognition is the key direction to improve accuracy and reliability. However, existing multi-modal emotion recognition technologies still need to be optimized in aspects such as data fusion, model selection, and processing flow. Specifically, the existing technologies have the following deficiencies: First, the emotion targeting is poor. Most systems recommend based on historical data and are difficult to meet the emotional needs in specific situations. Second, there is a lack of active adaptability. Users still need to manually search and select music, increasing the usage difficulty and psychological burden. Third, the emotion monitoring means are single, such as questionnaires or user-set tags, which are not precise enough and are easily affected by subjective factors. Fourth, the personalization depth is insufficient. It fails to provide a deep-level personalized music adjustment plan according to the complex and changeable emotional states of users. Fifth, the data information is limited. The limitations of single-modal data affect the overall accuracy of emotion recognition.

[0024] In view of the problems and defects existing in the related technologies, the embodiments of this application use multi-modal data (such as respiration, pulse, and voice data), combined with deep learning models, to achieve accurate recognition of users' emotional states, effectively improving the accuracy of emotion recognition and being able to capture users' real-time emotion changes more comprehensively. The system automatically creates a matching playlist according to the emotion recognition results and the users' music preferences, providing a highly personalized music experience that not only meets the users' current emotional needs but also actively adjusts their psychological states.

[0025] First, the implementation steps of the emotion-adaptive intelligent music playback method provided by the embodiments of this application will be elaborated in detail with reference to the accompanying drawings.

[0026] The emotion-adaptive intelligent music playback method proposed in the embodiments of this application can be applied to a terminal, a server, or software running on a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.

[0027] Referring to Figure 1 , the implementation process of the emotion-adaptive intelligent music playback method provided in the embodiments of this application includes but is not limited to the following steps.

[0028] Step 101, obtain the user's breathing data, pulse data, voice data, and music preference data.

[0029] It should be noted that the breathing data reflects the emotional state. For example, when the user feels anxious or nervous, the breathing rate will increase; when the user feels relaxed or calm, the breathing rate will decrease. Breathing data is a non-invasive and easily collectible physiological indicator that can monitor the user's emotional changes in real time and is suitable for long-term emotional monitoring. Pulse data can reflect the user's heart rate changes, and the heart rate changes are closely related to the emotional state. For example, when the user feels nervous or excited, the heart rate will increase; when the user feels relaxed or calm, the heart rate will decrease. Pulse data is relatively stable and not easily affected by the external environment, and can provide reliable physiological indicators, which is suitable for emotion recognition. Voice data can reflect the user's emotional expression, including pitch, loudness, etc. These characteristics are closely related to the emotional state. For example, when the user feels sad, the pitch may decrease; when the user feels angry, the pitch may increase. Voice data can capture the user's immediate reaction and is suitable for emotion recognition in various environments, such as at home, in the office, outdoors, etc.

[0030] Step 101, collect the user's breathing data, pulse data, voice data, and music preference data. These data are the basis for subsequent emotion recognition and music recommendation, and can comprehensively reflect the user's current emotional state and music preferences.

[0031] Optionally, collect the user's breathing data, pulse data, and voice data through wearable devices, such as smart bracelets and smart watches.

[0032] Optionally, guide the user to select a favorite music genre at the beginning as music preference data, such as rap, pop, rock, classical, folk, jazz, etc. The music genre also plays an important role in regulating the user's mood. Based on the user's favorite music genre, it is more conducive to providing the music the user needs, which is extremely helpful for regulating the user's mood.

[0033] Step 102: Perform feature processing on the breathing data, pulse data, and voice data to obtain breathing features, pulse features, and voice features.

[0034] In step 102, perform feature processing on the breathing data, pulse data, and voice data to obtain breathing features, pulse features, and voice features, ensuring the quality and consistency of the data and eliminating the influence of noise and outliers. By extracting breathing features, pulse features, and voice features, high-quality input data is provided for subsequent emotion recognition.

[0035] Step 103: Based on the breathing features, pulse features, and voice features, use a deep learning model to obtain the emotion recognition result of the user.

[0036] It should be noted that using a deep learning model to obtain the emotion recognition result of the user is mainly because the deep learning model has advantages such as high-precision recognition, multi-modal data fusion, automatic learning, and strong adaptability. The deep learning model can automatically extract complex features from the original data, process non-linear relationships, and improve the accuracy of emotion recognition. By integrating multiple modal data such as breathing, pulse, and voice, the limitations of single data are overcome, and the robustness of recognition is improved. In addition, the deep learning model can process data in real time, adapt to different users and situations, and provide personalized services. Generally speaking, the deep learning model performs excellently in emotion recognition, can achieve real-time and high-precision emotion monitoring, and provide more effective personalized music recommendations and emotion regulation services for users.

[0037] In step 103, use a deep learning model to analyze the breathing features, pulse features, and voice features to achieve high-precision emotion recognition. The use of multi-modal data improves the accuracy and robustness of emotion recognition and can more comprehensively reflect the user's emotional state.

[0038] Step 104: Establish a music track playlist according to the emotion recognition result and music preference data.

[0039] In step 104, a music track playlist is established based on the emotion recognition result and the user's music preference data. This not only takes into account the user's current emotional needs but also respects the user's long-term music taste, ensuring that the recommended music not only matches the current mood but also does not deviate from the user's personal preferences. Based on the emotion recognition result, the system can select the music that is most suitable for helping the user adjust their mood. For example, when it is detected that the user is in a state of stress or anxiety, the system can select soothing and relaxing music; while when the user is feeling frustrated or low, it can select inspiring and morale-boosting tracks. This precise matching helps to more effectively improve the user's emotional state. Moreover, as the user's mood changes, the system can continuously update the playlist to ensure that the music always remains in sync with the user's emotions. This means that even within the same time period, if the user's mood changes, the system can respond in a timely manner and provide corresponding music to support the calming or transformation of the mood.

[0040] Step 105, play music according to the music track playlist.

[0041] In step 105, after accurately identifying the user's emotion and establishing a personalized playlist, directly playing the adapted music can immediately provide emotional support to the user. This immediate response can quickly help the user adjust their mood, for example, from anxiety to relaxation, or from low to positive, thus effectively improving the psychological state. The automatic play function avoids the process of the user having to manually select music, especially when the user is in a bad mood or unwilling to perform additional operations, reducing the burden on the user. The system actively provides music, making the whole process smoother and more natural, enhancing the coherence and comfort of the user experience.

[0042] In some embodiments of the present application, with reference to Figure 2 , in step 102, the implementation process of performing feature processing on the respiration data, pulse data, and voice data to obtain respiration features, pulse features, and voice features includes but is not limited to the following steps.

[0043] Step 201, preprocess the respiration data to obtain respiration frequency features, respiration depth features, and respiration rhythm features.

[0044] It should be noted that the respiration frequency, as an indicator of the number of breaths per unit time, directly reflects the activity speed of the respiratory system and is an important parameter for evaluating the emotional state. In a state of emotional excitement, anxiety, or stress, the respiration frequency often increases significantly, reflecting the body's immediate response to an emergency. On the contrary, in a calm or relaxed state, the respiration frequency tends to be stable and decrease.

[0045] Respiratory depth, which is the volume of air inhaled or exhaled in each breath, not only measures the amount of breathing but also reflects the psychological state. Deep breathing is usually associated with relaxation and calmness because it promotes more oxygen intake and carbon dioxide exhalation, helping to relieve stress and tension. A shallow and rapid breathing pattern often occurs during moments of anxiety or panic, when the respiratory depth decreases and the breathing efficiency is reduced.

[0046] Respiratory rhythm describes the time distribution of the breathing cycle, including the proportional relationship of the inhalation, breath-holding, and exhalation times, and is one of the key factors in emotion recognition. A regular and slow respiratory rhythm is usually associated with relaxation and concentration, while an irregular or rapid respiratory rhythm may indicate anxiety, tension, or other negative emotions.

[0047] In step 201, by preprocessing the respiratory data, respiratory frequency features, respiratory depth features, and respiratory rhythm features are obtained, and these features together provide a comprehensive description of the user's respiratory state.

[0048] In step 202, feature fusion processing is performed on the respiratory frequency features, respiratory depth features, and respiratory rhythm features to obtain respiratory features.

[0049] In step 202, the respiratory frequency, depth, and rhythm features are fused through methods such as weighted average or principal component analysis to form a comprehensive respiratory feature vector. This process not only simplifies the feature dimension but also improves the accuracy and reliability of emotion recognition, providing a more comprehensive assessment of the respiratory state.

[0050] In some embodiments of the present application, the implementation process of using the median filtering method to preprocess the respiratory data to obtain respiratory frequency features, respiratory depth features, and respiratory rhythm features includes but is not limited to the following steps.

[0051] First, for the respiratory data sequence , set the median filtering window size to , , usually an odd number, and use the median filtering window to filter the respiratory data sequence to obtain the filtered respiratory sequence . The -th data point in the filtered respiratory sequence satisfies the following formula (1): (1); In formula (1), represents the operation of taking the median value. Specifically, sort all the values within the corresponding median filtering window of , and find the median value, and take the median value as the -corresponding filtered data point .

[0052] Secondly, perform max - min normalization on the filtered respiration sequence to obtain a normalized respiration sequence , The th data point in satisfies the following formula (2): In formula (2), represents the minimum value of the data points in the filtered respiration sequence , represents the maximum value of the data points in the filtered respiration sequence , is a very small constant used to avoid division - by - zero errors.

[0053] Finally, extract respiration frequency features, respiration depth features, and respiration rhythm features from the normalized respiration sequence .

[0054] In some embodiments of the present application, a peak detection algorithm is used to extract respiration frequency features from the normalized respiration sequence. Specifically, the peak detection algorithm is used to identify the peaks in the inhalation phase or the valleys in the exhalation phase. The respiratory cycle can be obtained by calculating the time interval between adjacent peaks (or valleys). For each respiratory cycle, calculate its duration, and take the reciprocal of the duration as the respiration frequency. Calculate the respiration frequency of each respiratory cycle, and finally form a respiration frequency sequence as the respiration frequency feature.

[0055] In some embodiments of the present application, an amplitude measurement method is used to extract respiration depth features from the normalized respiration sequence. Specifically, calculate the amount of air inhaled or exhaled during each respiratory cycle, that is, the change from the lowest point to the highest point as the respiration depth value. Calculate the respiration depth value of each respiratory cycle, and finally form a respiration depth sequence as the respiration depth feature.

[0056] In some embodiments of the present application, a phase synchrony method is used to extract respiration rhythm features from the normalized respiration sequence. Specifically, calculate the inhalation time and exhalation time of each respiratory cycle, and calculate their ratio. According to the ratio difference of the inhalation time and exhalation time between adjacent respiratory cycles, quantify the consistency or variability of the respiration rhythm, thereby obtaining the respiration rhythm feature.

[0057] Step 203: Pre - process the pulse data to obtain heart rate features and heart rate variability features.

[0058] It should be noted that the pulse data includes the pulse wave signal. The pulse wave signal is the pressure fluctuation propagating in the arterial system during each heartbeat of the heart. It has a periodicity synchronized with the heartbeat, contains complex waveforms such as the ascending branch and peak, and carries cardiovascular health information such as vascular elasticity and resistance. It is vulnerable to the influence of movement and respiration, so appropriate signal processing is required to extract useful information.

[0059] The heart rate feature is used to measure the number of heartbeats per minute and is an important physiological indicator reflecting the body's immediate response and emotional state. During emotional fluctuations, such as experiencing anxiety, anger, or excitement, the heart rate often rises rapidly, while in a relaxed or calm state, the heart rate tends to be stable and at a lower level.

[0060] The heart rate variability feature is used to characterize the variation of the heartbeat interval time and is also an important physiological indicator reflecting the body's immediate response and emotional state. High heart rate variability usually means good psychological resilience and stable emotional management ability, and the individual is in a relaxed state. On the contrary, low heart rate variability may mean long-term stress or negative emotions.

[0061] In step 203, the pulse data is preprocessed to obtain the heart rate feature and the heart rate variability feature, which together provide a comprehensive description of the user's pulse state.

[0062] In step 204, the heart rate feature and the heart rate variability feature are subjected to feature fusion processing to obtain the pulse feature.

[0063] In step 204, the heart rate feature and the heart rate variability feature are fused by methods such as weighted average or principal component analysis to form a comprehensive pulse feature vector. This process not only simplifies the feature dimension but also improves the accuracy and reliability of emotion recognition, providing a more comprehensive pulse state assessment.

[0064] In some embodiments of the present application, the implementation process of preprocessing the pulse wave signal data to obtain the heart rate feature and the heart rate variability feature includes but is not limited to the following steps.

[0065] First, wavelet transform is used to denoise the pulse wave signal and remove interferences such as baseline drift in the pulse wave signal. Specifically, given the pulse wave signal data sequence , assume the pulse wave signal data sequence , at time , the corresponding pulse wave signal is , the wavelet function is , and assume the decomposition level is . After wavelet decomposition of the pulse wave signal , wavelet transform coefficients at different scales and displacements are obtained.

[0066] The scale parameter is used to represent the degree of compression or stretching of the wavelet function. A larger scale parameter corresponds to a wider wavelet function, which can capture low-frequency components; a smaller scale parameter corresponds to a narrower wavelet function, which can capture high-frequency components. The displacement parameter is used to represent the position of the wavelet function along the time axis. The wavelet function can slide within the entire signal range to capture features at different time points. The wavelet transform coefficients are local feature descriptors of the pulse wave signal at different scales and positions. By analyzing and processing these coefficients, effective denoising and feature extraction of the pulse wave signal can be achieved.

[0067] According to the pulse wave signal and the wavelet function , the wavelet transform coefficients based on the scale and displacement are obtained, which satisfy the following formula (3): (3); In formula (3), represents the wavelet function based on the scale and displacement , represents 's conjugate function, represents the differential of the time variable .

[0068] Then, the high-frequency wavelet transform coefficients are set to zero, that is, the high-frequency noise is removed, and the denoised wavelet transform coefficients are obtained. By reconstructing the signal through the inverse wavelet transform, the pulse wave signal after removing the baseline drift and high-frequency noise is obtained, which satisfies the following formula (4): (4); Finally, the pulse wave signal data sequence after removing the baseline drift and high-frequency noise is obtained through the above steps. According to the aforementioned formula (2), is subjected to min-max normalization processing to obtain the normalized pulse wave signal data sequence . The heart rate feature and heart rate variability feature are extracted from the normalized pulse wave signal data sequence .

[0069] In some embodiments of the present application, the implementation process of extracting the heart rate feature from the normalized pulse wave signal data sequence includes: using the adaptive threshold method to locate each heartbeat cycle, calculating the time interval between adjacent heartbeat cycles, converting these time intervals into heart rate values, and finally forming a heart rate value sequence as the heart rate feature.

[0070] In some embodiments of the present application, the implementation process of extracting heart rate variability features from the normalized pulse wave signal data sequence includes: calculating the standard deviation and root mean square deviation between all adjacent heartbeat cycles as heart rate variability features, which comprehensively consider the long-term variability and short-term variability of the heart rate.

[0071] Step 205: Perform voice activity detection processing and endpoint localization processing on the voice data to obtain voice segment data.

[0072] It should be noted that the main purpose of voice activity detection is to distinguish voice segments from non-voice segments and remove background noise and silent segments. The main purpose of endpoint localization is to accurately locate the start point and end point of the voice segment. Through the cross-correlation method or other methods, the boundaries of the voice segment can be accurately determined, avoiding misjudgment and missed detection, ensuring the accuracy of subsequent processing. Especially in multi-modal data fusion and emotion recognition, accurate voice segment data can provide more reliable input, thereby improving the accuracy and robustness of emotion recognition.

[0073] In step 205, perform voice activity detection processing and endpoint localization processing on the voice data to distinguish voice and non-voice paragraphs and remove background noise. At the same time, accurately mark the start and end positions of the voice segment, extract effective voice data, and ensure the accuracy of subsequent feature extraction. Moreover, only process the detected voice segments, avoiding full-scale analysis of the entire audio stream, saving computational resources and time.

[0074] Step 206: Preprocess the voice segment data to obtain pitch features and loudness features.

[0075] It should be noted that pitch features are used to measure the high and low of sound. It not only reflects the fundamental frequency of speech but also deeply affects emotional expression. In the field of emotion recognition, pitch features can reveal the emotional state of the speaker. For example, a high pitch may indicate tension, excitement, or anger, while a low pitch is often associated with calmness, confidence, or sadness. In addition, the changing pattern of pitch, such as a sudden increase or decrease, may also indicate a drastic fluctuation in emotion.

[0076] Loudness features are used to characterize the intensity or volume of sound and are another key emotion recognition indicator. Changes in loudness can reflect the internal emotional state of the speaker. In strong emotional experiences, such as anger, fear, or surprise, people tend to unconsciously increase the loudness of their voices; while in a calm or depressed state, the voice may become softer.

[0077] Step 207: Perform feature fusion processing on the pitch features and loudness features to obtain sound features.

[0078] In step 207, the pitch feature and the loudness feature are fused by methods such as weighted average or principal component analysis to form a comprehensive voice feature vector. This process not only simplifies the feature dimension, but also improves the accuracy and reliability of emotion recognition, providing a more comprehensive assessment of the voice state.

[0079] In some embodiments of the present application, a double-threshold method based on short-time energy and short-time zero-crossing rate is used to perform voice activity detection processing and endpoint localization processing on the voice data to obtain voice segment data.

[0080] It should be noted that the energy of the voice signal is usually higher than the background noise. Therefore, an energy threshold can be set, and when the energy of the signal exceeds this threshold, voice activity is considered to exist. Short-time energy is an index for measuring the intensity of an audio signal, usually obtained by dividing the audio signal into multiple short frames and calculating the sum of the energy within each frame. The zero-crossing rate refers to the number of times the signal crosses the zero level per unit time. For voice signals, especially in the vowel part, the zero-crossing rate is relatively low; while in voiceless consonants or noise, the zero-crossing rate is relatively high. The short-time zero-crossing rate reflects the number of times the audio signal crosses the zero point within a frame and is used to capture the frequency change characteristics of the signal. The double-threshold method is used to reduce false detection, and two different thresholds, a high threshold for starting voice detection and a low threshold for closing voice detection, are adopted.

[0081] Optionally, the implementation process of using a double-threshold method based on short-time energy and short-time zero-crossing rate to perform voice activity detection processing and endpoint localization processing on the voice data to obtain voice segment data includes but is not limited to the following steps.

[0082] First, calculate the short-time energy of the voice data. The short-time energy corresponding to the nth frame audio signal of the voice data satisfies the following formula (5): (5); In formula (5), N is the number of samples per frame, n is the sample index within the frame, and the value range of n is from 0 to N - 1.

[0083] Secondly, calculate the short-time zero-crossing rate of the voice data. The short-time zero-crossing rate corresponding to the nth frame audio signal of the voice data satisfies the following formula (6): (6); In formula (6), sgn is the sign function used to judge the sign change between adjacent sample values.

[0084] Then, set the short-time energy upper threshold , the short-time energy lower threshold , the short-time zero-crossing rate upper threshold and the short-time zero-crossing rate lower threshold . According to , , and , perform speech segment detection on the voice data. Specifically, if the short-time energy corresponding to the audio signal of the th frame and the short-time zero-crossing rate , then it is considered that the audio signal of the th frame belongs to the speech segment. If the short-time energy corresponding to the audio signal of the th frame or the short-time zero-crossing rate , then it is considered that the audio signal of the th frame does not belong to the speech segment.

[0085] For other audio signal frames between the high and low thresholds, the information of the front and back audio signal frames needs to be combined for further judgment. Specifically, if the audio signal of the th frame exists and , then within the judgment window of size corresponding to the audio signal of the th frame ( is odd), if the number of frames belonging to the speech segment is greater than the number of frames belonging to the non-speech segment, then it is considered that the audio signal of the th frame belongs to the speech segment, otherwise it is considered that the audio signal of the th frame does not belong to the speech segment. This can more sensitively capture the presence of speech at the beginning of the speech and maintain a certain tolerance at the end of the speech to avoid premature closing of the detection.

[0086] Finally, through the above steps, detect whether each frame of the audio signal in the voice data belongs to the speech segment to obtain the speech segment data.

[0087] In some embodiments of the present application, the implementation process of preprocessing the speech segment data by using the spectral subtraction method and the zero-mean unit-variance normalization method to obtain the pitch feature and the loudness feature includes but is not limited to the following steps.

[0088] First, before voice activity detection, the power spectral density estimate of noise is obtained by analyzing non-speech segments. For the noisy speech segment a short-time Fourier transform is performed to obtain its complex spectrum Assume that the estimated speech spectrum after denoising is where, represents the frequency, represents the time frame.

[0089] Secondly, the amplitude spectrum subtraction method is used to process the amplitude spectrum at each frequency point to obtain the amplitude spectrum after denoising , which satisfies the following formula (7): (7); In formula (7), represents the amplitude spectrum of the noisy speech segment ; is a regulation parameter, usually taking values between 2 and 4, used to control the degree of noise subtraction; The function is used to take the maximum value to ensure that the result will not be negative because the amplitude spectrum is non-negative.

[0090] Then, when using the spectral subtraction method for calculation, it usually refers to processing the amplitude spectrum while keeping the original phase unchanged. Therefore, the complex spectrum after denoising satisfies the following formula (8): (8): In formula (8), represents the phase angle of.

[0091] Furthermore, the corrected complex spectrum after denoising is converted back to the time-domain signal through the inverse short-time Fourier transform , that is, the denoised speech segment data in the form of a time series is obtained.

[0092] Then, for the denoised speech segment data, zero-mean unit-variance normalization processing is performed to obtain the normalized speech segment data sequence , The th frame audio signal in (9); In formula (9), represents the th frame audio signal in the denoised speech segment data sequence , represents The standard deviation of represents the mean value of

[0093] Finally, pitch features and loudness features are extracted from the normalized speech segment data sequence among them.

[0094] In some embodiments of the present application, in order to extract pitch features from the normalized speech segment data sequence, first the speech signal is segmented into short-time frames, and then a fundamental frequency estimation algorithm, such as the autocorrelation method and the cepstrum method, is applied to each frame to determine its fundamental frequency, that is, the high and low changes of the sound. Then, noise and discontinuities in the fundamental frequency estimation are reduced through smoothing processing, and finally the average fundamental frequency of each speech segment is calculated to obtain an average fundamental frequency sequence as the pitch feature.

[0095] In some embodiments of the present application, when extracting loudness features from the normalized speech segment data sequence, first the short-time energy of each frame is calculated, and different weights are assigned to the energies of different frequencies. Considering the weighted energies comprehensively, the average loudness of each speech segment is calculated to obtain an average loudness sequence as the loudness feature.

[0096] In some embodiments of the present application, deep learning includes a breathing data model, a pulse data model, and a voice data model. Referring to Figure 3 , in step 103, the implementation process of obtaining the user's emotion recognition result based on the breathing feature, pulse feature, and voice feature by using the deep learning model includes but is not limited to the following steps.

[0097] Step 301, according to the breathing feature, combined with the breathing data model, obtain a breathing emotion probability vector.

[0098] It should be noted that the breathing feature can reflect the changes in breathing frequency, depth, and rhythm, and can reflect the user's emotional state. The breathing feature has obvious time series characteristics and can capture the dynamic process of emotional changes.

[0099] In step 301, by combining the breathing data model to analyze the breathing feature, a breathing emotion probability vector is obtained. The function of this step is to initially identify the user's emotional state by using the stability and time series characteristics of the breathing feature, providing a reliable basis for subsequent multi-modal data fusion and final emotion recognition.

[0100] Step 302, according to the pulse feature, combined with the pulse data model, obtain a pulse emotion probability vector.

[0101] In step 302, by analyzing the pulse characteristics and combining with the pulse data model, a pulse emotion probability vector is obtained. The function of this step is to utilize the stability and local feature extraction ability of the pulse characteristics to further identify the user's emotional state, providing a reliable basis for subsequent multi-modal data fusion and final emotion recognition. Through the analysis of the pulse characteristics, the accuracy and robustness of emotion recognition can be improved, ensuring that the system can comprehensively understand the user's emotional changes.

[0102] In step 303, according to the voice characteristics and combining with the voice data model, a voice emotion probability vector is obtained.

[0103] In step 303, by analyzing the voice characteristics and combining with the voice data model, a voice emotion probability vector is obtained. The function of this step is to utilize the emotional expression ability and real-time nature of the voice characteristics to further identify the user's emotional state, providing a reliable basis for subsequent multi-modal data fusion and final emotion recognition. Through the analysis of the voice characteristics, the accuracy and meticulousness of emotion recognition can be improved, ensuring that the system can comprehensively understand the user's emotional changes, especially more effectively when expressing complex emotions.

[0104] In step 304, the respiratory emotion probability vector, the pulse emotion probability vector, and the voice emotion probability vector are fused to obtain the emotion recognition result of the user.

[0105] In step 304, by fusing the respiratory emotion probability vector, the pulse emotion probability vector, and the voice emotion probability vector, the final emotion recognition result is obtained. It can integrate the advantages of multiple modal data, overcome the limitations of single modal data, and improve the accuracy and robustness of emotion recognition. The final emotion recognition result is the emotion category with the highest probability among each emotion category, providing accurate emotion state recognition for the user.

[0106] Optionally, the weighted summation method is used to fuse the respiratory emotion probability vector, the pulse emotion probability vector, and the voice emotion probability vector to improve the accuracy and reliability of emotion recognition. Specifically, let the respiratory emotion probability vector be , the pulse emotion probability vector be , the voice emotion probability vector be , and let the fused emotion probability vector be , where is the number of emotion categories, then the th element in satisfies , where , and is the weight coefficient, which is determined according to the importance of different modality data for emotion recognition. The emotion category with the largest probability value in the fused emotion probability vector is the final recognized emotion result, that is, if , then the current emotional state is recognized as the th emotion category.

[0107] In some embodiments of the present application, referring to Figure 4 , in step 401, the implementation process of establishing a music track playlist according to the emotion recognition result and music preference data includes but is not limited to the following steps.

[0108] Step 401, establish a user's emotional music library according to the music preference data.

[0109] In step 401, according to the user's music preference data, such as genre preference, artist selection, and listening history, construct a highly personalized user emotional music library. The music is classified according to different emotion labels (such as happy, sad, relaxed) to ensure that each emotion has corresponding tracks. This library not only reflects the user's long-term preferences but also adapts to their changing tastes through continuous updates, laying a solid foundation for subsequent emotion matching.

[0110] Step 402, screen out the matching music tracks from the emotional music library according to the emotion recognition result and establish a music track playlist.

[0111] In step 402, according to the emotion recognition result, analyze the user's current emotional state, and accurately select the music tracks that can best reflect or regulate the emotion from the emotional music library to create a personalized playlist. This process emphasizes real-time and diversity, ensuring that the recommendation not only meets the user's immediate emotional needs but also provides a rich music experience. At the same time, allow the user to give feedback (like, skip, rate, etc.) on the recommendation result, so as to help the system better understand the user's true feelings and make more appropriate recommendations in the future.

[0112] In some embodiments of the present application, referring to Figure 4 , in step 401, the implementation process of establishing a user's emotional music library according to the music preference data includes but is not limited to the following steps.

[0113] Step 501, obtain the user's favorite tracks according to the music preference data.

[0114] In step 501, the system identifies and collects the user's preferred tracks by analyzing the user's music preference data. This data may come from multiple sources, such as the user's historical play records, favorite lists, ratings, and comments. Obtaining the user's favorite tracks is the basis for building a personalized mood music library, ensuring that subsequent mood classification and recommendations can closely match the user's actual preferences. This process not only improves the user experience but also provides accurate and relevant materials for the subsequent mood feature classification.

[0115] In step 502, perform mood feature classification and annotation on the favorite tracks to obtain the mood labels corresponding to the favorite tracks.

[0116] In step 502, conduct in-depth analysis on the obtained favorite tracks and classify and annotate each track according to the mood it conveys. This step generates the mood label corresponding to each track, such as "happy", "sad", "relaxed", or "excited", etc. The accuracy of the mood label is crucial for realizing personalized recommendations, enabling the system to provide the most suitable music choices for users in different emotional states and enhancing the connection between music and users' emotions.

[0117] In step 503, establish the user's mood music library based on the favorite tracks and their corresponding mood labels.

[0118] In step 503, integrate the results of the previous two steps and build a comprehensive and personalized user mood music library based on the user's favorite tracks and their corresponding mood labels. This library not only stores the user's favorite music but also details the mood attributes of each track. As the core resource for subsequent mood matching, the mood music library ensures that the system can quickly and accurately respond to the user's emotional changes and provide music recommendations that match the current mood. In addition, this library can be continuously updated with the user's newly added favorite tracks to maintain the freshness and relevance of the recommendations.

[0119] In some embodiments of the present application, referring to Figure 4 , in step 402, the implementation process of screening out matching music tracks from the mood music library according to the mood recognition result and establishing a music track playlist includes but is not limited to the following steps.

[0120] In step 601, according to the mood recognition result, screen out matching music tracks from the mood music library to obtain a set of music tracks suitable for the current mood.

[0121] In step 601, according to the emotion recognition result, the music tracks that can best reflect or regulate the emotion are selected from the pre-constructed emotion music library. This process ensures that the recommended music not only meets the user's immediate emotional needs but also effectively enhances or adjusts their mood. Through precise emotion matching, the system can provide a highly personalized music experience, enhancing the user's emotional resonance and satisfaction. The set of music tracks suitable for the current emotion provides the basic material for the creation of the subsequent playlist, ensuring the relevance and pertinence of the playlist.

[0122] In step 602, based on the set of music tracks suitable for the current emotion and combined with the preset sorting rules, a playlist of music tracks is established.

[0123] In step 602, based on the set of music tracks suitable for the current emotion, the preset sorting rules are applied to construct the final playlist of music tracks. The sorting rules can include various factors, such as the popularity of the song, the user's historical preferences, the emotional intensity of the music, and the time length, etc., to ensure the diversity and smoothness of the playlist. In addition, the sorting rules can also consider the smoothness of the transition between music, avoiding sudden mood changes and providing a more coherent auditory experience. In this way, the system not only realizes personalized recommendation but also optimizes the playing order, enabling users to feel the natural flow and support of emotions while enjoying the music. The finally generated playlist aims to maximize the user's emotional experience and music enjoyment.

[0124] In some embodiments of the present application, the breathing data model includes a long short-term memory network model.

[0125] The long short-term memory network model (Long Short-Term Memory, LSTM) is a special recurrent neural network, specifically designed to process and predict time series data. By introducing memory cells and gating mechanisms, LSTM can effectively capture long-term dependencies and solve the problems of gradient disappearance and gradient explosion in traditional RNNs when dealing with long sequence data.

[0126] Since breathing data has obvious time series characteristics, that is, there are temporal dependencies between data points. LSTM can capture these dependencies and identify changes in breathing patterns, thus more accurately reflecting the user's emotional state. The gating mechanism of LSTM makes the model more stable during the training process, enabling it to converge better and having good robustness to noise and outliers, improving the generalization ability of the model.

[0127] In addition, LSTM can automatically extract complex features in time series data, capture subtle changes in respiratory data, and improve the accuracy of emotion recognition. In multi-modal data fusion, LSTM can be combined with other models to jointly improve the accuracy and robustness of emotion recognition. Moreover, LSTM can achieve real-time data processing and emotion recognition, which is suitable for real-time monitoring of users' emotional changes and providing timely feedback and suggestions. Therefore, it is reasonable and effective to choose LSTM as the respiratory data model.

[0128] Optionally, the implementation process of obtaining the respiratory emotion probability vector by combining LSTM according to respiratory features includes but is not limited to the following steps.

[0129] First, take the respiratory features in the form of time series as the input of LSTM to obtain the predicted emotion label corresponding to the respiratory features .

[0130] Secondly, train the LSTM model. During the training process, the cross-entropy loss function is used to measure the difference between the model prediction result and the true result.

[0131] Optionally, the cross-entropy loss function satisfies the following formula (10): (10); In formula (10), is the number of samples, is the corresponding true emotion label, is the corresponding predicted emotion label.

[0132] Then, during the training process, continuously adjust the weights and biases of the LSTM model through the backpropagation algorithm to optimize the model performance.

[0133] Optionally, the backpropagation algorithm gradually updates the parameters according to the gradient of the loss function with respect to the model parameters to reduce the loss. Let the parameters of the model be , then the update formula of satisfies the following formula (11): In formula (11), is the learning rate; represents the gradient, indicating the change direction of the loss function at the current parameter values.

[0134] Finally, continuously train the LSTM model, calculate the gradient of the loss function with respect to the parameters, and update the parameters along the opposite direction of the gradient to minimize the loss function until the accuracy of the LSTM model on the validation set reaches a stable state.

[0135] In some embodiments of the present application, the pulse data model includes a convolutional neural network model.

[0136] A Convolutional Neural Network (CNN) is a deep learning model specifically designed to process data with grid structures (such as images and signals). Through structures such as convolutional layers, pooling layers, and fully connected layers, CNN can automatically extract local features and spatial relationships in the data, capturing the key information in the data.

[0137] In the pulse data model, the pulse data has obvious local features, such as heart rate changes and morphological features of the pulse waveform. CNN can extract these local features through convolutional layers, reduce the data dimension through pooling layers, retain key information, improve the computational efficiency and robustness of the model. When processing pulse data, CNN can effectively capture the subtle changes in heart rate and pulse waveform, improving the accuracy of emotion recognition. In addition, CNN has good robustness to noise and outliers, can handle imperfect data, and improve the generalization ability of the model. Therefore, it is reasonable and effective to choose CNN as the pulse data model, which can provide high-precision emotion recognition results and is suitable for real-time monitoring and multi-modal data fusion.

[0138] Optionally, the implementation process of obtaining the pulse emotion probability vector by combining CNN according to pulse features includes but is not limited to the following steps.

[0139] First, set the hyperparameters of CNN.

[0140] Optionally, CNN contains multiple convolutional layers and pooling layers. For the convolutional layer, randomly set the convolutional kernel size of each convolutional layer to 3x3 or 5x5, set the convolutional stride of the convolutional layer to 1, and set the padding method to SAME to ensure that the output feature map has the same size as the input. Then, use the ReLU function as the activation function to enhance the non-linear expression ability of CNN. For the pooling layer, use the max-pooling method, set the pooling window size to 2x2, and set the pooling stride to 2. Through convolutional and pooling operations, gradually reduce the data dimension and extract the key feature patterns in the pulse signal. These settings help the model efficiently capture the local features and spatial relationships of the pulse signal, improving the accuracy and robustness of emotion recognition.

[0141] Second, use the pulse features and their corresponding true emotion labels as the input of CNN to obtain the predicted emotion labels corresponding to the pulse features. Then, train CNN according to the true emotion labels and predicted emotion labels corresponding to the pulse features.

[0142] Optionally, during the training process, according to the aforementioned formula (10), the cross-entropy loss function is used to measure the difference between the true emotion label and the predicted emotion label corresponding to the pulse feature.

[0143] Finally, according to the aforementioned formula (11), the CNN is optimized through the backpropagation algorithm.

[0144] In some embodiments of the present application, the sound data model includes a deep belief network model.

[0145] A deep belief network (DBN) is a deep generative model composed of multiple restricted Boltzmann machines (RBMs) stacked together. An RBM is a shallow unsupervised learning model mainly used to learn feature representations from data. In a DBN, multiple RBMs can learn complex feature representations in the data through layer-by-layer unsupervised pre-training, and then further optimize the model through supervised fine-tuning.

[0146] Through unsupervised pre-training and supervised fine-tuning, DBN can learn complex feature representations in the data. In the sound data model, the sound data contains rich information such as pitch, loudness, etc., and these features are closely related to the emotional state. DBN can automatically extract high-level abstract features in the sound data during the unsupervised pre-training stage, capturing the subtle changes and complex patterns in the sound. Subsequently, through the supervised fine-tuning stage, DBN can further optimize the model and improve the accuracy of emotion recognition. DBN has good robustness to noise and outliers, can handle imperfect data, and improve the generalization ability of the model. In addition, DBN can capture the emotional fluctuations of users, provide more detailed emotion recognition, and is particularly effective when expressing complex emotions. Therefore, it is reasonable and effective to choose DBN as the sound data model, which can provide high-precision emotion recognition results and is suitable for real-time monitoring and multi-modal data fusion.

[0147] Optionally, the implementation process of obtaining the sound emotion probability vector by combining DBN according to the sound features includes but is not limited to the following steps.

[0148] First, in the unsupervised pre-training stage, each RBM is trained in turn to maximize the likelihood of the data, and the parameters are updated through the contrastive divergence algorithm. Specifically, the energy function of the RBM satisfies the following formula (12): (12); In formula (12), represents the number of nodes in the visible layer, represents the number of nodes in the hidden layer, represents the The state of the visible layer nodes, represents the state of the th hidden layer node, represents the bias term of the th visible layer node, represents the bias term of the th hidden layer node, represents the weight of the connection between the th visible layer node and the th hidden layer node. These variables together constitute the energy function of the RBM , where , and . The energy function reflects the interaction strength between the visible layer and the hidden layer and is one of the optimization objectives during the RBM training process.

[0149] Then, in the supervised fine-tuning stage, the entire DBN is regarded as a deep neural network, and the sound features are used as the input of the DBN to obtain the predicted emotion labels corresponding to the sound features. According to the aforementioned formulas (10) and (11), based on the predicted emotion labels and the true emotion labels corresponding to the sound features, the DBN is trained using the cross-entropy loss function.

[0150] In summary, the emotion-adaptive intelligent music playback method provided by the embodiments of the present application has the following technical effects.

[0151] The emotion-adaptive intelligent music playback method provided by the embodiments of the present application effectively improves the effect of personalized music recommendation. Through precise emotion recognition technology and in-depth analysis of users' music preferences, it can respond to users' emotional changes in real time, select the most suitable tracks from a carefully constructed emotion music library, ensuring the effectiveness and pertinence of the recommendation. This method not only considers users' long-term preferences but also can dynamically update the music library, continuously optimize the recommendation model, and maintain the freshness and relevance of the content. When creating a playlist, the system combines multiple preset sorting rules, such as song popularity, historical preferences, and emotional intensity, to ensure the diversity and smooth transition of the playlist, providing a coherent and rich auditory experience. In addition, the intelligent selection and arrangement of music tracks not only improve users' emotions but also help relieve stress and improve mood, becoming an effective tool for improving the quality of life, providing a more intelligent, emotional, and user-friendly music experience, and having broad application prospects and social value.

[0152] Secondly, referring to Figure 5 , the embodiments of the present application provide an emotion-adaptive intelligent music playback system, including: a data acquisition module, a feature processing module, an emotion recognition module, and a music playback module.

[0153] The data acquisition module is used to acquire the user's breathing data, pulse data, voice data, and music preference data.

[0154] The feature processing module is used to perform feature processing on the breathing data, pulse data, and voice data to obtain breathing features, pulse features, and voice features.

[0155] The emotion recognition module is used to obtain the user's emotion recognition result according to the breathing features, pulse features, and voice features by using a deep learning model.

[0156] The music playback module is used to establish a music track playlist according to the emotion recognition result and music preference data, and play music according to the music track playlist.

[0157] Furthermore, referring to Figure 6 , an emotion-adaptive intelligent music playback device provided by an embodiment of the present application includes: a processor and a memory. The memory is used to store a program. When the program is executed by the processor, the processor implements the foregoing emotion-adaptive intelligent music playback method.

[0158] In addition, an embodiment of the present application provides a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to implement the foregoing emotion-adaptive intelligent music playback method when executed by the processor.

[0159] Similarly, the content in the foregoing method embodiments is applicable to system embodiments, device embodiments, and medium embodiments. The functions specifically implemented by the system embodiments, device embodiments, and medium embodiments are the same as those in the foregoing method embodiments, and the beneficial effects achieved are also the same as those in the foregoing method embodiments.

[0160] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation schematic diagram. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously, or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.

[0161] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0162] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0163] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable programs for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by a program execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can retrieve and execute programs from the program execution system, apparatus, or device), or in conjunction with these program execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with a program execution system, apparatus, or device.

[0164] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program is printed, as the program can be obtained from the paper or other media by optical scanning or the like and stored in a computer memory.

[0165] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for performing logical operations on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0166] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0167] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0168] The above has specifically described the best implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. An intelligent music playing method adapted to emotions, characterized in that: The steps include: Obtain the user's breathing data, pulse data, sound data, and music preference data; Performing feature processing on the breathing data, the pulse data and the sound data to obtain breathing features, pulse features and sound features; Obtaining an emotion recognition result of the user using a deep learning model according to the breathing feature, the pulse feature and the sound feature; Creating a music playlist based on the emotion recognition result and the music preference data; Music is played according to the music track playlist.

2. The method for playing music adapted to emotions according to claim 1, characterized in that: The step of performing feature processing on the breathing data, the pulse data and the sound data to obtain breathing features, pulse features and sound features includes: Preprocessing the respiratory data to obtain respiratory frequency characteristics, respiratory depth characteristics and respiratory rhythm characteristics; Performing feature fusion processing on the breathing frequency feature, the breathing depth feature and the breathing rhythm feature to obtain the breathing feature; Preprocessing the pulse data to obtain heart rate characteristics and heart rate variability characteristics; Performing feature fusion processing on the heart rate feature and the heart rate variability feature to obtain the pulse feature; Performing voice activity detection processing and endpoint location processing on the sound data to obtain voice segment data; Preprocessing the speech segment data to obtain pitch features and loudness features; The tone feature and the loudness feature are subjected to feature fusion processing to obtain the sound feature.

3. The method for playing music adapted to emotions according to claim 1, characterized in that: The deep learning includes a breathing data model, a pulse data model and a sound data model; The step of obtaining the emotion recognition result of the user by using a deep learning model according to the breathing feature, the pulse feature and the sound feature includes: According to the breathing characteristics, combined with the breathing data model, a breathing emotion probability vector is obtained; According to the pulse characteristics, combined with the pulse data model, a pulse emotion probability vector is obtained; According to the sound features, combined with the sound data model, a sound emotion probability vector is obtained; The breathing emotion probability vector, the pulse emotion probability vector and the sound emotion probability vector are fused to obtain an emotion recognition result of the user.

4. The method for playing music adapted to emotions according to claim 1, characterized in that: The step of establishing a music track playlist according to the emotion recognition result and the music preference data comprises: Establishing an emotional music library of the user according to the music preference data; According to the emotion recognition result, matching music tracks are screened out from the emotion music library and a playlist of the music tracks is established.

5. The method for playing music adapted to emotions according to claim 4, characterized in that: The step of establishing the user's emotional music library according to the music preference data comprises: Obtaining the user's favorite music tracks according to the music preference data; Classify and label the favorite songs according to their emotional characteristics to obtain emotional labels corresponding to the favorite songs; An emotional music library of the user is established according to the favorite songs and their corresponding emotional tags.

6. The method for playing music adapted to emotions according to claim 4, characterized in that: The step of selecting matching music tracks from the emotion music library and establishing a playlist of the music tracks according to the emotion recognition result comprises: According to the emotion recognition result, matching music tracks are selected from the emotion music library to obtain a set of music tracks suitable for the current emotion; The music track playlist is established based on the music track set that suits the current mood and in combination with a preset sorting rule.

7. The method for playing music adapted to emotions according to claim 3, characterized in that: The respiratory data model includes a long short-term memory network model; the pulse data model includes a convolutional neural network model; and the sound data model includes a deep belief network model.

8. An intelligent music playing system adapted to emotions, characterized in that: include: Data acquisition module, feature processing module, emotion recognition module and music playing module; The data acquisition module is used to acquire the user's breathing data, pulse data, sound data and music preference data; The feature processing module is used to perform feature processing on the breathing data, the pulse data and the sound data to obtain breathing features, pulse features and sound features; The emotion recognition module is used to obtain the emotion recognition result of the user according to the breathing characteristics, the pulse characteristics and the sound characteristics by using a deep learning model; The music playing module is used to create a music track playlist according to the emotion recognition result and the music preference data; Music is played according to the music track playlist.

9. An intelligent music playing device adapted to emotions, characterized in that: include: Processor and memory; The memory is used to store programs; when the program is executed by the processor, the processor implements the intelligent music playing method adapted to emotions as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to implement the emotion-adaptive intelligent music playing method as described in any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Bluetooth earphone playing control method

    CN117041807A

  • Personalized music recommendation method and device, equipment, medium and wearable equipment

    CN118013071A