Emotional resonance audio system

Through the emotional resonance audio system that combines deep learning and physiological feedback, audio parameters can be identified and dynamically adjusted in real time, solving the problems of inaccurate emotion recognition and imprecise adjustment in existing technologies, and achieving a personalized and natural audio experience.

CN120708659APending Publication Date: 2025-09-26SHANGHAI ORIENT CHIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510784448.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing audio processing technology is unable to recognize user emotions in real time and make dynamic adjustments, resulting in poor emotional resonance effects and an inability to meet personalized emotional needs, especially in the lack of effective support in psychotherapy and emotion regulation scenarios.

Method used

An emotion recognition module based on deep learning is used, combined with a physiological feedback module and an emotion mapping module. The CNN and LSTM hybrid model is used to analyze audio signals in real time, dynamically adjust audio parameters such as volume, pitch, rhythm, etc., establish emotion-audio feature mapping tables and conversion rules, and integrate physiological feedback data for personalized adjustments.

Benefits of technology

It achieves multi-dimensional and precise audio emotion adjustment, and can adaptively adjust according to the user's immediate emotional changes, providing a natural, immersive and personalized audio experience, improving the accuracy of emotion recognition and the intelligence of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708659A_ABST
    Figure CN120708659A_ABST
Patent Text Reader

Abstract

The invention provides an emotion resonance audio system which comprises an emotion recognition module, an audio processing module and a user feedback module. The emotion recognition module is set to obtain a recognized emotion type according to the original audio signal; the audio processing module comprises an audio parameter adjusting unit which is set to obtain the adjusting quantity and the adjusting rate of the real-time audio parameter and process the original audio signal according to the adjusting quantity and the adjusting rate of the real-time audio parameter; the user feedback module is set to collect personalized data of a user, and the audio parameter adjustment unit dynamically adjusts the adjustment amount and the adjustment rate of audio parameters by using the personalized data of the user collected by the user feedback module. The emotion resonance audio system solves the problems that in the prior art, emotion recognition accuracy is insufficient, audio parameter adjustment is not fine enough, and the system cannot conduct self-adaptive adjustment according to instant emotion changes of a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of audio processing, and in particular relates to an emotional resonance audio system. Background Art

[0002] In today's digital age, audio content plays an increasingly important role in people's daily lives. From music and audiobooks to audio effects in movies and games, audio not only conveys information but also carries rich emotional expressions. However, traditional audio processing technologies have significant limitations in terms of emotional experience and are unable to meet users' growing demand for personalized emotions.

[0003] Existing audio processing systems typically adjust based on preset parameters, which are determined during audio production. These parameters lack the ability to analyze and dynamically adjust the emotional characteristics of audio content in real time. For example, when playing a sad song, even if the listener is currently in a happy mood, the audio's volume, pitch, tempo, and other parameters will not automatically adjust based on the listener's emotional state. This results in the audio failing to resonate with the listener's emotions, affecting their immersion and emotional engagement.

[0004] The emotional regulation role of audio is particularly important in specialized application scenarios, such as psychotherapy, emotional regulation, and rehabilitation training. However, existing audio processing technologies cannot effectively support these scenarios. For example, during psychotherapy, therapists need to adjust audio content in real time based on the patient's emotional state to help them relax, relieve anxiety, or stimulate positive emotions. However, current audio systems are unable to achieve this real-time emotional feedback and audio adjustments, making them unable to provide optimal emotional support for patients.

[0005] With the rapid development of artificial intelligence and machine learning technologies, emotion recognition and analysis have become possible. These technologies can extract emotional features such as pitch, rhythm, and speaking speed from audio signals and classify them into different emotional states. However, combining emotion recognition technology with audio processing technology to achieve emotional adaptive adjustment of audio content remains an underexplored area. Currently, there is no audio system on the market that can recognize audio emotions in real time and dynamically adjust them based on user emotional feedback, which provides an innovative opportunity for this patent application.

[0006] Existing technologies for emotional audio recognition and adjustment typically rely on simple emotion classification or manually annotated emotion tags, lacking a deep understanding and accurate recognition of complex emotional fluctuations. Furthermore, existing systems are typically limited to adjusting basic audio parameters such as volume and pitch, failing to achieve multi-dimensional, delicate, and dynamic adjustments to audio emotional expression. Existing technologies also have significant shortcomings in addressing users' personalized needs and emotional fluctuations. They are unable to accurately adapt to users' immediate emotional needs and fail to incorporate physiological feedback for real-time audio adjustments.

[0007] It is necessary to propose an emotional audio recognition and adjustment system based on deep learning to solve the problems in the existing technology, such as insufficient accuracy of emotion recognition, insufficient fine adjustment of audio parameters, and the inability of the system to adaptively adjust according to the user's immediate emotional changes. Summary of the Invention

[0008] The purpose of the present invention is to propose an emotional resonance audio system to solve the problems in the prior art such as insufficient accuracy of emotion recognition, insufficient fine adjustment of audio parameters, and the system's inability to adaptively adjust according to the user's immediate emotional changes.

[0009] In order to achieve the above-mentioned object, the present invention provides an emotional resonance audio system, comprising an emotion recognition module, an audio processing module and a user feedback module;

[0010] The emotion recognition module is configured to obtain a recognized emotion type based on the original audio signal;

[0011] The audio processing module includes an audio parameter adjustment unit, which is configured to obtain a real-time adjustment amount and adjustment rate of the audio parameter, and process the original audio signal according to the real-time adjustment amount and adjustment rate of the audio parameter;

[0012] The user feedback module is configured to collect personalized data of the user, and the audio parameter adjustment unit utilizes the personalized data collected by the user feedback module to dynamically adjust the adjustment amount and adjustment rate of the audio parameter.

[0013] The audio processing module directly obtains the adjustment amount and initial value of the adjustment rate of the real-time audio parameter according to the recognized emotion type and the predefined emotion type conversion rule;

[0014] Alternatively, the emotional resonance audio system also includes an emotion mapping module, which is configured to: establish an emotion-audio feature mapping table and an emotion type conversion rule based on the audio signal and its corresponding emotion type, and generate a real-time audio parameter adjustment amount and an initial value of the adjustment rate according to the emotion type of the original audio signal, the collection result of the conversion trigger condition, and the emotion type conversion rule.

[0015] The emotion mapping module is configured to perform the following steps:

[0016] Step S210: constructing an emotion-audio feature mapping table based on a large number of audio signals and their corresponding emotion types;

[0017] In the emotion-audio feature mapping table, the extracted audio features include volume, pitch, timbre, speaking rate, pause pattern, dynamic range, language energy distribution, and speech envelope;

[0018] Step S220: establishing conversion rules for emotion types, including conversion rules for different emotion types;

[0019] In step S220, conversion rules for different emotion types are established, specifically including:

[0020] Step S221: Establish an emotion conversion matrix to store conversion rules for different emotion types, where rows represent original emotion types, identified emotion types are used as original emotion types, and columns represent target emotion types. A set of original emotion types and target emotion types are converted as one emotion type.

[0021] Step S222: For each emotion type conversion, a conversion rule for smoothly transitioning from the original emotion to the target emotion is established based on the threshold range of the audio parameters of the original emotion type and the target emotion type in the emotion-audio feature mapping table, and the rule serves as an element corresponding to the emotion type conversion in the emotion conversion matrix; the emotion type conversion rule includes an adjustment amount of the audio parameter and an initial value of the adjustment rate;

[0022] Step S223: Setting a conversion trigger condition for each emotion type conversion as part of the conversion rules for different emotion types. The conversion trigger condition is used to obtain initial values ​​of the adjustment amount and adjustment rate of the audio parameter according to the corresponding emotion type conversion rule when activated;

[0023] Step S230: generating initial values ​​of the adjustment amount and adjustment rate of the real-time audio parameter according to the emotion type of the original audio signal, the collection result of the conversion trigger condition, and the conversion rule of the emotion type;

[0024] The user feedback module is configured to collect personalized data of the user and provide the collected data as a conversion trigger condition to the emotion mapping module.

[0025] The conversion rules of the emotion types also include conversion rules of different emotion intensities of the same emotion type;

[0026] Establish the conversion rules of different emotion intensities of the same emotion type, including:

[0027] Step S221': Divide each emotion type into multiple emotion intensities;

[0028] Step S222': for each emotion type, obtaining the adjustment amplitude of the audio parameter corresponding to the initial emotion intensity and the target emotion intensity under the emotion type according to the audio sample under the emotion type;

[0029] Step S223': for different emotional intensities of the same emotional type, setting the adjustment amplitude and adjustment rate of the audio parameter when switching between different emotional intensities according to the adjustment amplitude of the audio parameter corresponding to the initial emotional intensity and the target emotional intensity;

[0030] Step S224 ′: determining a triggering condition for the conversion of the emotion intensity, and determining a method for dynamically adjusting the audio parameters according to a dynamic temporal change trend of the emotion intensity when the conversion of the emotion intensity is triggered.

[0031] The emotion mapping module is further configured to: during the execution of steps S210 and S220, display the curves and constraint boundaries of the adjustment rates of the audio parameters according to the real-time monitoring interface, wherein the constraint boundaries are the boundaries of the typical combination of audio features corresponding to each emotion type; when necessary, import user personalized data through the user feedback module to adjust the curves and constraint boundaries of the adjustment rates of the audio parameters; and / or

[0032] In the emotion mapping module, after establishing the conversion rules for different emotion types, it also includes: establishing a neural network model, which is used to obtain a combination of adjusted values ​​of audio parameters corresponding to different target emotion types according to the currently input audio signal when user satisfaction is too low, and the difference between the adjusted value of the audio parameter and the initial value of the audio parameter is the adjustment amount of the audio parameter, thereby obtaining the real-time adjustment amount of the audio parameter and the initial value of the adjustment rate; the input of the neural network model is the audio signal, and the output is a combination of the target emotion type and the corresponding adjusted value of the audio parameter; when obtaining training samples, the audio parameter combination vector corresponding to each target emotion type includes the audio parameter combination vector finally confirmed by the user or evaluated by the system as satisfactory, and the high-quality audio parameter combination vector debugged by professionals based on the experience of emotional voice design.

[0033] The user feedback module includes a physiological feedback acquisition unit and an interactive interface feedback unit, and the personalized data includes physiological data collected by the physiological feedback acquisition unit and user evaluations and preferences collected by the interactive interface feedback unit;

[0034] The physiological feedback acquisition unit includes a heart rate sensor and an acceleration sensor built into a mobile phone and a skin galvanic response sensor of a smart watch; the interactive interface feedback unit adopts a dedicated application interface provided on the user's mobile phone or smart watch.

[0035] The audio parameter adjustment unit is configured to execute the following global optimization algorithm:

[0036] Step S310: Receive personalized data and historical adjustment behaviors collected by the user feedback module as a set of personalized adjustment samples; the historical adjustment behaviors include initial values ​​of the real-time audio parameter adjustment amount and adjustment rate, as well as the audio parameter adjustment amount and adjustment rate finally determined by the user;

[0037] Step S320: Clean, normalize, and fill missing values ​​in the set of personalized adjustment samples. A fused feature vector is constructed based on the personalized data from the historical adjustment behavior and the real-time audio parameter adjustment amount and adjustment rate as input. The final audio parameter adjustment amount and adjustment rate determined by the user in the historical adjustment behavior is output as the output, making it suitable for subsequent analysis and modeling.

[0038] Step S330: constructing the input layer, hidden layer, and output layer of the neural network model, and using samples for training to obtain the regularity between the personalized data and the adjustment amount and adjustment rate of the audio parameters.

[0039] The audio processing module further includes a dynamic adjustment algorithm unit, which processes the audio output by the audio parameter adjustment unit using a dynamic compression algorithm and an expander algorithm to ensure a smooth transition of the audio output.

[0040] The emotion recognition module is configured to perform the following steps:

[0041] Step S110: Acquire an audio signal, extract time domain and frequency domain features of the audio signal, and label the emotion category as an audio dataset;

[0042] Step S120: pre-training a deep learning model using an audio data set to obtain a sentiment classification model, wherein the sentiment classification model is used to output a corresponding sentiment type based on time domain and frequency domain features of the audio signal;

[0043] Step S140: obtaining the original audio signal, extracting its time domain features and frequency domain features, and using the emotion classification model to obtain the emotion type of the original audio signal.

[0044] The emotion recognition module is further configured to:

[0045] Step S130: Statistically analyzing the frequency domain features extracted in step S110 and the emotion type of each frequency domain feature to obtain the frequency domain energy of the audio signal for each emotion type. Based on the frequency domain energy threshold, a corresponding frequency domain energy threshold is set for each emotion type. The frequency domain energy threshold is used to determine whether the emotion type of the audio signal needs to be reviewed.

[0046] Step S140 further includes: calculating the frequency domain energy of the audio signal from the frequency domain features, and comparing the frequency domain energy with a frequency domain energy threshold to determine whether the emotion type of the audio signal needs to be reviewed.

[0047] Through innovative emotion mapping and dynamic audio parameter adjustment mechanisms, combined with deep learning models and user physiological feedback data, the present invention can adjust audio in real time and accurately in multiple dimensions, provide a personalized and natural audio experience, and perform adaptive adjustments based on the user's immediate emotional changes.

[0048] First, the present invention is based on a deep learning-based emotional audio recognition model, which extracts time and frequency domain features from audio signals to perform emotion recognition. Unlike traditional recognition methods based on static emotion labels, the present invention introduces a hybrid model combining CNN and LSTM, which can analyze emotional fluctuations in audio signals in real time, significantly improving the accuracy and robustness of emotion recognition. Traditional solutions usually rely on simple emotion classification or emotion labels based on manual annotation, which are difficult to handle complex and varied emotional expressions. Through the application of deep learning models, the present invention can more accurately capture subtle emotional changes.

[0049] Secondly, the present invention proposes an emotion mapping module for generating emotion conversion rules. Traditional audio processing solutions usually rely only on simple volume and pitch adjustments without involving emotion types. By combining the identified emotion types, the system can accurately and dynamically adjust the volume, pitch, rhythm and other dimensions of the audio in real time according to the emotion conversion rules to ensure that the emotional expression of the audio is always natural and smooth. Especially when emotions change drastically, the system can avoid drastic fluctuations in audio parameters, improve the listening experience, and prevent distortion caused by over-adjustment.

[0050] Furthermore, the present invention innovatively integrates a physiological feedback acquisition unit. By using smart devices (such as watches and smart headphones) to monitor the user's physiological data (such as heart rate and galvanic skin response) in real time, the present invention combines this data with the user's evaluation preferences for comprehensive analysis, dynamically adjusting the audio playback effect. Unlike traditional systems that rely solely on emotion recognition, the present invention, through the incorporation of physiological feedback, can more accurately adapt to the user's immediate emotional needs, thereby providing a more immersive and personalized audio experience.

[0051] Furthermore, this invention provides customized emotional audio adjustments based on users' evaluation preferences and physiological responses. Through continuous optimization using machine learning algorithms, it adapts to changing user emotions, improving the accuracy of understanding and responding to users' emotional needs over the long term. Compared to traditional emotional audio systems, this innovative mechanism avoids fixed emotion models and static parameter settings, making the system more intelligent and adaptable over the long term.

[0052] Another innovation of the present invention is the global optimization algorithm implemented by the audio parameter adjustment unit. This algorithm integrates multi-source emotional information from the original audio signal, user feedback, and physiological data, and comprehensively considers the accuracy of audio emotional expression (through the emotion-audio feature mapping table), user comfort (through the user feedback module), sound quality stability (through the dynamic adjustment algorithm unit), and emotional transmission effect (through the rules designed by the emotion mapping module). The present invention can more comprehensively analyze emotional changes, making audio adjustments more precise and adaptable, providing a richer and more delicate emotional experience. It also enables global optimization across multiple dimensions, avoiding excessive emotional exaggeration or loss of sound quality, and ensuring the optimal performance of the audio system.

[0053] The combination of these innovative technologies gives the present emotional audio system significant advantages in emotion recognition, audio adjustment, user experience, and system adaptability. Compared to existing technologies, this system not only improves the accuracy of emotion recognition but also intelligently adjusts to the user's individual needs and physiological responses, ultimately achieving a more natural, immersive, and personalized audio experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a structural block diagram of an emotional resonance audio system according to an embodiment of the present invention.

[0055] Figure 2 This is a flow chart of the steps performed by the emotion recognition module of the emotional resonance audio system.

[0056] Figure 3 This is a structural block diagram of the audio processing module of the emotional resonance audio system.

[0057] Figure 4 This is a structural block diagram of the user feedback module of the emotional resonance audio system.

[0058] Figure 5 It is a structural block diagram of the audio device in which the emotional resonance audio system of the present invention is located. DETAILED DESCRIPTION

[0059] like Figure 1The figure shows an emotional resonance audio system according to an embodiment of the present invention, which includes an emotion recognition module 10, an emotion mapping module 20, and an audio processing module 30 connected in sequence, and a user feedback module 40 connected to the emotion recognition module 10, the emotion mapping module 20, and the audio processing module 30 at the same time to provide the user's personalized data to the emotion recognition module 10, the emotion mapping module 20, and the audio processing module 30.

[0060] like Figure 2 As shown, the emotion recognition module 10 is configured to obtain the recognized emotion type based on the original audio signal.

[0061] like Figure 1 As shown, the emotion recognition module 10 is configured to perform the following steps:

[0062] Step S110: Acquire an audio signal, extract its time and frequency domain features, and label the emotion categories to form an audio dataset. This completes the collection and preprocessing of training data.

[0063] Before extracting the time domain and frequency domain features of the audio signal, the audio signal is denoised to eliminate background noise and ensure that the signal is clear and effective.

[0064] The time and frequency domain features of audio signals can be extracted through short-time Fourier transforms, energy calculations, and autocorrelation analysis. Common features include Mel-frequency cepstral coefficients and short-time energy, which are used to describe the spectral characteristics and time-domain dynamics of audio signals.

[0065] Short-time Fourier transform is a method that converts time-domain signals into time-frequency domain representation. With the help of sliding window function and Fourier transform, it can accurately capture the frequency distribution of audio signals in different time segments.

[0066] Energy calculation is used to obtain amplitude information from the time-domain characteristics of audio signals, typically by calculating the sum of squares or square root of the signal to reflect signal strength. In audio processing, short-time energy refers to the square sum of the amplitudes of all sampling points within each frame and is used to characterize the dynamic characteristics of audio signals.

[0067] The fundamental frequency component of an audio signal can be extracted through autocorrelation analysis. Autocorrelation analysis effectively identifies the periodicity of an audio signal by calculating the correlation between the signal and itself at different time delays, thereby determining its fundamental frequency.

[0068] Step S120: pre-training the deep learning model using the audio data set to obtain a sentiment classification model, where the sentiment classification model is used to output a corresponding sentiment type based on the time domain and frequency domain features of the audio signal.

[0069] In step S110, data collection and preprocessing collects a large amount of audio data labeled with emotion categories. The audio signal is denoised to eliminate background noise and ensure that the signal is clear and effective.

[0070] When selecting a deep learning model, the appropriate architecture is usually selected based on the task requirements, that is, the type of time domain and frequency domain features to be extracted. Among them, convolutional neural networks are suitable for automatically extracting local features from spectrograms, and are particularly capable of capturing frequency domain features in spectrograms. Recurrent neural networks or their variants, long short-term memory networks, are good at processing time series data and can learn long-term time dependencies in audio signals. Therefore, in this embodiment, the deep learning model adopts a hybrid model that combines CNN and LSTM, which can simultaneously process the frequency domain features and time domain features of audio signals in the same model, providing more powerful characterization capabilities. In other embodiments, Transformer can be used instead of CNN+RNN model (higher parallel computing capabilities).

[0071] Use the audio dataset to pre-train the deep learning model, specifically including: first, divide the audio dataset into training set, validation set and test set; then use the backpropagation algorithm to train the model with the training set, optimize the network weights, and minimize the loss function (such as cross entropy loss) through the gradient descent algorithm; use the validation set to achieve hyperparameter optimization to ensure that the model has good generalization ability; then, use the test set to evaluate the accuracy of the model, and adjust the network structure based on the results or use transfer learning methods to improve model performance, avoid overfitting, and achieve model evaluation and optimization.

[0072] Step S130 (optional): Statistics are performed based on the frequency domain features extracted in step S110 and the emotion type of each frequency domain feature to obtain the frequency domain energy of the audio signal of each emotion type (such as anger, happiness, sadness, etc.). Based on this, a corresponding frequency domain energy threshold is set for each emotion type. The frequency domain energy threshold is used to confirm whether the emotion type of the audio signal needs to be reviewed.

[0073] Although the features extracted in step S120 already cover the spectral features, and the trained emotion classification model is already able to output the corresponding emotion type based on the audio signal, the deep learning model shows good classification accuracy during the training process. However, in the actual deployment environment, the boundaries between emotion types are often blurred and overlapping, which makes the model prone to "overfitting" the emotion category. For example, in the subtle transition area between "sadness" and "calmness", the model may frequently make misjudgments, while the frequency domain energy of the actual audio is more consistent with the typical characteristics of the "calm" type. Secondly, under non-ideal conditions such as strong background noise, speech distortion, and echo interference, the output stability of the model is greatly reduced. As a physical observable, frequency domain energy is less affected by environmental changes and can provide an emotion verification mechanism independent of the model results. Therefore, by setting the frequency domain energy threshold, it can be used as a secondary filtering mechanism to constrain and correct the model output results.

[0074] On this basis, the present invention can also use the frequency domain energy threshold as an auxiliary judgment means. Thus, the emotion classification model first outputs the preliminary identified emotion type based on the time domain and frequency domain signal characteristics of the audio signal. Subsequently, the system calculates the frequency domain energy of the audio signal and compares the frequency domain energy with the energy threshold interval corresponding to the identified emotion type: if the frequency domain energy value is within the energy threshold range of the emotion type, the output result of the emotion classification model is confirmed; if it does not fall into the threshold interval of the emotion type, the recognition result is marked as pending review, user feedback is requested, or the energy threshold interval is adjusted to further review the emotion type of the audio signal. In other words, if an audio sample is identified as the emotion type by the emotion classification model, but its actual frequency domain energy value does not fall within the energy threshold interval of the emotion type, the system can trigger a round of review, request user feedback, or adjust the confidence of the model prediction to adjust the energy threshold interval to improve the accuracy of the recognition system.

[0075] Step S130 is based on the following principle: audio signals of different emotional types exhibit different energy distribution patterns in the frequency domain. For example, angry audio signals typically have higher high-frequency energy, while sad or calm audio signals tend to be more prominent in the low-frequency portion. By statistically analyzing these frequency domain characteristics, an appropriate energy threshold range can be set for each emotional type.

[0076] Therefore, in step S130, a corresponding frequency domain energy threshold is also required for each emotion type, so that the extracted frequency domain features can be more specifically applied and analyzed to meet the needs of emotion classification. This step further analyzes and processes the frequency domain features of the audio signal to statistically obtain the frequency domain energy distribution of the audio signal for each emotion type, and accordingly sets a corresponding frequency domain energy threshold for each emotion type.

[0077] Among them, the frequency domain energy distribution of the audio signal of each emotion type is obtained by statistics, and the corresponding frequency domain energy threshold is set for each emotion type accordingly, specifically including: calculating the frequency domain energy of each audio signal, and forming a statistical sample set with the frequency domain energies of all samples belonging to a certain emotion type, and obtaining the mean and standard deviation of the frequency domain energy of all audio signals of each emotion type based on the statistical sample set; for each emotion type, determining the frequency domain energy threshold according to its frequency domain energy mean and standard deviation.

[0078] The frequency domain energy threshold is determined by the confidence interval method. Therefore, the frequency domain energy threshold is determined specifically as follows: first, for each emotion type, a large number of historical audio samples are collected, and each audio sample is an emotion type that has been manually or model-labeled; for each audio sample, its spectral characteristics are obtained by short-time Fourier transform and other methods, and its frequency domain energy is calculated based on the upper and lower limits of the frequency band related to the emotion characteristics. , the calculation formula is: ,in Represents the power spectrum density of the i-th audio signal, which is the component unit of frequency domain energy. and is the upper and lower limits of the frequency band related to the emotional characteristics (determined according to the specific emotional characteristics, such as sadness is usually concentrated in the low frequency); when obtaining the frequency domain energy values ​​of all n samples of the emotional type {E1, E2, ..., E n}, calculate the frequency domain energy mean μ and standard deviation σ of the sample set; select the confidence level as 95%, determine the corresponding confidence coefficient Z≈1.96, and thus construct the confidence interval of the frequency domain energy value [μ-Z·σ, μ+Z·σ]; use this confidence interval as the frequency domain energy threshold of the current emotion type, which is used for a posteriori verification or auxiliary judgment of the recognition result.

[0079] That is to say, the frequency domain energy thresholds of the present invention are all calculated relative to a specified frequency domain range. This frequency range is determined according to different emotional characteristics and has a clear physical meaning (such as low frequency - sadness, bright frequency - happiness, etc.), thereby enhancing the pertinence and effectiveness of the discrimination.

[0080] In addition, the normalization effect can also be considered to eliminate the influence of different audio intensities on the energy results. The formula is used to normalize the frequency domain energy value. The normalized result of the frequency domain energy value represents the energy concentration ratio in the overall spectrum and is used for comparison with the frequency domain energy threshold, which is suitable for horizontal comparison between different audio samples.

[0081] The frequency domain energy threshold range can be refined and adjusted based on actual applications using manual debugging and experimental data to ensure that different emotion types can be effectively distinguished based on frequency domain energy. Therefore, by providing a clear frequency domain energy threshold for each emotion type, this application can provide more accurate emotion recognition in practical applications such as intelligent customer service, sentiment analysis, and music recommendations, further improving the user experience and the level of intelligent system response.

[0082] Step S140: obtaining an original audio signal, extracting its time domain features and frequency domain features, and obtaining the emotion type of the original audio signal using an emotion classification model;

[0083] When step S130 is executed, step S140 further includes: calculating the frequency domain energy of the audio signal from the frequency domain features, and comparing the frequency domain energy with a frequency domain energy threshold to determine whether the emotion type of the audio signal needs to be reviewed.

[0084] Audio signals include but are not limited to music, voice, etc., and cover a variety of emotional labels such as anger, happiness, sadness, etc.

[0085] For the original audio signal to be classified, the system first outputs the emotion category and its confidence based on the deep emotion classification model. At the same time, it extracts the frequency domain energy within the upper and lower limits of the frequency band related to its emotion characteristics and compares it with the pre-established frequency domain energy threshold range for each emotion type. If the frequency domain energy is consistent with the model output type, the identified emotion type does not need to be reviewed, and the recognition result is considered high confidence. If it is inconsistent, the identified emotion type needs to be reviewed, and the model output confidence is dynamically corrected, or a round of review is triggered, user feedback is requested, or the confidence of the model prediction is adjusted to adjust the energy threshold range to ensure the reliability of the output emotion label.

[0086] In addition, in step S110, the audio data set may also include text data extracted from the audio signal, so as to simultaneously use the audio signal and text data for training, so that the emotion classification model can not only learn the correspondence between audio features and emotion types but also learn the correspondence between text content and emotion types; and before step S140, it also includes: obtaining a large number of audio signals, extracting their time domain features and frequency domain features while extracting text data, so as to obtain the corresponding emotion types according to the time domain features, frequency domain features and text data using the emotion classification model, and using the large number of audio signals and their corresponding emotion types as input samples of the emotion mapping module 20. The input samples of the emotion mapping module thus obtained refer to the rules of text emotion expression, so that the accuracy of emotion recognition is higher, and it is ensured that the constructed emotion-audio feature mapping table meets the requirements of overall emotion expression.

[0087] The emotion mapping module 20 is configured to establish an emotion-audio feature mapping table, conversion rules for different emotion types, and conversion rules for different emotion intensities of the same emotion type based on the audio signal and its corresponding emotion type, and generate the initial values ​​of the adjustment amount and adjustment rate of the real-time audio parameters according to the emotion type of the original audio signal, the collection results of the conversion trigger conditions, and the conversion rules of the emotion type. The initial value here refers to the adjustment amount and adjustment rate of the real-time audio parameters that have not yet been adjusted through user feedback. Thus, precise enhancement of emotional expression is achieved through quantitative modeling and collaborative control.

[0088] like Figure 3 As shown, the emotion mapping module 20 is configured to perform the following steps:

[0089] Step S210: constructing an emotion-audio feature mapping table based on a large number of audio signals and their corresponding emotion types;

[0090] In the emotion-audio feature mapping table, the extracted audio features include basic audio parameters such as volume and pitch, as well as timbre (by changing the harmonic structure of the sound to make it sound softer, warmer or colder and sharper), speaking speed (adjusting the speaking speed according to the emotional state), pause pattern (appropriately adjusting the pause duration and rhythm in the speech to make it more in line with natural emotional expression), dynamic range (controlling the dynamic range of the sound (that is, the difference between the minimum and maximum volume)), language energy distribution (by adjusting the energy distribution of the speech in the frequency domain, strengthening certain frequency components, making the sound sound more powerful or softer), and voice envelope (adjusting the fluctuation pattern of the sound to make it more layered), so as to adapt to the needs of different scenarios.

[0091] The step S210 includes:

[0092] Step S211: data collection; that is, taking a large number of audio signals and their corresponding emotion types as input samples;

[0093] Step S212: Feature extraction; that is, extracting audio features from each audio sample.

[0094] Step S213: analysis and association; that is, counting the combinations of audio features of audio samples of different emotion types, and finding a typical combination of audio features corresponding to each emotion type.

[0095] Typical combinations of acoustic features are usually those that appear frequently in statistics. For example, anger corresponds to high volume, fast tempo, and sharp tone.

[0096] Step S214: establishing an emotion-audio feature mapping table; that is, storing typical combinations of emotion types and their corresponding audio features in association, thereby establishing an emotion-audio feature mapping table.

[0097] In this embodiment, the established emotion-audio feature mapping table is shown in Table 1 below.

[0098] Table 1: Emotion-audio feature mapping table

[0099] Emotional Type Audio feature combination Threshold range anger High volume, fast tempo, sharp pitch Volume>80dB, Tempo>120BPM, Pitch>200Hz hapiness Medium volume, medium tempo, bright tone Volume 60-80dB, tempo 80-120BPM, pitch 150-200Hz sad Low volume, slow tempo, low pitch Volume <60dB, tempo <80BPM, pitch <150Hz calm Low volume, slow tempo, steady tone Volume <50dB, tempo <70BPM, pitch <130Hz fear High volume, fast tempo, irregular pitch Volume > 70dB, tempo > 100BPM, large pitch changes excited High volume, fast tempo, high pitch Volume>85dB, tempo>120BPM, high pitch

[0100] In other embodiments, dynamic parameter adjustment based on reinforcement learning may be used to replace the preset emotion-audio feature mapping table.

[0101] Step S220: establishing conversion rules for emotion types, including conversion rules for different emotion types and conversion rules for different emotion intensities of the same emotion type.

[0102] In step S220, conversion rules for different emotion types are established, specifically including:

[0103] Step S221: establishing an emotion conversion matrix with conversion rules for different emotion types, where rows represent original emotion types (i.e., recognized emotion types are used as original emotion types), and columns represent target emotion types, and converting a set of original emotion types and target emotion types into one emotion type;

[0104] Step S222: For each emotion type conversion, a conversion rule for smoothly transitioning from the original emotion to the target emotion (including an adjustment amount and an initial value of the adjustment rate of the audio parameter) is established based on the threshold range of the audio parameters of the original emotion type and the target emotion type in the emotion-audio feature mapping table, and serves as an element corresponding to the emotion type conversion in the emotion conversion matrix;

[0105] For example, when expressing excitement, the tempo is increased rapidly and then gradually stabilized at a higher level at a moderate rate.

[0106] Different transition effects can be used for different types of emotional conversions. The transition effect makes the adjustment rate change differently over time. The adjustment rate is determined based on the adjustment amount of the audio parameters and the transition effect to meet the requirements for the delicateness of emotional conversion in different scenarios.

[0107] In this embodiment, transition effects include linear transition (where the adjustment rate remains constant over time), nonlinear transition (where the adjustment rate changes over time), and step-wise transition (where the adjustment rate suddenly changes over time). For example, in an intelligent customer service system, when a user changes from "angry" to "calm," a nonlinear transition can be used to quickly reduce the volume and tempo before slowly adjusting them back to normal levels. This allows the user to experience the system's rapid response to their emotional changes and gradual soothing.

[0108] In this embodiment, the adjustment rate of the audio parameters is obtained by using a multi-dimensional nonlinear adjustment model. Among them, for volume adjustment, an exponential decay envelope control based on emotion intensity is designed, and its mathematical model can be expressed as:

[0109] V(t)=V0+ΔVmax(1-e -t / τ )

[0110] The time constant τ is set differently based on the emotion type: a rapid 150ms response for happy emotions and a gradual 300ms response for sad emotions. For rhythm adjustment, a dynamic time warping algorithm is introduced to ensure timing alignment accuracy, achieving a certain rhythm pattern matching rate on the test set. Phase compensation technology is also used to keep harmonic distortion below 1.5%. A two-stage pitch adjustment strategy is designed for melancholy: a 2Hz / s down-tune rate for the first 500ms, followed by a 0.5Hz / s fine-tuning phase during the subsequent maintenance phase. This ensures significant emotional expression while avoiding auditory abruptness.

[0111] Step S223: Set a conversion trigger condition for each emotion type conversion. As part of the conversion rules for different emotion types, the conversion trigger condition is used to obtain the initial value of the adjustment amount and adjustment rate of the audio parameter according to the conversion rules of the corresponding emotion type when activated, and then output the audio parameters that change over time to achieve a smooth transition of emotions.

[0112] The triggering conditions for emotional type conversion can be a sudden change in user emotion, an audio preference expressed by the user through evaluation, or a shift in the semantic conversational scene (including triggering events based on context-recognized policy rules and active conversion instructions from the user feedback module 40). This multi-dimensional triggering mechanism dynamically drives the audio emotion expression module to generate emotional audio output that matches the current context, enhancing the system's natural interaction capabilities and the user's emotional resonance experience.

[0113] "Context-based policy rule triggering events" may include the following scenarios: When the system identifies a user as being in a sad mood and continues to play sad-themed audio for longer than a set threshold (e.g., 5 minutes), and if the user's mood does not improve during this period (no transition instructions, no playback skipping, etc.), the system automatically determines, based on preset emotion regulation policy rules, that the user's mood may be at risk of a persistent negative mood, triggering a transition. In such cases, the system gradually adjusts audio output parameters (increasing volume, raising audio baseline, speeding up tempo, etc.) through a gradual adjustment method, allowing the user to subtly transition from a negative to a positive emotional state.

[0114] Among them, the migration of the semantic dialogue scene includes two situations. One is the change of interaction state identified by the speech recognition and semantic understanding module based on semantic intention or context. The other is the active conversion instruction (such as "switch to a certain emotion") obtained by the user feedback module 40. The latter structurally achieves a direct response to the user's intention through event monitoring and semantic analysis, and forms a control logic link, thereby ensuring the controllability and pertinence of the audio emotion switching.

[0115] Establish the conversion rules of different emotion intensities of the same emotion type, including:

[0116] Step S221': Divide each emotion type into multiple emotion intensities;

[0117] For example, "anger" can be divided into emotional intensities such as mild anger, moderate anger, and extreme anger.

[0118] Step S222 ′: for each emotion type, according to the audio samples of the emotion type, obtain the adjustment range of the audio parameters corresponding to the initial emotion intensity and the target emotion intensity under the emotion type.

[0119] For each combination of initial and target emotional intensity within a given emotional type, the corresponding audio parameter adjustment amplitude must be determined. When transitioning from one emotional type to a new one, the initial emotional intensity is typically defaulted to light. If there's no clear information to determine the user's current intensity, the initial emotional intensity is defaulted to light. If the system detects (via the watch sensor) or the user explicitly specifies an intensity, the actual intensity is used. It's important to note that each target emotional intensity doesn't directly correspond to a fixed audio parameter adjustment result (e.g., the absolute range of adjusted volume). Instead, the range between the initial and target emotional intensity corresponds to a fixed audio parameter adjustment amplitude (e.g., the volume adjustment increment ΔV). This refers to the "amplitude" of the volume, not the "absolute value." For example, different volume adjustment amplitudes can be set based on the change in emotional intensity. The higher the emotional intensity, the stronger the emotion, and the larger the volume adjustment amplitude. For example, for "happiness," mild happiness is accompanied by a moderate increase in volume, a slightly faster tempo, and a slightly higher pitch. Extreme happiness, on the other hand, is accompanied by a significant increase in volume, a very fast tempo, and a significantly higher pitch. Even upbeat sound effects can be added to enhance the emotional expression. Volume adjustment adopts relative adjustment strategy (increase or decrease ), the system adjusts the current playback volume And the combination S of the initial emotion intensity and the target emotion intensity, calculate the final output volume . Volume adjustment formula: ,in, The current audio playback volume. The adjustment amplitude of the audio parameter corresponding to the initial emotional intensity and the target emotional intensity can be a positive or negative value.

[0120] In this embodiment, the establishment of the adjustment amplitude of the audio parameters corresponding to different initial emotional intensities and target emotional intensities does not originate from original manual labeling, but is achieved through automatic grading or manual labeling through audio feature analysis during the system construction process.

[0121] Automatic classification specifically includes: clustering or interval division of the audio features (frequency domain energy, tone amplitude, etc.) of samples under the same emotion type to form pseudo labels, such as: "low emotion intensity": the energy distribution is lower than the average value of samples of this type ; "Medium emotional intensity": falls into [ ]; "High emotional intensity": higher than Then, the adjustment range of the audio parameter corresponding to the initial emotional intensity and the target emotional intensity is determined according to the difference between the audio parameters of different emotional intensities.

[0122] Manual labeling: If the training samples are limited or cannot be clustered, the division rules can be preset based on business experience. For example, the intervals of audio parameters corresponding to different emotional intensities can be manually set according to the spectrum peaks, and then the adjustment range of the audio parameters corresponding to the initial emotional intensity and the target emotional intensity can be determined.

[0123] Step S223': for different emotional intensities of the same emotional type, setting the adjustment amplitude and adjustment rate of the audio parameter when switching between different emotional intensities according to the adjustment amplitude of the audio parameter corresponding to the initial emotional intensity and the target emotional intensity;

[0124] Therefore, by determining the change range and change speed of each parameter, the changes are made more natural and smooth.

[0125] Step S224 ′: determining a triggering condition for the conversion of the emotion intensity, and determining a method for dynamically adjusting the audio parameters according to a dynamic temporal change trend of the emotion intensity when the conversion of the emotion intensity is triggered.

[0126] The dynamic adjustment method of audio parameters can be, for example, for a piece of speech, the dynamic change trend of the emotional intensity over time can be that the emotional intensity gradually increases to a peak and then gradually weakens. The system adjusts the audio parameters accordingly according to the dynamic change trend.

[0127] The triggering conditions for emotional intensity conversion include: when the system detects a sudden change in user emotion, the user's audio preferences expressed through evaluation, or a shift in the semantic conversational scene (including a policy rule triggering event based on context recognition, or the user's active conversion instruction in the user feedback module 40). Thus, when emotional intensity conversion is triggered, the emotional intensity conversion phase is triggered. Emotional intensity is divided into multiple discrete levels (e.g., low emotional intensity, medium emotional intensity, and high emotional intensity). The following two types of detection conditions are set to determine whether emotional intensity conversion has been executed. Numerical judgment conditions based on characteristic mutations: The system tracks the changes in frequency domain energy in real time and defines the trigger formula: Intensity rise trigger condition: ; Trigger conditions for strength reduction: .in, , The energy change threshold determined for training can be set separately under different emotion types. Time window condition based on statistical trend: To prevent short-term disturbances from triggering frequent switching, the system introduces a statistical criterion for the sample ratio within the time window: in the latest W frames (such as 4 frames within 1 second), if there is If the intensity level of the frame is higher than the current level, it is determined that the emotional intensity has increased; similarly, if the number of frames below the current level exceeds a certain proportion, it can also be determined that the emotional intensity has decreased.

[0128] Thus, the present invention can determine the amplitude and speed of change for each audio parameter, making the changes more natural and smooth. For example, the amplitude and speed of volume adjustment can be set based on the type and intensity of emotion, making the volume changes more natural and smooth. For rhythm adjustment, the amplitude and speed of rhythm speed can be determined based on the type and intensity of emotion. For pitch adjustment, the amplitude and speed of pitch change can be set based on the type and intensity of emotion. For example, when expressing melancholy, the pitch can be slowly lowered and maintained at a low level for a period of time.

[0129] In the emotion mapping module 20, after establishing the conversion rules for different emotion types, it can also include: establishing a neural network model, wherein the neural network model is used to obtain a combination of adjusted values ​​of audio parameters corresponding to different target emotion types according to the currently input audio signal when user satisfaction is too low, and the difference between the adjusted value of the audio parameter and the initial value of the audio parameter is the adjustment amount of the audio parameter, thereby obtaining the real-time adjustment amount of the audio parameter and the initial value of the adjustment rate. The real-time adjustment amount of the audio parameter and the initial value of the adjustment rate refer to the values ​​that have not yet been adjusted through real-time user feedback.

[0130] Thus, a combination of adjustment amounts of audio parameters can be generated based on the relationship between the adjustment amounts of multiple audio parameters learned by the neural network model, thereby avoiding the adverse effects of single feature adjustment on the overall audio effect. For example, while increasing the volume, the pitch and rhythm are appropriately adjusted to make the audio sound more coordinated and consistent, and better meet the needs of emotional expression. Among them, the emotion intensity conversion rule defined in step S220 can be used for traditional rule-driven audio adjustment logic, and at the same time serve as the basis for parameter label generation in the neural network model training process here; and, the emotion intensity conversion rule provides interpretable boundary constraints and abnormal adjustment strategies after the combination of the adjusted numerical values ​​of the model output, forming a dual guarantee mechanism for emotion adjustment.

[0131] Specifically, the neural network model takes an audio signal as input and outputs a combination of the target emotion type and the adjusted values ​​of the corresponding audio parameters. During training, the neural network model automatically extracts and learns various audio features, including volume, rhythm, pitch, spectrum, and other features, and understands the relationship between these and emotion types. In this way, the neural network model generates a combination of adjusted audio parameter values ​​that are overall consistent for different target emotion types, thereby comprehensively considering the coordination of multiple audio parameters during audio adjustment.

[0132] In this embodiment, a multi-output neural network model is constructed and trained. The model takes the audio signal as input and jointly outputs two sub-results: the label of the target emotion type and the target emotion type. ; The audio parameter combination vector corresponding to each target emotion type: , They represent the adjusted values ​​of audio parameters such as volume, pitch, and spectrum shape.

[0133] When obtaining training samples, first, the audio parameter combination vector corresponding to the recommended target emotion type can be fine-tuned and screened in combination with expert listening adjustment and user feedback scores, and the audio parameter combination vector corresponding to the high-quality target emotion type whose subjective satisfaction and physiological indicator response are consistent with the target emotion expression can be selected as the final adjustment value; then, the final satisfaction adjustment value and the corresponding user feedback score of the user in actual use can be recorded as the audio parameter combination vector corresponding to the "user-approved" target emotion type, so as to further correct the audio parameter combination vector corresponding to each target emotion type.

[0134] That is to say, when obtaining training samples, the audio parameter combination vectors corresponding to each target emotion type mainly include the following two categories: one is the audio parameter combination vectors that are finally confirmed by the user or evaluated by the system as satisfactory by recording the user's feedback data on audio adjustment during the actual interaction process, which is used as the training target to achieve personalized customization of the model; the other is the high-quality audio parameter combination vectors debugged by professionals based on their experience in emotional voice design, which are used to construct a high-precision training sample set.

[0135] In this embodiment, the input audio signal Divide into frames and extract frame-level features to form the model input matrix: , T: the number of frames corresponding to the audio duration; d: the dimension of the acoustic features extracted for each frame. The internal structure of the model includes a set of feature encoding modules for temporal modeling and feature fusion of the input sequence. The encoder can be a CNN, Bi-LSTM, or a combination thereof, outputting an embedded representation h of the entire audio. On this basis, the model connects two output branches in parallel: the sentiment classification branch uses a Softmax layer to output the probability distribution of sentiment labels, and the audio parameter regression branch outputs a combined vector of audio parameters through a fully connected layer, thereby achieving joint reasoning. The overall mapping function of the neural network model can be expressed as: in is the weight matrix of the sentiment classification branch, represents the predicted sentiment label probability vector, are the linear mapping coefficients of the audio parameter regression branch, Represents the predicted audio adjustment parameter combination, corresponding to volume, pitch intensity, and spectral tilt or centroid frequency, respectively. All parameters can be used to control the emotional audio adjustment strategy in the playback module. During training, the model is optimized using a joint loss function, where the emotion classification loss is the cross entropy loss and the audio parameter prediction loss is the mean squared error loss. The combination of the two is: in, is the predicted probability of the i-th category emotion, is the encoding of the true emotional label, λ is the loss weight coefficient, and The target parameters and predicted parameter vectors are used to generate a set of adjusted values ​​of audio parameters matching each target emotion type, which are then used to guide the playback engine to generate an audio style consistent with the emotion.

[0136] To further study the coordinated regulation relationship between various audio parameters, the present invention can also establish a multivariate regression model based on the output data of a large number of training samples to indirectly derive the correlation function between various audio parameters, for example: , ,in 、 is the regression fitting parameter, 、 is the residual term, v represents the volume adjustment of the current audio; p represents the rhythm parameter, such as speech speed or music beat speed; Indicates the offset amplitude of the pitch; e is the target emotion intensity value (normalized, 0-1); Indicates the rate of change of volume; 、 are the regression residual values ​​of the volume adjustment amount and its adjustment rate, respectively, which are used to indicate that the model fails to fully fit some error factors. The established correlation function between the audio parameters can replace the above-mentioned neural network model and achieve coordination between the adjustment amounts of various audio parameters on the basis of lightweight design.

[0137] The aforementioned correlation function can serve as an approximate alternative to neural network output, enabling basic emotional audio adjustment in low-resource scenarios. The established correlation function between audio parameters enables the system to automatically generate consistent combinations of adjustment amounts for multiple audio parameters based on the established correlation function and the target emotion category. This ensures that the audio output is consistent with the identified emotional state in multiple dimensions, such as volume, speech rate, intonation, and spectrum, ultimately achieving a more natural and resonant sound expression.

[0138] According to the above correlation function, the adjusted values ​​of audio parameters such as volume, pitch, and spectrum shape The calculation formula can be:

[0139] Volume adjustment amount: , is the regression fitting parameter, p represents the rhythm parameter, such as speech speed or music beat speed, represents the shift amplitude of the pitch, and e is the target emotion intensity value (normalized, 0-1);

[0140] Tone adjustment amount: , is the default base frequency, is the tone emotion adjustment gain, f is the adjusted tone frequency, and e is the target emotion intensity value (normalized, 0-1);

[0141] Spectral shape: ,in, is the original spectrum energy distribution, is the spectrum shape after adjustment, The spectrum weighting function related to the target emotion category assigns different enhancement coefficients to different frequency bands. For example, for high-frequency enhancement, it can be set as ,in, is the gain coefficient, The upper limit frequency of the audio bandwidth.

[0142] In addition, the present invention also provides multimodal learning, that is, combining audio with other modal information (such as text, video, etc.) to understand the comprehensive effect of multimodal features in emotional expression and learn multimodal features other than audio under different emotional types. Specific references [Zhou Ping. Music emotion classification algorithm based on multimodal deep learning [J]. Intelligent Computers and Applications, 2022, 12(9):5].

[0143] Step S230: Generate the initial value of the adjustment amount and adjustment rate of the real-time audio parameter according to the emotion type of the original audio signal, the collection result of the conversion trigger condition, and the conversion rule of the emotion type.

[0144] Furthermore, the emotion mapping module 20 is configured to display the adjustment rate curves and constraint boundaries of each audio parameter on a real-time monitoring interface during steps S210 and S220 for intuitive observation by the operator. The constraint boundaries represent the boundaries of typical combinations of audio features corresponding to each emotion type. If necessary, user-specific data (such as user adjustment preferences) can be imported through the user feedback module 40 to adjust the adjustment rate curves and constraint boundaries of the audio parameters, achieving a smooth transition from standard emotion to personalized expression.

[0145] Therefore, users can import user personalized data based on the real-time monitoring interface and the user feedback module 40 to develop an emotional guidance model suitable for psychotherapy. For example, by adjusting the curve of the adjustment rate of each audio parameter, the gradual change of the emotional type can be further achieved, especially for special groups such as autistic patients. Combining emotional visualization with an interactive interface expands the application value of the system in the fields of medicine and rehabilitation.

[0146] In some other embodiments, the emotion mapping module 20 can be omitted, and the audio processing module 30 directly obtains the real-time audio parameter adjustment amount and the initial value of the adjustment rate (i.e., it has not yet been adjusted through user feedback) based on the identified emotion type and the pre-defined emotion type conversion rules, so as to make real-time adjustments to the audio volume, pitch, rhythm and other parameters.

[0147] like Figure 3 As shown, the audio processing module 30 includes an audio parameter adjustment unit 31 and a dynamic adjustment algorithm unit 32 .

[0148] The audio parameter adjustment unit 31 is configured to process the original audio signal according to the real-time audio parameter adjustment amount and adjustment rate, thereby adjusting the time domain and frequency domain characteristics of the audio to ensure that the audio playback effect can accurately meet the needs of emotional expression.

[0149] The audio parameter adjustment unit 31 includes an audio filter, a volume controller and a transmission. The audio parameter adjustment unit 31 uses an audio filter to perform frequency domain processing on the audio signal, adjusts the frequency distribution of the audio according to the adjustment amount and adjustment rate of the tone output by the emotion mapping module, and enhances or weakens the energy of a specific frequency band to achieve tone adjustment. The volume controller adjusts the volume of the audio in real time, and accurately controls the volume change according to the adjustment amount and adjustment rate of the volume determined by the emotion mapping module. The transmission is used to change the playback speed of the audio, thereby adjusting the rhythm of the audio, while ensuring that the sound quality of the audio is not affected.

[0150] The dynamic adjustment algorithm unit 32 uses a dynamic compression algorithm and an expander algorithm to process the audio output by the audio parameter adjustment unit 31 to ensure a smooth transition of the audio output.

[0151] In this embodiment, the dynamic adjustment algorithm unit uses a dynamic compression algorithm to compress the dynamic range of the audio signal, making the audio volume change smoother and avoiding the situation where the volume is too loud or too soft. In particular, when the emotional expression is relatively strong, the algorithm can effectively control the peak value of the volume and prevent audio distortion. The dynamic adjustment algorithm can be found in the literature [Giannoulis D, Massberg M, Reiss JD. Digital Dynamic Range Compressor Design-A Tutorial and Analysis [J]. Journal of the Audio Engineering Society, 2012, 60 (6): p. 399-408]. In addition, the expander algorithm is used to expand the dynamic range of the audio signal, enhance the dynamic effect of the audio, make the audio details richer, and make the emotional expression more delicate. For example, when expressing soothing emotions, the expander can appropriately increase the dynamic range of the audio to make the audio sound softer and more natural. The expander algorithm can be found in the literature [Press F. Mastering audio: the art and the science third edition [J]. 2025-05-07].

[0152] As a result, the audio parameter adjustment unit 31 can make precise real-time adjustments to various audio parameters based on the adjustment amount and adjustment rate of the audio parameters determined by the emotion mapping module. For example, when playing audio with tense emotions, the audio parameter adjustment unit will quickly increase the volume and speed up the rhythm, while the dynamic adjustment algorithm unit 32 ensures that the volume and rhythm increase process is smooth and natural, allowing users to better feel the tense atmosphere conveyed by the audio. In addition, the module effectively avoids problems such as sound quality distortion and parameter mutations that may occur during the audio adjustment process by using professional audio processing tools and dynamic processing algorithms, thereby improving the quality and stability of audio processing and further enhancing the user's listening experience.

[0153] The user feedback module 40 is configured to collect personalized data of the user, and provide the collection results as conversion trigger conditions to the emotion mapping module 20, and provide them to the audio parameter adjustment unit 31 to dynamically adjust the adjustment amount and adjustment rate of the audio parameters, thereby adjusting the audio output.

[0154] In some other embodiments, the emotion mapping module 20 can be omitted, and the user feedback module 40 is configured to collect the user's personalized data and provide it to the audio parameter adjustment unit 31 to dynamically adjust the adjustment amount and adjustment rate of the audio parameters, thereby adjusting the audio output.

[0155] The user feedback module 40 is further configured to activate the conversion trigger condition in real time through a sudden change in the user's emotion, the user's audio preference expressed through evaluation means, or the migration of the semantic dialogue scene, so as to convert the audio signal into a specific emotion.

[0156] Among them, the migration of the semantic dialogue scene includes two situations. One is the change of interaction state identified by the speech recognition and semantic understanding module based on semantic intention or context. The other is the active conversion instruction (such as "switch to a certain emotion") obtained by the user feedback module 40. The latter structurally achieves a direct response to the user's intention through event monitoring and semantic analysis, and forms a control logic link, thereby ensuring the controllability and pertinence of the audio emotion switching.

[0157] like Figure 4 As shown, the user feedback module 40 includes a physiological feedback acquisition unit 41 and an interactive interface feedback unit 42. The personalized data includes physiological data collected by the physiological feedback acquisition unit 41 and user evaluations and preferences collected by the interactive interface feedback unit 42, as well as user dialogue scenarios.

[0158] In this embodiment, the physiological feedback acquisition unit 41 includes a heart rate sensor and accelerometer built into a mobile phone, and a galvanic skin response sensor from a smartwatch. The user feedback module 40 uses the physiological feedback acquisition unit 41 to monitor the user's heart rate, activity data, and galvanic skin response in real time as physiological data, and analyzes the changing trends and characteristics of this physiological data to determine the user's emotional state. For details, see [Emotion Detection Based on Wearable Devices, https: / / www.docin.com / p-4659553527.html] and [Shukla J, Barreda-Angeles M, Oliver J, et al. Feature Extraction and Selection for Emotion Recognition from Electrodermal Activity[J]. IEEE Transactions on Affective Computing, 2019:1-1. DOI:10.1109 / TAFFC.2019.2901673]. In other embodiments, the physiological feedback acquisition unit 41 may employ an emotion recognition device based on electroencephalogram (EEG).

[0159] The interactive interface feedback unit 42 provides an interactive interface for the user to collect the user's evaluation and preferences for the audio. The interactive interface feedback unit 42 uses a dedicated application interface located on the user's mobile phone or smart watch. In the dedicated application interface, the user can evaluate the intensity of the audio emotional atmosphere, the appropriateness of the volume, the speed of the rhythm, etc. through operations such as clicking and sliding; at the same time, the dedicated application interface also provides a user preference setting function, allowing the user to set the desired transition effect for different emotional types of audio in advance according to their preferences, as well as the real-time recording of the adjustment amount and adjustment rate of the audio parameters as user preferences. The user's evaluation and preferences will be transmitted to the processing system in real time as an important basis for adjusting the adjustment amount and adjustment rate of the audio parameters.

[0160] In order to enable the audio parameter adjustment unit 31 to dynamically adjust the adjustment amount and adjustment rate of the audio parameters by using the user feedback module 40 to collect the user's personalized data, the audio parameter adjustment unit 31 is configured to execute the following global optimization algorithm:

[0161] Step S310: Receive personalized data collected by the user feedback module 40 (including physiological data collected by the physiological feedback collection unit 41 and user evaluations and preferences collected by the interactive interface feedback unit 42) and historical adjustment behaviors (including initial values ​​of the real-time audio parameter adjustment amount and adjustment rate, and the audio parameter adjustment amount and adjustment rate finally determined by the user) as a set of personalized adjustment samples;

[0162] The initial values ​​of the adjustment amount and adjustment rate of the real-time audio parameter refer to the adjustment amount and adjustment rate of the real-time audio parameter before adjustment by the user feedback module 40 .

[0163] The optimization algorithm uses personalized data collected by the user feedback module 40 as its core input information, including physiological data (such as heart rate, skin conductance, and pulse rate) acquired by the physiological feedback acquisition unit 41, and user evaluations and preferences (including user ratings of audio tempo, volume, and pitch, selective feedback, and historical interaction records) collected by the interactive interface feedback unit 42. Historical adjustment behavior includes audio adjustment behavior data accessed by the system in real time, including the adjustment amplitude (such as the absolute change in tempo or volume increase) and adjustment rate (i.e., the speed of change of audio parameters controlled by user operations), as well as the parameter settings ultimately accepted or confirmed by the user. This data constitutes the set of personalized adjustment samples required for model training.

[0164] Step S320: Clean, normalize, and fill missing values ​​in the set of personalized adjustment samples. A fused feature vector is constructed based on the personalized data from the historical adjustment behavior and the real-time audio parameter adjustment amount and adjustment rate as input. The final audio parameter adjustment amount and adjustment rate determined by the user in the historical adjustment behavior is output as the output, making it suitable for subsequent analysis and modeling.

[0165] To unify the feature representation, the system uses a weighted fusion strategy to construct a fused feature vector, in which each type of original feature is assigned a corresponding weight according to its importance in a specific scenario. The fused feature vector can be expressed as: ,

[0166] in represents the i-th original feature (such as heart rate value, rhythm feedback score, volume adjustment trajectory), The corresponding weight factor can be determined according to empirical rules or dynamically optimized by sorting the feature importance during model training. For example, if user feedback data is more important in the current scenario, it will be given a higher weight in the feature vector.

[0167] Therefore, during the training process, the model input is the fused feature vector, and the output is the audio parameter adjustment amount and adjustment rate finally confirmed by the user; for example, the audio parameter adjustment amount and adjustment rate include the rhythm adjustment amplitude , volume adjustment range , volume adjustment rate , then the output vector is recorded as: .

[0168] Step S330: constructing the input layer, hidden layer, and output layer of the neural network model, and using samples for training to obtain the regularity between the personalized data and the adjustment amount and adjustment rate of the audio parameters;

[0169] When the dimension of the feature vector is d, the number of neurons in the input layer of the neural network model is set to d. The neural network model described in the present invention adopts a standard feedforward architecture, including an input layer, at least one hidden layer, and an output layer. The hidden layer can optionally be a fully connected layer with a ReLU or Tanh activation function, and the output layer uses a linear activation function to output continuous-valued parameters.

[0170] The model training uses supervised learning, and the loss function uses mean square error (MSE): Where m is the output dimension, The values ​​of the adjustment parameters confirmed by users in the actual sample. The optimization algorithm can use Adam or SGD. The model training uses batch gradient descent to update the weights, iterating until the validation set error converges or meets the preset performance indicators.

[0171] Step S340 (optional): After training is completed, perform model evaluation and optimization.

[0172] After training, the system uses an independent validation set to evaluate the generalization ability of the neural network model to ensure its stability and practical application value on non-training data. The validation process is based on a comprehensive evaluation of multiple metrics, including but not limited to mean squared error (MSE), mean absolute percentage error (MAPE), dynamic time warping (DTW), and Pearson correlation coefficient.

[0173] The mean absolute percentage error is calculated as follows: , where m represents the number of validation samples, and ε is a very small constant (such as 10 -8 If the model predicts a time series parameter, the system will further calculate the predicted sequence With the real sequence Dynamic Time Warping (DTW) distance: , where W represents the set of all possible time alignment paths. In addition, the system can also calculate the Pearson correlation coefficient between two sequences: ,in 、 are the means of the predicted value and the true value, respectively.

[0174] Based on the evaluation results, the structure and parameters of the model are adjusted and optimized, such as increasing or decreasing the number of hidden layer neurons, adjusting the learning rate, etc., to improve the performance of the model.

[0175] Step S350: Subsequently, the trained neural network model is used to predict the audio parameter adjustment amount and adjustment rate ultimately determined by the user based on the real-time personalized data, the real-time audio parameter adjustment amount and adjustment rate, so as to dynamically adjust the audio parameter adjustment amount and adjustment rate.

[0176] Therefore, the audio parameter adjustment unit 31 is configured to execute a global optimization algorithm based on personalized data and historical adjustment behavior to dynamically predict and adjust the adjustment amplitude and rate of the audio output parameters. During the deployment phase, the audio parameter adjustment unit receives real-time information about the user's current physiological state and interactive feedback. After feature fusion, it is input into the trained neural network model. The model predicts the adjustment amplitude and rate of change of the target audio parameter, driving the downstream audio playback control module to perform dynamic adjustments, thereby achieving highly consistent and responsive emotional audio matching. For example, if the system detects an increase in the user's current skin conductance index (indicating a high arousal state), i.e., physiological feedback indicates that the user is in a state of excitement, while interface feedback indicates that the user considers the tempo too fast and assigns a negative rating, the system will output a balanced and modulated tempo parameter combination that reduces the tempo adjustment rate while maintaining a smaller tempo increase. This achieves dual compatibility between emotional state and subjective preferences, ensuring that the audio output is consistent with the user's current physiological and emotional state and satisfies the user's subjective preferences.

[0177] Therefore, the present invention integrates physiological data and user evaluations and preferences as historical adjustment behaviors for training, thereby achieving a fusion of physiological data, user evaluations, and preferences to comprehensively judge the user's emotional needs and audio adjustment direction, and provide a customized emotional audio adjustment solution. Through continuous optimization of machine learning algorithms, it can adapt to user emotional changes, improve the understanding of and response accuracy of user emotional needs in the long term, and ensure that the audio system can better adapt to the user's personalized needs. Compared with traditional emotional audio systems, this innovative mechanism avoids fixed emotional models and static parameter settings, making the system more intelligent and adaptable in the long term.

[0178] In summary, the user feedback module 40 can comprehensively collect personalized user data from both physiological and psychological perspectives, thereby more accurately understanding the user's perceptions and needs for audio emotional effects. This method dynamically adjusts the adjustment amount and rate of audio parameters by collecting and processing personalized user data in real time, making the output audio playback more aligned with the user's real-time emotional needs, significantly improving the audio system's adaptability and user experience satisfaction. Furthermore, the user feedback module 40 can utilize devices used daily by users (such as mobile phones or smartwatches) to collect physiological data and user evaluations and preferences, eliminating the need for additional physiological sensors. This reduces costs and user barriers to entry, and improves the system's practicality and scalability.

[0179] Corresponding to the emotional resonance audio system described above, the present invention further provides an emotional resonance audio method, comprising:

[0180] Step S100: using the emotion recognition module 10 to obtain the recognized emotion type according to the original audio signal;

[0181] The step S100 includes:

[0182] Step S110 , obtaining an audio signal, extracting time domain and frequency domain features of the audio signal, and labeling the emotion category as an audio data set.

[0183] Step S120 , pre-training the deep learning model using the audio data set to obtain a sentiment classification model, wherein the sentiment classification model is used to output a corresponding sentiment type according to the time domain and frequency domain features of the audio signal.

[0184] In step S130, statistics are performed based on the frequency domain features extracted in step S110 and the emotion type of each frequency domain feature to obtain the frequency domain energy of the audio signal of each emotion type (such as anger, happiness, sadness, etc.), and accordingly set the corresponding frequency domain energy threshold for each emotion type.

[0185] Step S140: Obtain the original audio signal, extract its time domain features and frequency domain features, and use the emotion classification model to obtain the emotion type of the original audio signal; and obtain the frequency domain energy of the audio signal from the frequency domain feature statistics, and compare it with the frequency domain energy threshold to obtain the emotion type of the original audio signal.

[0186] Step S200: Using the emotion mapping module 20, based on the audio signal and its corresponding emotion type, establish an emotion-audio feature mapping table, conversion rules for different emotion types, and conversion rules for different emotion intensities of the same emotion type; and generate real-time audio parameter adjustment amounts and adjustment rates based on the emotion type of the original audio signal, the collected results of the conversion trigger conditions, and the emotion type conversion rules;

[0187] The step S200 includes:

[0188] Step S210: constructing an emotion-audio feature mapping table based on a large number of audio signals and their corresponding emotion types;

[0189] Step S220: establishing conversion rules for different emotion types and conversion rules for different emotion intensities of the same emotion type.

[0190] The user feedback module 40 is used to collect the user's personalized data, and the collected data is provided to the emotion mapping module 20 as the conversion trigger condition.

[0191] Step S230: Generate a real-time audio parameter adjustment amount and adjustment rate according to the emotion type of the original audio signal, the collection result of the conversion trigger condition, and the conversion rule of the emotion mapping module.

[0192] Step S300: Using the audio parameter adjustment unit 31 to process the original audio signal according to the real-time audio parameter adjustment amount and adjustment rate, thereby adjusting the time domain and frequency domain characteristics of the audio to ensure that the audio playback effect can accurately meet the requirements of emotional expression;

[0193] The user feedback module 40 is used to collect the user's personalized data and provide it to the audio parameter adjustment unit 31 to dynamically adjust the adjustment amount and adjustment rate of the audio parameters, thereby adjusting the audio output.

[0194] The personalized data includes physiological data collected by the physiological feedback collection unit 41 and user evaluations and preferences collected by the interactive interface feedback unit 42, as well as user dialogue scenarios. The user feedback module 40 collects the user's personalized data and provides it to the audio parameter adjustment unit 31 to dynamically adjust the adjustment amount and adjustment rate of the audio parameters. Specifically, the audio parameter adjustment unit 31 executes the following global optimization algorithm:

[0195] Step S310: receiving personalized data collected by the user feedback module 40 (including physiological data collected by the physiological feedback collection unit 41 and user evaluations and preferences collected by the interactive interface feedback unit 42), the real-time audio parameter adjustment amount and adjustment rate, and the audio parameter adjustment amount and adjustment rate finally determined by the user as historical adjustment behavior;

[0196] Step S320: The collected historical adjustment behaviors are cleaned and normalized to make them suitable for subsequent analysis and modeling. Based on the importance of the personalized data and the real-time audio parameter adjustment amount and adjustment rate, different weights are assigned to them to combine and generate a feature vector. If user feedback data is more important in the current scenario, it is given a higher weight in the feature vector.

[0197] Step S330: Constructing the input layer, hidden layer, and output layer of the neural network model. The model uses the personalized data from the historical adjustment behavior and the real-time audio parameter adjustment amount and adjustment rate as input data, and uses the audio parameter adjustment amount and adjustment rate determined by the user in the historical adjustment behavior as output data. Training is performed using a gradient descent method to obtain a regularity between the personalized data and the audio parameter adjustment amount and adjustment rate.

[0198] The number of neurons in the input layer is determined by the fused feature dimensions, while the number of neurons in the output layer is determined by the number of audio parameter adjustment amplitudes and rate values ​​to be predicted. During training, the model's prediction error is measured using the mean squared error between the model's predicted output and the actual label. Training is iterated until satisfactory model performance is achieved.

[0199] Step S340: After training is complete, model evaluation and optimization are performed. The trained model is evaluated using an independent validation set to check its generalization and prediction accuracy. Based on the evaluation results, the model structure and parameters are adjusted and optimized, such as increasing or decreasing the number of hidden layer neurons and adjusting the learning rate, to improve model performance.

[0200] Step S350: Subsequently, the trained neural network model is used to predict the audio parameter adjustment amount and adjustment rate ultimately determined by the user based on the real-time personalized data, the real-time audio parameter adjustment amount and adjustment rate, so as to dynamically adjust the audio parameter adjustment amount and adjustment rate.

[0201] In addition, step S300 further includes: utilizing the dynamic adjustment algorithm unit 32 to process the audio output by the audio parameter adjustment unit 31 using algorithms such as dynamic compression and expansion to ensure smooth transition of the audio output.

[0202] like Figure 5 As shown, the present invention also provides an audio device, which includes a microphone, an audio processor and a chip, on which an input audio processing module and the emotional resonance audio system described above are installed.

[0203] In this way, a collaborative working relationship between the integrated hardware devices and software systems is achieved. The hardware devices (i.e. microphones, audio processors and chips) provide the necessary execution environment and data input for the software system, while the software system (i.e. input audio processing module and emotional resonance audio system) provides intelligent processing and algorithm support to jointly complete complex tasks.

[0204] In this embodiment, the intelligent customer service system receives voice instructions through a microphone and transmits raw audio data to the audio processing module 30 .

[0205] The input audio processing module is configured to perform preliminary processing on the audio signal (such as noise reduction, echo cancellation, etc.), and then transmit the processed data to the emotional resonance audio system.

[0206] The specific structure and functions of the emotional resonance audio system are as described above.

[0207] The chip is also integrated with a speech recognition and semantic understanding module. The speech recognition and semantic understanding module is located on the chip of the audio processor and can be integrated with the input audio processing module of the audio processor and the emotional resonance audio system described above, and embedded in the chip together to form a hardware audio processor as a whole. The speech recognition and semantic understanding module is mainly used to obtain semantic information in the user's voice, that is, to convert the user's voice commands into text or understandable semantics such as instructions. It can also obtain some feature information of the audio signal, such as the user's voice features, intonation, speaking speed, etc., in order to further optimize the recognition accuracy and the response effect of the system.

[0208] The software system, in turn, provides algorithmic support for the hardware devices, optimizing their efficiency. For example, based on the data features generated by the audio processing algorithm, machine learning models can be used to train and optimize the speech recognition and semantic understanding modules, improving the response accuracy and fluency of the intelligent customer service system. The audio emotional resonance audio system can improve the audio effect of audio signal processing based on historical adjustment behavior, thus achieving deep integration and collaborative work between hardware and software to meet the diverse needs of users.

[0209] Furthermore, the audio device can also include supporting smart home devices to achieve linkage with these devices. These smart home devices include lighting systems, air conditioning systems, and other devices. By pre-defining parameter values ​​for different types of emotional conversions for smart home devices, the emotional resonance audio system can adjust environmental parameters such as lighting and temperature based on the type of emotional conversion, providing a multi-sensory emotional resonance experience and enhancing the user's immersion and emotional comfort.

[0210] The smart home device is in communication with the user feedback module 40 to receive the user's personalized data (such as evaluation and preferences), and further optimizes the parameter values ​​of the smart home device when converting different types of emotion types based on the personalized data.

[0211] Example 1: Emotional Resonance Audio System

[0212] In this embodiment, the core of the emotional resonance audio system lies in the collaborative processing mechanism of emotion recognition and emotion mapping. When an audio signal is input into the system, the emotion recognition module first analyzes the audio signal, extracting its time and frequency domain features, such as amplitude, frequency, and fundamental frequency. Using these features, the emotion recognition module utilizes a pre-trained emotion classification model to accurately identify the emotion type in the audio signal, such as anger, anxiety, or satisfaction. After emotion recognition is complete, the emotion mapping module intervenes and, based on the identified emotion type, retrieves corresponding acoustic parameter adjustment rules from the emotion-audio feature mapping unit. These rules are trained based on a standard emotional speech database and can map emotion types to specific audio parameter adjustments. For example, happiness corresponds to characteristics such as concentrated high-frequency energy, an upward shift in fundamental frequency, and an accelerated tempo. Based on these rules, the emotion mapping module dynamically adjusts audio parameters such as volume, pitch, and tempo to enhance the audio's emotional expression. Based on the output of the emotion mapping module, the audio processing module adjusts various audio parameters in real time to ensure that the audio playback effect accurately meets the requirements of emotional expression. Finally, the user feedback module collects user feedback information, such as physiological data and evaluation of the emotional effect of audio. Based on this feedback, the system further optimizes the audio output and improves the user experience.

[0213] Example 2: Emotional Resonance Audio System

[0214] In this embodiment, the emotional resonance audio system focuses on the integrated optimization mechanism of audio processing and user feedback. After the audio signal is input into the system, the emotion recognition module also performs emotion recognition, extracts audio features, and identifies the emotion type. However, after emotion recognition is complete, the audio processing module directly adjusts audio parameters such as volume, pitch, and tempo in real time based on the identified emotion type and pre-defined conversion rules. The audio processing module is equipped with professional audio processing tools and dynamic processing algorithms, effectively preventing problems such as sound quality distortion and parameter abrupt changes that may occur during the audio adjustment process. Simultaneously, the user feedback module comes into play, collecting user physiological data and evaluations of the audio's emotional effects through a physiological feedback acquisition unit and an interactive interface unit. The system integrates this user feedback with the audio processing process and further optimizes the audio parameters through a dynamic adjustment algorithm unit to better meet the user's personalized needs. For example, if a user's physiological feedback indicates excitement, but the interactive interface feedback indicates that the audio tempo is too fast, the system will appropriately slow down the audio tempo to ensure that the audio output is consistent with the user's physiological and emotional state and meets the user's subjective preferences. Through this fusion optimization mechanism of audio processing and user feedback, the system can adjust the audio output in real time, significantly improving the adaptability of the audio system and the satisfaction of the user experience.

[0215] Through innovative emotion mapping and dynamic audio parameter adjustment mechanisms, combined with deep learning models and user physiological feedback data, the present invention can adjust audio in real time and accurately in multiple dimensions, provide a personalized and natural audio experience, and perform adaptive adjustments based on the user's immediate emotional changes.

[0216] First, the present invention is based on a deep learning-based emotional audio recognition model, which extracts time and frequency domain features from audio signals to perform emotion recognition. Unlike traditional recognition methods based on static emotion labels, the present invention introduces a hybrid model combining CNN and LSTM, which can analyze emotional fluctuations in audio signals in real time, significantly improving the accuracy and robustness of emotion recognition. Traditional solutions usually rely on simple emotion classification or emotion labels based on manual annotation, which are difficult to handle complex and varied emotional expressions. Through the application of deep learning models, the present invention can more accurately capture subtle emotional changes.

[0217] Secondly, the present invention proposes an emotion mapping module for generating emotion conversion rules. Traditional audio processing solutions usually rely only on simple volume and pitch adjustments without involving emotion types. By combining the identified emotion types, the system can accurately and dynamically adjust the volume, pitch, rhythm and other dimensions of the audio in real time according to the emotion conversion rules to ensure that the emotional expression of the audio is always natural and smooth. Especially when emotions change drastically, the system can avoid drastic fluctuations in audio parameters, improve the listening experience, and prevent distortion caused by over-adjustment.

[0218] Furthermore, the present invention innovatively integrates a physiological feedback acquisition unit. By using smart devices (such as watches and smart headphones) to monitor the user's physiological data (such as heart rate and galvanic skin response) in real time, the present invention combines this data with the user's evaluation preferences for comprehensive analysis, dynamically adjusting the audio playback effect. Unlike traditional systems that rely solely on emotion recognition, the present invention, through the incorporation of physiological feedback, can more accurately adapt to the user's immediate emotional needs, thereby providing a more immersive and personalized audio experience.

[0219] Furthermore, this invention provides customized emotional audio adjustments based on users' evaluation preferences and physiological responses. Through continuous optimization using machine learning algorithms, it adapts to changing user emotions, improving the accuracy of understanding and responding to users' emotional needs over the long term. Compared to traditional emotional audio systems, this innovative mechanism avoids fixed emotion models and static parameter settings, making the system more intelligent and adaptable over the long term.

[0220] Another innovation of the present invention is the global optimization algorithm implemented by the audio parameter adjustment unit. This algorithm integrates multi-source emotional information from the original audio signal, user feedback, and physiological data, and comprehensively considers the accuracy of audio emotional expression (through the emotion-audio feature mapping table), user comfort (through the user feedback module), sound quality stability (through the dynamic adjustment algorithm unit), and emotional transmission effect (through the rules designed by the emotion mapping module). The present invention can more comprehensively analyze emotional changes, making audio adjustments more precise and adaptable, providing a richer and more delicate emotional experience. It also enables global optimization across multiple dimensions, avoiding excessive emotional exaggeration or loss of sound quality, and ensuring the optimal performance of the audio system.

[0221] The combination of these innovative technologies gives the present emotional audio system significant advantages in emotion recognition, audio adjustment, user experience, and system adaptability. Compared to existing technologies, this system not only improves the accuracy of emotion recognition but also intelligently adjusts to the user's individual needs and physiological responses, ultimately achieving a more natural, immersive, and personalized audio experience.

[0222] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the present invention. Various modifications are possible. Any simple, equivalent changes and modifications made in accordance with the claims and description of the present invention are within the scope of protection of the patent claims. Anything not fully described in this invention is conventional technology.

Claims

1. An emotional resonance audio system, characterized in that: Includes emotion recognition module, audio processing module and user feedback module; The emotion recognition module is configured to obtain a recognized emotion type based on the original audio signal; The audio processing module includes an audio parameter adjustment unit, which is configured to obtain a real-time adjustment amount and adjustment rate of the audio parameter, and process the original audio signal according to the real-time adjustment amount and adjustment rate of the audio parameter; The user feedback module is configured to collect personalized data of the user, and the audio parameter adjustment unit utilizes the personalized data collected by the user feedback module to dynamically adjust the adjustment amount and adjustment rate of the audio parameter.

2. The emotional resonance audio system according to claim 1, characterized in that The audio processing module directly obtains the adjustment amount and initial value of the adjustment rate of the real-time audio parameter according to the recognized emotion type and the predefined emotion type conversion rule; Alternatively, the emotional resonance audio system also includes an emotion mapping module, which is configured to: establish an emotion-audio feature mapping table and an emotion type conversion rule based on the audio signal and its corresponding emotion type, and generate a real-time audio parameter adjustment amount and an initial value of the adjustment rate according to the emotion type of the original audio signal, the collection result of the conversion trigger condition, and the emotion type conversion rule.

3. The emotional resonance audio system according to claim 1, characterized in that The emotion mapping module is configured to perform the following steps: Step S210: constructing an emotion-audio feature mapping table based on a large number of audio signals and their corresponding emotion types; In the emotion-audio feature mapping table, the extracted audio features include volume, pitch, timbre, speaking rate, pause pattern, dynamic range, language energy distribution, and speech envelope; Step S220: establishing conversion rules for emotion types, including conversion rules for different emotion types; In step S220, conversion rules for different emotion types are established, specifically including: Step S221: Establish an emotion conversion matrix to store conversion rules for different emotion types, where rows represent original emotion types, identified emotion types are used as original emotion types, and columns represent target emotion types. A set of original emotion types and target emotion types are converted as one emotion type. Step S222: For each emotion type conversion, a conversion rule for smoothly transitioning from the original emotion to the target emotion is established based on the threshold range of the audio parameters of the original emotion type and the target emotion type in the emotion-audio feature mapping table, and the rule serves as an element corresponding to the emotion type conversion in the emotion conversion matrix; the emotion type conversion rule includes an adjustment amount of the audio parameter and an initial value of the adjustment rate; Step S223: Setting a conversion trigger condition for each emotion type conversion as part of the conversion rules for different emotion types. The conversion trigger condition is used to obtain initial values ​​of the adjustment amount and adjustment rate of the audio parameter according to the corresponding emotion type conversion rule when activated; Step S230: generating initial values ​​of the adjustment amount and adjustment rate of the real-time audio parameter according to the emotion type of the original audio signal, the collection result of the conversion trigger condition, and the conversion rule of the emotion type; The user feedback module is configured to collect personalized data of the user and provide the collected data as a conversion trigger condition to the emotion mapping module.

4. The emotional resonance audio system according to claim 3, characterized in that The conversion rules of the emotion types also include conversion rules of different emotion intensities of the same emotion type; Establish the conversion rules of different emotion intensities of the same emotion type, including: Step S221': Divide each emotion type into multiple emotion intensities; Step S222': for each emotion type, obtaining the adjustment amplitude of the audio parameter corresponding to the initial emotion intensity and the target emotion intensity under the emotion type according to the audio sample under the emotion type; Step S223': for different emotional intensities of the same emotional type, setting the adjustment amplitude and adjustment rate of the audio parameter when switching between different emotional intensities according to the adjustment amplitude of the audio parameter corresponding to the initial emotional intensity and the target emotional intensity; Step S224 ′: determining a triggering condition for the conversion of the emotion intensity, and determining a method for dynamically adjusting the audio parameters according to a dynamic temporal change trend of the emotion intensity when the conversion of the emotion intensity is triggered.

5. The emotional resonance audio system according to claim 3, characterized in that The emotion mapping module is further configured to: during the execution of steps S210 and S220, display the curves and constraint boundaries of the adjustment rates of the audio parameters according to the real-time monitoring interface, wherein the constraint boundaries are the boundaries of the typical combination of audio features corresponding to each emotion type; when necessary, import user personalized data through the user feedback module to adjust the curves and constraint boundaries of the adjustment rates of the audio parameters; and / or In the emotion mapping module, after establishing the conversion rules for different emotion types, it also includes: establishing a neural network model, which is used to obtain a combination of adjusted values ​​of audio parameters corresponding to different target emotion types according to the currently input audio signal when user satisfaction is too low, and the difference between the adjusted value of the audio parameter and the initial value of the audio parameter is the adjustment amount of the audio parameter, thereby obtaining the real-time adjustment amount of the audio parameter and the initial value of the adjustment rate; the input of the neural network model is the audio signal, and the output is a combination of the target emotion type and the corresponding adjusted value of the audio parameter; when obtaining training samples, the audio parameter combination vector corresponding to each target emotion type includes the audio parameter combination vector finally confirmed by the user or evaluated by the system as satisfactory, and the high-quality audio parameter combination vector debugged by professionals based on the experience of emotional voice design.

6. The emotional resonance audio system according to claim 1, characterized in that The user feedback module includes a physiological feedback acquisition unit and an interactive interface feedback unit, and the personalized data includes physiological data collected by the physiological feedback acquisition unit and user evaluations and preferences collected by the interactive interface feedback unit; The physiological feedback acquisition unit includes a heart rate sensor and an acceleration sensor built into a mobile phone and a skin galvanic response sensor of a smart watch; the interactive interface feedback unit adopts a dedicated application interface provided on the user's mobile phone or smart watch.

7. The emotional resonance audio system according to claim 6, characterized in that The audio parameter adjustment unit is configured to execute the following global optimization algorithm: Step S310: Receive personalized data and historical adjustment behaviors collected by the user feedback module as a set of personalized adjustment samples; the historical adjustment behaviors include initial values ​​of the real-time audio parameter adjustment amount and adjustment rate, as well as the audio parameter adjustment amount and adjustment rate finally determined by the user; Step S320: Clean, normalize, and fill missing values ​​in the set of personalized adjustment samples. A fused feature vector is constructed based on the personalized data from the historical adjustment behavior and the real-time audio parameter adjustment amount and adjustment rate as input. The final audio parameter adjustment amount and adjustment rate determined by the user in the historical adjustment behavior is output as the output, making it suitable for subsequent analysis and modeling. Step S330: constructing the input layer, hidden layer, and output layer of the neural network model, and using samples for training to obtain the regularity between the personalized data and the adjustment amount and adjustment rate of the audio parameters.

8. The emotional resonance audio system according to claim 1, characterized in that The audio processing module further includes a dynamic adjustment algorithm unit, which processes the audio output by the audio parameter adjustment unit using a dynamic compression algorithm and an expander algorithm to ensure a smooth transition of the audio output.

9. The emotional resonance audio system according to claim 1, characterized in that The emotion recognition module is configured to perform the following steps: Step S110: Acquire an audio signal, extract time domain and frequency domain features of the audio signal, and label the emotion category as an audio dataset; Step S120: pre-training a deep learning model using an audio data set to obtain a sentiment classification model, wherein the sentiment classification model is used to output a corresponding sentiment type based on time domain and frequency domain features of the audio signal; Step S140: obtaining the original audio signal, extracting its time domain features and frequency domain features, and using the emotion classification model to obtain the emotion type of the original audio signal.

10. The emotional resonance audio system according to claim 9, characterized in that The emotion recognition module is further configured to: Step S130: Statistically analyzing the frequency domain features extracted in step S110 and the emotion type of each frequency domain feature to obtain the frequency domain energy of the audio signal for each emotion type. Based on the frequency domain energy threshold, a corresponding frequency domain energy threshold is set for each emotion type. The frequency domain energy threshold is used to determine whether the emotion type of the audio signal needs to be reviewed. Step S140 further includes: calculating the frequency domain energy of the audio signal from the frequency domain features, and comparing the frequency domain energy with a frequency domain energy threshold to determine whether the emotion type of the audio signal needs to be reviewed.