Audio adaptation method, wearable device, storage medium and computer program product
By collecting biometric data locally on wearable devices to generate emotional state vectors, calculating emotional changes, and adjusting audio data in conjunction with user feedback, the problem of lagging emotional feedback and untimely strategy updates in existing technologies is solved, enabling continuous and timely optimization of audio content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-27
AI Technical Summary
Existing audio adaptation solutions lack the ability to dynamically respond to continuous changes in user status, resulting in delayed emotional feedback and untimely updates to adaptation strategies.
By collecting users' biometric data locally through wearable devices, generating emotional state vectors, calculating emotional changes, and dynamically adjusting audio data based on user feedback, the audio adaptation strategy can be continuously and timely optimized.
Ensuring that music feedback is synchronized with the user's current physiological state solves the problems of delayed emotional feedback and untimely strategy updates, achieving continuous and accurate adaptation of audio content.
Smart Images

Figure CN121751055A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an audio adaptation method, wearable device, storage medium, and computer program product. Background Technology
[0002] With the widespread application of smart terminals and wearable devices, personalized audio services based on users' physiological states are gradually becoming an important direction for improving user experience. Current technologies typically recommend matching playlists or adjust music playback rhythm by collecting users' biometric data such as heart rate, thus achieving personalized audio content adaptation. However, most of these solutions rely on static emotion mapping models and fixed audio adjustment strategies, lacking the ability to dynamically respond to continuous changes in user states. Therefore, existing audio adaptation solutions suffer from lagging emotional feedback and untimely updates to adaptation strategies, making it difficult to achieve continuous and accurate music adaptation when user states change.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide an audio adaptation method, electronic device, storage medium, and computer program product, which aims to solve the technical problems of delayed emotional feedback and untimely updates of adaptation strategies in existing audio adaptation solutions.
[0005] To achieve the above objectives, this application proposes an audio adaptation method for wearable devices, the audio adaptation method comprising: The wearable device collects the user's biometric data at preset time intervals. In the current cycle, a first emotional state vector is generated based on the first biometric data collected in the current cycle; Calculate the change in emotion between the first emotional state vector and the second emotional state vector, wherein the second emotional state vector is generated in the previous adjacent cycle of the current cycle; Based on the emotional change and the received user feedback, the historical audio data of the previous adjacent period is adjusted to obtain target audio data, and the target audio data is output, wherein the target audio data is adapted to the first biometric data.
[0006] Optionally, the first biometric data includes physiological data and voice data, and the step of generating a first emotional state vector based on the first biometric data collected in the current period includes: The physiological emotional component is calculated based on the physiological data, and the corresponding speech emotional component is calculated based on the speech data. The physiological emotional component and the voice emotional component are weighted and summed based on the first preset fusion weight to generate a first emotional state vector.
[0007] Optionally, the steps of calculating the corresponding physiological emotion component based on the physiological data and calculating the corresponding speech emotion component based on the speech data include: The physiological emotional component corresponding to the physiological data is calculated based on the physiological data, preset normalization parameters, and preset weighting coefficients. The physiological data includes heart rate, resting heart rate, and heart rate variability. The speech data is input into a preset neural network to obtain the speech emotion component corresponding to the speech data, wherein the speech data includes speech rate, fundamental frequency and sound pressure level features.
[0008] Optionally, the step of adjusting the historical audio data of the previous adjacent period based on the emotional change and the received user feedback information to obtain the target audio data includes: The first preset fusion weight is updated based on the emotional change amount, and the first preset music parameter mapping table is updated based on the received user feedback information to obtain the second preset fusion weight and the second preset music parameter mapping table. The first preset fusion weight is used to generate the first emotional state vector, the second preset fusion weight is used to generate the third emotional state vector of the next adjacent cycle of the current cycle, and the first preset music parameter mapping table is used to obtain the historical audio data of the previous adjacent cycle. Based on the second preset music parameter mapping table, the first music parameter corresponding to the current period is determined, and the historical audio data is adjusted based on the first music parameter to obtain the target audio data.
[0009] Optionally, the steps of updating the first preset fusion weight based on the emotional change amount and updating the first preset music parameter mapping table based on the received user feedback information include: If the change in the emotional state is inconsistent with the trend of the change in the first biometric data, the first preset fusion weight is reduced. If user feedback is received and the feedback is unsatisfactory, the correlation strength between the second emotional state vector and the second music parameter in the first preset music parameter mapping table is reduced, wherein the second emotional state vector and the second music parameter correspond to the previous adjacent cycle.
[0010] Optionally, the step of adjusting the historical audio data based on the first music parameters to obtain the target audio data includes: Based on the first music parameters, adjust the preset audio synthesizer to generate the target audio data corresponding to the first music parameters; or The audio segment with the highest similarity to the first music parameter is retrieved from the preset music material library, and the audio segment is used as the target audio data corresponding to the music parameter.
[0011] Optionally, the step of outputting the target audio data includes: Upon receiving a music switching signal, the target audio data and the historical audio data are cross-faded, and the historical audio data is switched to the target audio data. The cross-fading process is used to smoothly transition the switching process of the target audio data.
[0012] Furthermore, to achieve the above objectives, this application also proposes an audio adaptation device, which includes: The data acquisition module is used to collect the user's biometric data through the wearable device at preset time intervals; The vector calculation module is used to generate a first emotional state vector based on the first biometric data collected in the current cycle. The difference calculation module is used to calculate the amount of emotional change between the first emotional state vector and the second emotional state vector, wherein the second emotional state vector is generated in the previous adjacent cycle of the current cycle; An audio adjustment module is used to adjust the historical audio data of the previous adjacent period based on the emotional change amount and the received user feedback information to obtain target audio data, and output the target audio data, wherein the target audio data is adapted to the first biometric data.
[0013] In addition, to achieve the above objectives, this application also proposes a wearable device, the wearable device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio adaptation method as described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the audio adaptation method described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the audio adaptation method described above.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: Since this solution is applied to wearable devices, it is implemented locally on the wearable device, eliminating the need for cloud data transmission and remote computing. This eliminates cloud computing and network transmission links, minimizing end-to-end latency and ensuring that music feedback is synchronized with the user's current physiological state, effectively overcoming the problem of delayed emotional feedback. The solution continuously collects the user's biometric data at preset intervals. In the current interval, a first emotional state vector is generated locally based on the collected first biometric data. By comparing the emotional changes of the second emotional state vector with the first emotional state vector, the audio data generation strategy is dynamically adjusted, quantifying emotional changes. Combined with user feedback, it can perceive the continuous changes in the user's state and update the audio adaptation strategy accordingly, solving the adaptation bias problem caused by the rigidity of traditional emotional models. Compared to existing audio adaptation solutions that rely on static emotion mapping models and fixed audio adjustment strategies, this solution eliminates cloud dependence through terminal localization, solving the feedback lag problem caused by data transmission and computation delays. At the same time, by periodically calculating and integrating emotion changes, it breaks the limitations of static model fixation, enabling continuous and timely optimization of audio adaptation strategies. This effectively solves the technical problems of delayed emotion feedback and untimely strategy updates in existing audio adaptation solutions. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the first embodiment of the audio adaptation method of this application. Figure 2 This is a flowchart illustrating the second embodiment of the audio adaptation method of this application. Figure 3 A flowchart illustrating the overall audio adaptation method provided in the second embodiment of this application; Figure 4 This is a schematic diagram of the module structure of the audio adapter device of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the audio adaptation method in the embodiments of this application.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0023] Current mainstream solutions typically recommend matching playlists or adjust the playback rhythm of music by collecting users' biometric data such as heart rate in order to achieve personalized adaptation of audio content. However, they mostly rely on static emotion mapping models and fixed audio adjustment strategies, lacking the ability to dynamically respond to continuous changes in user status.
[0024] This application provides a solution that, since it is applied to wearable devices, is implemented locally on the wearable device. Therefore, it eliminates the need for cloud data transmission and remote computation, thus minimizing end-to-end latency and ensuring that music feedback is synchronized with the user's current physiological state, effectively overcoming the problem of delayed emotional feedback. The solution continuously collects the user's biometric data at preset intervals. In the current interval, a first emotional state vector is generated locally based on the collected first biometric data. By comparing the emotional changes of the second emotional state vector and the first emotional state vector, the audio data generation strategy is dynamically adjusted, quantifying emotional changes. Combined with user feedback, it can perceive the continuous changes in the user's state and update the audio adaptation strategy accordingly, solving the adaptation deviation problem caused by the rigidity of traditional emotional models. Compared to existing audio adaptation solutions that rely on static emotion mapping models and fixed audio adjustment strategies, this solution eliminates cloud dependence through terminal localization, solving the feedback lag problem caused by data transmission and computation delays. At the same time, by periodically calculating and integrating emotion changes, it breaks the limitations of static model fixation, enabling continuous and timely optimization of audio adaptation strategies. This effectively solves the technical problems of delayed emotion feedback and untimely strategy updates in existing audio adaptation solutions.
[0025] It should be noted that the execution subject of each embodiment of the audio adaptation method of this application can be a wearable device with data processing, network communication and program running functions, such as a smart bracelet, headphones, etc., and the wearable device has a built-in or external microphone.
[0026] Reference Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the audio adaptation method of this application.
[0027] In this embodiment, the audio adaptation method is applied to a wearable device, and the audio adaptation method includes steps S10~S40: Wearable devices are hardware entities that integrate multimodal biosensors and data processing units, enabling functions such as local data acquisition, data fusion, emotion modeling, audio generation, and dynamic adaptation. These multimodal biosensors may include photoplethysmography (PPG) sensors, microphones, and electrocardiogram (ECG) sensors for local data acquisition. The data processing unit has a built-in lightweight emotion analysis model, which is a pre-trained neural network model that can be used to achieve functions such as data fusion, emotion modeling, audio generation, and dynamic adaptation.
[0028] Step S10: Collect the user's biometric data through the wearable device at preset time intervals; Biometric data is a quantitative representation of a user's current emotional state, including physiological and vocal data. Physiological data is a quantitative representation of a user's physiological signals, reflecting real-time changes in the user's autonomic nervous system, motor state, and metabolic activity, which may include heart rate, heart rate variability, and the user's resting heart rate. Vocal data consists of the acoustic features and semantic content of the user's voice signals, including explicit commands (such as "turn down the volume") and implicit emotional cues (such as trembling tone and increased speech rate).
[0029] The preset duration interval refers to the time interval between data sampling and processing (e.g., 5 seconds), which is triggered by the system clock and represents the length of one cycle.
[0030] Optionally, physiological data can be collected using a fixed frequency (e.g., 10Hz) via PPG. PPG is an optical-based biosensor widely used in smart bracelets, medical devices, and other fields, and can be used for non-invasive monitoring of cardiovascular activity-related physiological parameters (such as heart rate and blood oxygen saturation).
[0031] Optionally, voice data can be acquired via a microphone. During the acquisition process, Voice Activity Detection (VAD) technology can be incorporated, initiating a power-intensive voice feature extraction process only when the microphone detects the user's voice, thereby obtaining the voice data. VAD is a technology used to automatically identify speech segments and non-speech segments (such as silence, background noise, and non-speech sounds) in an audio signal. It can accurately segment the effective speech portion from a continuous audio stream, for example, retaining only speech segments as the effective speech portion and removing non-speech segments, thereby accurately extracting the user's spoken content.
[0032] Understandably, by collecting biometric data through wearable devices at preset intervals, localized data collection is achieved. Compared to cloud-based data transmission scenarios, this allows data to be directly collected to the corresponding processing modules, significantly reducing data transmission latency.
[0033] Step S20: In the current cycle, generate a first emotional state vector based on the first biometric data collected in the current cycle; The current period is the baseline time period, which can be the initial time window when the wearable device is first started (i.e., t=0 to t=T, where T is a preset duration interval), or the Nth time window after the wearable device is started (i.e., t=N to t=N+T). The biometric data collected in the current period is the first biometric data. The first emotional state vector is a low-dimensional emotional representation generated by an emotional analysis model, used to represent the user's emotional state in the current period, and is represented in the form of a normalized numerical vector.
[0034] Understandably, because this step is performed on wearable devices, it is possible to complete the process from biometrics to emotion vectors directly on the user's terminal side, avoiding cloud interaction. This not only further reduces end-to-end latency but also ensures service continuity in environments without network coverage.
[0035] Step S30: Calculate the change in emotion between the first emotional state vector and the second emotional state vector, where the second emotional state vector is generated in the previous adjacent cycle of the current cycle; The previous adjacent cycle is the adjacent cycle before the current cycle. The emotional state vector generated by it is the second emotional state vector. The second emotional state vector is similar to the first emotional state vector. It is also a low-dimensional emotional representation generated by the emotional analysis model. It is used to represent the user's emotional state in the previous adjacent cycle and is represented in the form of a normalized numerical vector.
[0036] The change in emotional state is the difference between the second and first emotional state vectors, and can be calculated using Euclidean distance or Dynamic Time Warping (DTW) algorithms. The calculation expression is:
[0037] in, For the measure of emotional change, This is the first emotional state vector. This is the second emotional state vector.
[0038] Step S40: Adjust the historical audio data of the previous adjacent period based on the amount of emotional change and the received user feedback information to obtain the target audio data, and output the target audio data, wherein the target audio data is adapted to the first biometric data.
[0039] User feedback is a quantitative record of user actions (such as skipping songs) or implicit behaviors (such as playback completion rate and volume adjustment) on audio data. For example, if a user skips a song or cuts it halfway through playback, it indicates that the user is dissatisfied with the audio data. Conversely, if the user plays the entire song or loops it, it indicates that the user is satisfied with the audio data.
[0040] Audio data is digital music content generated or scheduled based on music parameters, which can be pre-stored files (such as MP3s) or real-time synthesized waveforms (such as FM synthesizer output). Historical audio data is audio generated from the previous adjacent period based on the second emotional state vector, and is the audio currently playing on the wearable device's speaker, adapted to the user's biometric data (second biometric data) from the previous adjacent period. The new audio data obtained after adjusting it based on emotional changes and user feedback is the target audio data, which can be adapted to the user's first biometric data in the current period.
[0041] It should be noted that because users' biometric data changes in real time, there may be significant fluctuations between the current period and the previous adjacent period. For example, a user might have been quite excited in the previous period, but after receiving bad news, their mood might become calm or even depressed in the next period, resulting in significant emotional fluctuations. Since historical audio data is adapted to the second biometric data, if the user's biometric data changes significantly in the current period, the historical audio data from the previous adjacent period will no longer be suitable and needs to be adjusted to match the user's first biometric data in the current period. The adjustment is based on the magnitude of emotional change.
[0042] Additionally, it should be noted that even if a user's biometric data has not changed significantly, there may still be cases where the user dislikes a certain audio data (e.g., doesn't like listening to the currently playing song). In this case, the audio data can be adjusted directly based on the received user feedback, and the adjusted audio data must also be adapted to the user's biometric data in the current period.
[0043] Understandably, by calculating the emotional changes between adjacent periods and integrating user feedback, it is possible to quantify the changing trends of the user's state and adjust the audio adaptation strategy accordingly in a timely manner. This solves the problem of untimely strategy updates caused by the static nature of existing technologies, ensuring that audio content can dynamically adapt to continuous changes in the user's state.
[0044] In one feasible implementation, the first biometric data includes physiological data and voice data. Step S20, which generates a first emotional state vector based on the first biometric data collected in the current period, includes: Step S201: Calculate the corresponding physiological emotion component based on physiological data, and calculate the corresponding speech emotion component based on speech data; Physiological emotion components are emotion representation vectors calculated from physiological data, representing the contribution of the user's physiological state to their emotions. Speech emotion components are emotion representation vectors extracted from speech features, representing the contribution of the user's speech information to their emotions.
[0045] It should be noted that the calculation of physiological emotional components and phonological emotional components can be performed in parallel.
[0046] In one feasible implementation, step S201 includes: Step A10: Calculate the physiological emotional component corresponding to the physiological data based on physiological data, preset normalization parameters, and preset weighting coefficients. The physiological data includes heart rate, resting heart rate, and heart rate variability. The preset normalization parameters are standardized coefficients used to eliminate individual differences, including maximum heart rate, minimum heart rate, and maximum heart rate variability. The preset weighting coefficients are weighting parameters used in the calculation of physiological and emotional components to adjust the contribution of different physiological indicators; their initial values are preset.
[0047] Heart rate refers to the number of times the heart beats per unit of time, such as 72 beats per minute. Resting heart rate refers to the user's baseline heart rate value in a resting state, calculated through long-term monitoring. Heart rate variability is a quantitative indicator of the difference between consecutive heartbeat cycles, used to reflect the user's autonomic nervous system's regulatory ability.
[0048] It's important to note that physiological data such as heart rate and heart rate variability (HRV) are related to user emotions. For example, when a user experiences high-arousal emotions such as tension, anxiety, excitement, or fear, their sympathetic nervous system is activated, releasing substances like adrenaline, leading to increased cardiac contractility and a faster heart rate. Conversely, in low-arousal emotional states such as relaxation and calmness, the parasympathetic nervous system (primarily responsible for "rest and digestion") dominates, and the heart rate slows down. High HRV usually indicates strong parasympathetic activity, flexible and adaptive cardiac regulation, and a relaxed, restorative state, associated with lower emotional stress. Low HRV usually indicates dominant sympathetic activity or weakened parasympathetic activity, stiffened cardiac regulation, and a state of stress, tension, or fatigue, highly correlated with emotional states such as anxiety, depression, and high stress.
[0049] Optionally, the formula for calculating the physiological emotional component can be:
[0050] in, As for physiological and emotional components, and For preset weighting coefficients, Heart rate, Resting heart rate For heart rate variability, , as well as To preset normalization parameters, For maximum heart rate, Minimum heart rate and This represents the maximum heart rate variability value. This is the heart rate normalization term, used to quantify the degree to which the heart rate deviates from the resting heart rate. When the heart rate is significantly higher than the resting heart rate, the output of this term increases, indicating that the user may be in a state of excitement or tension. The HRV normalization term reflects the user's autonomic nervous system balance through heart rate variability (HRV). A low HRV leads to an increase in this term value, indicating that the user may be in a state of high stress and needs to enhance emotional arousal.
[0051] Step A20: Input the speech data into a preset neural network to obtain the speech emotion component corresponding to the speech data. The speech data includes speech rate, fundamental frequency and sound pressure level features.
[0052] The preset neural network is a pre-trained lightweight deep learning model, such as 1D-CNN or miniature Transformer. The input dimensions of the preset neural network are [speech rate, fundamental frequency, sound pressure level], and the output is speech emotion components, which can map speech data into speech emotion components that represent the user's current emotional state.
[0053] Speech rate is the number of syllables pronounced per unit time, such as 4.2 syllables / second; fundamental frequency is the basic frequency of vocal cord vibration, which determines the pitch of speech; sound pressure level is used to quantify the intensity of speech signals, in decibels.
[0054] It should be noted that speech data such as speech rate, fundamental frequency, and sound pressure level (SPL) characteristics are related to user emotions. For example, an increased speech rate is positively correlated with a high arousal emotional state, while a decreased speech rate is often associated with a low arousal emotional state; an increased fundamental frequency and a wider fluctuation range are associated with a high arousal emotional state, while a decreased fundamental frequency and smoother fluctuations are associated with a low arousal state; an increased SPL characteristic (i.e., increased volume) is associated with a high arousal emotional state, while a decreased SPL characteristic is associated with a low arousal emotional state.
[0055] Optionally, the calculation process for speech emotion components can be as follows:
[0056] in, For the emotional component of speech, For the preset neural network, For speaking speed, For the fundamental frequency, This is a characteristic of sound pressure level.
[0057] Understandably, by fusing multidimensional physiological data such as heart rate, resting heart rate, and heart rate variability into a single physiological-emotional component, the problem of misjudgment of emotions caused by relying on a single heart rate indicator in traditional solutions is solved, significantly improving the accuracy of physiological state assessment. By using a pre-set neural network to perform nonlinear transformations on speech features such as speech rate, fundamental frequency, and sound pressure level to generate speech emotion components, the problem of poor adaptability and limited accuracy of traditional rule-based speech emotion analysis is solved, enabling the extraction of deep emotional patterns from complex acoustic features. By computing physiological and speech emotion components in parallel, the perceptual blind spots of a single modality are compensated for, improving the accuracy of emotion assessment.
[0058] Step S202: Based on the first preset fusion weight, the physiological emotional component and the voice emotional component are weighted and summed to generate the first emotional state vector.
[0059] The first preset fusion weight is the fusion weight used to generate the first emotional state vector of the current period. It includes the weighting coefficients for the fusion of physiological emotional components and voice emotional components, and is used to adjust the contribution of the two.
[0060] Optionally, the formula for calculating the first emotional state vector is:
[0061] in, This is the first emotional state vector. and They are respectively physiological and emotional components. and the emotional weight of voice The preset fusion weights.
[0062] In this embodiment, by collecting biometric features locally at preset time intervals and generating emotion vectors, cloud transmission delays are eliminated. Based on the amount of emotion change and user feedback, the audio data is dynamically adjusted, which can continuously track the user's state transition. This overcomes the rigidity of static model adaptation and solves the problems of delayed emotional feedback and untimely strategy updates in traditional audio adaptation solutions, ensuring that the audio content is always accurately synchronized with the user's dynamically changing physiological state.
[0063] Based on the first embodiment described above, a second embodiment of the audio adaptation method of this application is proposed. In this embodiment, content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. (Refer to...) Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the audio adaptation method of this application. In this embodiment, step S30, which involves adjusting the historical audio data of the previous adjacent period based on the amount of emotional change and the received user feedback information to obtain the target audio data, includes: Step S301: Update the first preset fusion weight based on the emotional change amount, update the first preset music parameter mapping table based on the received user feedback information, and obtain the second preset fusion weight and the second preset music parameter mapping table. The first preset fusion weight is used to generate the first emotional state vector, the second preset fusion weight is used to generate the third emotional state vector of the next adjacent cycle of the current cycle, and the first preset music parameter mapping table is used to obtain the historical audio data of the previous adjacent cycle. The preset music parameter mapping table is a mapping lookup table stored locally on the wearable device, which defines the mapping relationship between emotional state vectors and music parameters. The first preset music parameter mapping table is applied to the previous adjacent cycle to determine the music parameters corresponding to the second emotional state vector (i.e., the second music parameters), thereby obtaining historical audio data. The second preset music parameter mapping table is applied to the current cycle to determine the first music parameters corresponding to the first emotional state vector, thereby obtaining the target audio data.
[0064] For example, a preset music parameter mapping table could be:
[0065] Specifically, when the emotional state vector E is less than -0.5, it indicates that the user is in a "deeply relaxed" state, and the target music parameters are a target rhythm of 60-70 BPM, a target harmony consisting of major keys and long chords, and low-pass filtering to soften the timbre. When E is between -0.5 and 0, it indicates that the user is in a "relaxed" state, and the corresponding music parameters are a rhythm of 70-90 BPM, a major key harmony, and moderate timbre brightness. When E is between 0 and 0.5, it indicates that the user is in a "neutral" state, and the music parameters are a rhythm of 90-110 BPM, a mixed major and minor key harmony, and moderate timbre brightness. When E is greater than or equal to 0.5, it indicates that the user is in an "anxious / excited" state, and the music parameters are a target rhythm of 110-130 BPM, a target harmony consisting of minor keys and short chords, and high-pass filtering to brighten the timbre.
[0066] In one feasible embodiment, step S301, which involves updating the first preset fusion weight based on the amount of emotional change and updating the first preset music parameter mapping table based on the received user feedback information, includes: Step B10: If the change trend of the emotional change is inconsistent with that of the first biometric data, reduce the first preset fusion weight. The trend refers to the direction of change of the emotional component value corresponding to the biometric data within a preset time window. For example, if the emotional component of speech in the current period is 0.8 and that in the previous period was 0.5, then the emotional component of speech is determined to be showing an upward trend.
[0067] Optionally, the preset adjustment formula for the fusion weights is:
[0068] in, Preset fusion weights corresponding to biometric data, =1or2, The preset fusion weights are the physiological and emotional components. Preset fusion weights corresponding to the speech emotion components. The adjusted preset fusion weights, The learning rate controls the step size for weight updates. The larger the value, the greater the magnitude of a single weight adjustment. The smaller the size, the more precise and smooth the adjustments. For the measure of emotional change, This is a symbolic function used to determine the direction of overall emotional change. A value greater than 0 indicates a shift in overall emotion towards greater excitement / anxiety. If the value is less than 0, the overall mood shifts towards a more relaxed direction. = 0 indicates no change in overall emotion. To indicate the weight of emotions, As for physiological and emotional components, For the emotional component of speech, It is the average value of the emotional components (physiological emotional components or vocal emotional components) within a preset time window.
[0069] For example, if the user's voice emotion component is detected to decrease from 0.7 to 0.3 ( =-0.4), indicating a relaxation trend, the physiological emotional component increased from 0.4 to 0.8 ( =+0.4), representing the excitement trend, and after fusion based on the current weights, the overall change in emotion is calculated. =+0.15>0, indicating the user's overall emotion is trending towards excitement. At this point, a consistency diagnosis is performed: the physiological modality change trend (+) and (+) consistent, speech modality change trend (-) and (+) is inconsistent, therefore the change in emotional quantity is inconsistent with the change trend of voice data. The weight update mechanism is activated to reduce the preset fusion weight corresponding to the voice emotional component.
[0070] Step B20: If the user feedback is unsatisfactory, reduce the correlation strength between the second emotional state vector and the second music parameter in the first preset music parameter mapping table, wherein the second emotional state vector and the second music parameter correspond to the previous adjacent cycle.
[0071] Association strength refers to the confidence index used in a preset music parameter mapping table to quantify the binding relationship between emotional state vectors and music parameters.
[0072] For example, the emotion state vector When the value is -0.3 (in the relaxed state range), music parameters M_1 {tempo: 75 BPM, harmony: major, timbre: moderate} are selected according to the preset music parameter mapping table, and the corresponding audio data is played. If the user gives negative feedback of "dislike," indicating dissatisfaction, the preset music parameter mapping table optimization mechanism is activated to reduce the emotional vector. =0.3 reduces the correlation strength between the parameter set M_1 and the parameter set M_1 (e.g., from the initial value of 0.9 to 0.7), thereby reducing the probability of selecting the unpopular parameter again in subsequent cycles by decreasing the original correlation strength.
[0073] Understandably, by monitoring the consistency between the change in emotional state and the trend of biometric data, the weight of modal data that contradicts the overall trend is identified and reduced, thus solving the problem of emotional misjudgment caused by data distortion or noise interference from a single sensor in multimodal systems, ensuring the accuracy and robustness of emotional state calculation. By directly incorporating explicit user feedback into the music strategy optimization loop, when negative feedback is received, the mapping relationship between the current emotional vector and music parameters is weakened, breaking through the limitations of the fixed static mapping model. It can adjust the output strategy in real time according to the user's subjective preferences, effectively solving the problem of untimely updates of adaptation strategies caused by fixed mapping relationships.
[0074] Step S302: Determine the first music parameter corresponding to the current period based on the second preset music parameter mapping table, and adjust the historical audio data based on the first music parameter to obtain the target audio data.
[0075] Musical parameters are a set of core variables that control audio generation or scheduling, and may include target tempo (BPM), target harmony, and target timbre brightness, among other things. The first musical parameter is applied to the current cycle, and the second musical parameter is applied to the previous adjacent cycle.
[0076] It should be noted that after obtaining the second preset music parameter mapping table, a first music parameter can be generated based on it. This music parameter acts on the current cycle and can be used to adjust the historical audio data in the previous adjacent cycle. The newly generated target audio data can be adapted to the user's first biometric data in the current cycle. In the current cycle, the first music parameter can be associated with the user's emotional state in the current cycle. Furthermore, the audio data obtained based on the first music parameter can also take into account the user's emotional state in the current cycle, that is, the target audio data can be adapted to the user's first biometric data in the current cycle.
[0077] In one feasible implementation, step S302, which involves adjusting historical audio data based on the first music parameter to obtain the target audio data, includes: Step C10: Adjust the preset audio synthesizer based on the first music parameters to generate the target audio data corresponding to the first music parameters; or Preset audio synthesizers are locally integrated lightweight audio synthesizers (such as FM synthesizers, wavetable synthesizers, or rule-based generative models) that receive music parameters and generate audio data corresponding to the music parameters through real-time signal processing.
[0078] Optionally, a preset audio synthesizer is driven according to the first musical parameters, adjusting the configuration parameters of its internal signal processing unit in real time. For example, the frequency modulation index of the oscillator is adjusted according to the target rhythm parameters, the note sequence and chords are set according to the target harmony parameters, and the cutoff frequency and formant characteristics of the filter are configured according to the target timbre brightness parameters. This generates a new audio waveform that conforms to the first emotional state vector, i.e., the target audio data.
[0079] Step C11: Retrieve the audio segment with the highest similarity to the first music parameter from the preset music material library, and use the audio segment as the target audio data corresponding to the music parameter.
[0080] The preset music library is a locally stored database of pre-generated audio clips. Each clip is labeled with music parameter tags (such as BPM, emotion vector, and timbre type), and supports quick retrieval based on the similarity of parameter tags.
[0081] Audio clips are the smallest callable units in a preset music library. They are usually short audio loops (such as 2-8 seconds of drum beats, melodies, or chord progressions), have complete musical characteristics, and are easy to splice together.
[0082] Optionally, the audio segment with the highest matching degree is retrieved from a preset music library based on the first music parameter. For example, by calculating the similarity between the first music parameter and the music parameter tags marked on each audio segment in the preset music library, the audio segment with the highest similarity is selected as the target audio data corresponding to the music parameter.
[0083] Understandably, by driving a preset audio synthesizer and adjusting its oscillator and filter parameters in real time, it can synthesize entirely new audio waveforms based on music parameters, ensuring synchronization between music feedback and changes in emotional state. This enables fine and continuous adjustment of music parameters, significantly improving the accuracy of audio adaptation. By using music parameters obtained based on emotion vector mapping, the most matching audio segments are quickly retrieved from a pre-loaded tagged material library, ensuring the continuous relevance of audio content to the user's state. Furthermore, since all audio segments are pre-loaded locally, there is no need to rely on cloud recommendations, effectively solving the experience gap caused by untimely updates to adaptation strategies.
[0084] In one feasible implementation, the step of outputting the target audio data in step S40 includes: Step S401: Upon receiving a music switching signal, crossfade processing is performed on the target audio data and historical audio data to switch the historical audio data to the target audio data. The crossfade processing is used to smoothly transition the switching process of the target audio data.
[0085] The music switching signal is the instruction that triggers the audio switching. It is triggered under preset scenarios, which may include: timed switching scenarios, emotional change scenarios, and user feedback scenarios.
[0086] Specifically, the timed switching scenario refers to the scenario where audio data is generated at the end of the current cycle. For example, if audio data is obtained at the end of the current cycle, a music switching signal will be triggered to switch the currently playing audio. The emotional change scenario refers to the scenario where the change in emotional state is inconsistent with the trend of the user's biometric data. For example, if the change in emotional state indicates that the user's emotional state is gradually relaxing, but the user's heart rate continues to rise, it means that the currently playing audio does not match the user's emotions. In this case, a music switching signal will be triggered to switch the currently playing audio. The user feedback scenario is the scenario where a user feedback instruction is received (such as clicking the "dislike" button or issuing a voice command to "switch music"). In this case, it is necessary to respond to the user's instruction and trigger a music switching signal to switch the currently playing audio.
[0087] Crossfading is a time-function-based dual-channel gain dynamic control algorithm that can continuously adjust the amplitude weights of the currently playing audio and the audio data to achieve a smooth transition from the currently playing audio to the audio data and eliminate auditory gaps during switching.
[0088] Optionally, the calculation formula for cross-fading can be expressed as:
[0089] in, These are the sample values of the audio signal finally output to the speaker after crossfading. Indicates the current time point, and The gain coefficient for cross-fading processing. For the currently playing audio, This is the audio data, and it's about to switch to the new audio stream that's currently playing. The crossfade processing calculation formula achieves a smooth energy transfer from one audio stream to another by linearly and complementaryly adjusting the gain of the two audio signals.
[0090] Optionally, crossfading can be performed in a preset circular buffer. The preset circular buffer is a ring-shaped storage area with its ends connected. When any pointer moves to the end of the circular buffer, it will automatically circle back to the beginning. This structure determines that when the circular buffer is full, if writing continues, new data will overwrite the oldest data (if that data has been read) or wait (if the old data has not been read).
[0091] It's important to note that when crossfading an audio segment, two different data segments need to be read and processed simultaneously: the end of the previous audio segment (historical sample) and the beginning of the next audio segment (future sample). Due to the circular nature of the pre-defined circular buffer, it logically becomes an infinitely reusable data loop. When the write pointer reaches the end of the buffer, it automatically loops back to the beginning, overwriting old data. This structure ensures that the currently playing audio and the audio data can coexist continuously in the same memory block. Therefore, during crossfading, independent read and write pointers can stably access any historical or future sample from both streams, ensuring continuous, latency-free mixing calculations of a large number of audio samples within a short fade window. This is something other data structures (such as stacks and queues) cannot achieve. For example, while queues also support FIFO (First-In, First-Out), their linear structure causes the memory of dequeued old data to be released, making it impossible to support backtracking historical data reading or parallel access to multiple streams. The LIFO (Last-In, First-Out) characteristic of stacks completely violates the sequential playback nature of audio streams, failing to guarantee that data is consumed in chronological order.
[0092] Understandably, when the music switching signal is triggered, the crossfading processing mechanism solves the problem of auditory stuttering and experience interruption caused by the instantaneous music switching in traditional audio adaptation solutions. It eliminates the silent gaps and signal abrupt changes during the switching process, ensures the continuity of music immersion, and improves the user experience.
[0093] For example, refer to Figure 3 , Figure 3This is an overall flowchart of the audio adaptation method provided in the second embodiment of this application. After the process begins, biometric data is first collected, and voice and physiological data are acquired in parallel. Then, emotional state vectors are calculated, and these vectors are converted into specific music parameters through music parameter mapping. The system then generates and plays music based on these music parameters. The process continuously monitors whether a feedback cycle has been reached. If not, it continues playing; if so, it initiates an optimization loop, sequentially re-collecting biometric data, calculating emotional changes, updating preset fusion weights and the preset music parameter mapping table, and finally generating a new set of music parameters to start the next adaptation cycle, thereby achieving continuous dynamic optimization of the audio content.
[0094] In this embodiment, by detecting the consistency between the trend of biometric data and the overall emotional change, the modality fusion weights are dynamically adjusted to suppress unreliable data sources and improve the accuracy and real-time performance of emotional state recognition. By utilizing negative user feedback to directly weaken the association of unpopular music parameters, online adaptive correction of the mapping relationship is achieved. The combination of these two approaches forms a closed-loop optimization, ensuring that the system can output accurately adapted audio content in the current cycle based on the latest state and user preferences. This overcomes the slow response of traditional static models and effectively solves the problems of delayed emotional feedback and untimely strategy updates in audio adaptation.
[0095] This application also provides an audio adapter device, please refer to... Figure 4 The audio adapter includes: Data acquisition module 10 is used to collect the user's biometric data through a wearable device at preset time intervals; Vector calculation module 20 is used to generate a first emotional state vector based on the first biometric data collected in the current cycle. The difference calculation module 30 is used to calculate the amount of emotional change between the first emotional state vector and the second emotional state vector, wherein the second emotional state vector is generated in the previous adjacent cycle of the current cycle; The audio adjustment module 40 is used to adjust the historical audio data of the previous adjacent period based on the amount of emotional change and the received user feedback information to obtain the target audio data and output the target audio data, wherein the target audio data is adapted to the first biometric data.
[0096] The audio adaptation device provided in this application adopts the audio adaptation method in the above embodiments. Compared with the prior art, the beneficial effects of the audio adaptation device provided in this application are the same as the beneficial effects of the audio adaptation method provided in the above embodiments. Moreover, the other technical features in the audio adaptation device are the same as the features disclosed in the method of the above embodiments, and will not be described in detail here.
[0097] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the audio adaptation method in the first embodiment described above.
[0098] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0099] like Figure 5 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0100] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0101] Compared with the prior art, the beneficial effects of the electronic device provided in this application embodiment are the same as the beneficial effects of the audio adaptation method provided in the above embodiment, and other technical features in the electronic device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0102] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0103] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the audio adaptation method in the above embodiments.
[0104] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0105] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0106] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to perform the functions defined in the methods of the embodiments disclosed in this application.
[0107] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0109] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0110] The readable storage medium provided in this application embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for performing the above-described audio adaptation method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the audio adaptation method provided in the above-described embodiments, and will not be repeated here.
[0111] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the audio adaptation method described above.
[0112] Compared with the prior art, the beneficial effects of the computer program product provided in this application embodiment are the same as the beneficial effects of the audio adaptation method provided in the above embodiments, and will not be repeated here.
[0113] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. An audio adaptation method, characterized in that, The audio adaptation method, applied to wearable devices, includes: The wearable device collects the user's biometric data at preset time intervals. In the current cycle, a first emotional state vector is generated based on the first biometric data collected in the current cycle; Calculate the change in emotion between the first emotional state vector and the second emotional state vector, wherein the second emotional state vector is generated in the previous adjacent cycle of the current cycle; Based on the emotional change and the received user feedback, the historical audio data of the previous adjacent period is adjusted to obtain target audio data, and the target audio data is output, wherein the target audio data is adapted to the first biometric data.
2. The audio adaptation method as described in claim 1, characterized in that, The first biometric data includes physiological data and voice data. The step of generating a first emotional state vector based on the first biometric data collected in the current period includes: The physiological emotional component is calculated based on the physiological data, and the corresponding speech emotional component is calculated based on the speech data. The physiological emotional component and the voice emotional component are weighted and summed based on the first preset fusion weight to generate a first emotional state vector.
3. The audio adaptation method as described in claim 2, characterized in that, The steps of calculating the corresponding physiological emotion component based on the physiological data and calculating the corresponding speech emotion component based on the speech data include: The physiological emotional component corresponding to the physiological data is calculated based on the physiological data, preset normalization parameters, and preset weighting coefficients. The physiological data includes heart rate, resting heart rate, and heart rate variability. The speech data is input into a preset neural network to obtain the speech emotion component corresponding to the speech data, wherein the speech data includes speech rate, fundamental frequency and sound pressure level features.
4. The audio adaptation method as described in claim 1, characterized in that, The step of adjusting the historical audio data of the previous adjacent period based on the emotional change and the received user feedback information to obtain the target audio data includes: The first preset fusion weight is updated based on the emotional change amount, and the first preset music parameter mapping table is updated based on the received user feedback information to obtain the second preset fusion weight and the second preset music parameter mapping table. The first preset fusion weight is used to generate the first emotional state vector, the second preset fusion weight is used to generate the third emotional state vector of the next adjacent cycle of the current cycle, and the first preset music parameter mapping table is used to obtain the historical audio data of the previous adjacent cycle. Based on the second preset music parameter mapping table, the first music parameter corresponding to the current period is determined, and the historical audio data is adjusted based on the first music parameter to obtain the target audio data.
5. The audio adaptation method as described in claim 4, characterized in that, The steps of updating the first preset fusion weight based on the emotional change amount and updating the first preset music parameter mapping table based on the received user feedback information include: If the change in the emotional state is inconsistent with the trend of the change in the first biometric data, the first preset fusion weight is reduced. If user feedback is received and the feedback is unsatisfactory, the correlation strength between the second emotional state vector and the second music parameter in the first preset music parameter mapping table is reduced, wherein the second emotional state vector and the second music parameter correspond to the previous adjacent cycle.
6. The audio adaptation method as described in claim 4, characterized in that, The step of adjusting the historical audio data based on the first music parameter to obtain the target audio data includes: Based on the first music parameters, adjust the preset audio synthesizer to generate the target audio data corresponding to the first music parameters; or The audio segment with the highest similarity to the first music parameter is retrieved from the preset music material library, and the audio segment is used as the target audio data corresponding to the music parameter.
7. The audio adaptation method as described in claim 1, characterized in that, The step of outputting the target audio data includes: Upon receiving a music switching signal, the target audio data and the historical audio data are cross-faded, and the historical audio data is switched to the target audio data. The cross-fading process is used to smoothly transition the switching process of the target audio data.
8. A wearable device, characterized in that, The wearable device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio adaptation method as described in any one of claims 1 to 7.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the audio adaptation method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the audio adaptation method as described in any one of claims 1 to 7.