Real-time dynamic control system for music parameters based on multimodal physiological signals
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]本申请的目的在于提供基于多模态生理信号的实时音乐参数动态调控系统,以解决因盲目同步融合导致的控制过冲与系统震荡问题,以提升系统的长期稳定性与个性化适配能力
本申请通过预标定各生理通道固有传导延时,构建由传导时间差严格约束深度的特征缓存队列,在计算域内彻底消除了多模态生理响应的物理传导相位差,使皮电与心电特征在因果意义上对齐至同一历史音频刺激周期,从根本上解决了现有技术中因盲目同步融合导致的控制过冲与系统震荡问题;
Smart Images

Figure CN122551810A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the interdisciplinary field of digital signal processing and biofeedback regulation, specifically a real-time dynamic control system for music parameters based on multimodal physiological signals. Background Technology
[0002] Music-assisted relaxation systems that adjust music playback parameters in real time based on the user's physiological state are gaining increasing attention. A typical implementation path for existing technologies involves simultaneously acquiring multimodal physiological signals within the current time window, directly weighting and averaging the signals to calculate the joint state difference, and then mapping this difference to audio control parameters. The underlying assumption of this approach is that the responses of various physiological systems in the human body to the current audio stimulus are instantaneous and synchronous. However, the above assumptions deviate significantly from physical reality under actual working conditions. The sympathetic nerve-dominated skin conductance response to music stimulation typically requires only 1 to 3 seconds of delay, while the vagus nerve-dominated heart rate variability response requires 10 to 20 seconds. This means that the skin conductance signal and heart rate variability signal collected at the same physical moment are not induced by the same audio stimulation cycle, and forcibly averaging and fusing the two will lead to serious systematic misjudgments. Specifically, during the ten-second misalignment period when skin conductance has reached the target threshold but heart rate variability has not yet responded, the system misjudges insufficient music adjustment intensity due to the lag in heart rate variability, thus continuously outputting excessively high audio parameter increments. When heart rate variability finally responds, the excessively high music parameters in turn trigger sympathetic nerve stress, causing the system to fall into control oscillations, severely disrupting the user's relaxation experience. Existing technology urgently needs improvement to address these issues. Summary of the Invention
[0003] The purpose of this application is to provide a real-time dynamic music parameter control system based on multimodal physiological signals to solve the problems of control overshoot and system oscillation caused by blind synchronization fusion, so as to improve the long-term stability and personalized adaptability of the system.
[0004] The objective of this application can be achieved through the following technical solution: Firstly, a real-time dynamic control system for music parameters based on multimodal physiological signals, comprising the following modules: The time difference acquisition module is used to acquire the skin conduction delay and electrocardiogram conduction delay of the target object in response to the preset test audio, and to obtain the conduction time difference based on the skin conduction delay and the electrocardiogram conduction delay; The queue construction module is used to construct a feature cache queue corresponding to the target depth based on the transmission time difference and the preset system sampling rate; The first extraction module is used to synchronously acquire the skin electrophysiological signal and electrocardiographic signal of the target object at the current moment according to the preset system sampling rate during the playback of the target audio, and extract the skin electrophysiological feature value and electrocardiographic feature value at the current moment respectively. The second extraction module is used to push the current electrodermal feature value into the feature cache queue and extract the electrodermal feature value of the target historical time, wherein the time interval between the target historical time and the current time is equal to the conduction time difference. The error value evaluation module is used to perform causal alignment fusion processing on the skin electrophysiological characteristic values of the target historical time and the electrocardiological characteristic values of the current time to obtain the joint state error value of the current time. The parameter control module is used to generate audio control parameters based on the joint state error value, update the playback parameters of the corresponding target audio according to the parameters, obtain feedback optimization parameters that characterize the music control effect, and update the preset system sampling rate according to the parameters.
[0005] Secondly, a real-time dynamic control method for music parameters based on multimodal physiological signals includes the following steps: Obtain the skin conduction delay and electrocardiogram conduction delay of the target object in response to a preset test audio, and obtain the conduction time difference based on the skin conduction delay and the electrocardiogram conduction delay; Construct a feature cache queue corresponding to the target depth based on the transmission time difference and the preset system sampling rate; During the playback of the target audio, the skin electrophysiological signal and electrocardiographic signal of the target object at the current moment are acquired synchronously according to the preset system sampling rate, and the skin electrophysiological feature value and electrocardiographic feature value at the current moment are extracted respectively. The current electrodermal feature value is pushed into the feature cache queue, and the electrodermal feature value of the target historical time is extracted. The time interval between the target historical time and the current time is equal to the conduction time difference. The skin electrophysiological characteristic values at the target historical moment and the electrocardiological characteristic values at the current moment are subjected to causal alignment and fusion processing to obtain the joint state error value at the current moment; Audio control parameters are generated based on the joint state error value, and the playback parameters of the corresponding target audio are updated accordingly to obtain feedback optimization parameters that characterize the music control effect, and the preset system sampling rate is updated accordingly.
[0006] Thirdly, a computer storage medium stores computer-executable instructions, which, when executed, realize the real-time dynamic music parameter control system based on multimodal physiological signals described in the field of skin conductance.
[0007] Compared with the prior art, the beneficial effects of this application are: This application constructs a feature buffer queue with strictly constrained depth by pre-calibrating the inherent conduction delay of each physiological channel, and completely eliminates the physical conduction phase difference of multimodal physiological responses in the computational domain, so that the skin conduction and electrocardiogram features are causally aligned to the same historical audio stimulation cycle, fundamentally solving the control overshoot and system oscillation problems caused by blind synchronization fusion in the prior art; By introducing a proportional-integral controller and an adaptive sampling rate adjustment mechanism, while ensuring control smoothness, the preset system sampling rate and feature buffer queue depth can be dynamically adjusted according to the feedback error variance. This enables the system to have good adaptability under different users and different wearing conditions, significantly improving the system's long-term stability and personalized adaptability. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the modules of the real-time dynamic music parameter control system based on multimodal physiological signals according to this application; Figure 2 This is a schematic diagram illustrating the steps of the real-time dynamic control method for music parameters based on multimodal physiological signals according to this application. Detailed Implementation
[0009] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but only to illustrate selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item has been defined in one figure, it does not need to be further defined and explained in subsequent figures. The terms "first", "second", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0010] In existing biofeedback-based real-time audio parameter control systems, the system synchronously acquires skin conductance signals and electrocardiogram (ECG) signals at the same physical moment, directly weighting and fusing them to generate audio control parameters. This approach implicitly assumes that the responses of various physiological channels to audio stimuli are instantaneous and synchronous. However, the conduction velocity of skin conductance, dominated by the sympathetic nervous system, is much faster than that of ECG, dominated by the vagus nerve. This results in the skin conductance signal reflecting the physiological response to the current audio stimulus at the same sampling moment, while the ECG signal reflects the physiological response to a historical stimulus tens of seconds ago. The two signals are severely misaligned in terms of physical causality, and the joint state error obtained by forcibly merging them synchronously is not the true joint response of the human body to the same music parameters, inevitably leading to control overshoot and system oscillation.
[0011] For example, in a typical music-assisted relaxation scenario, when the audio system executes a command to lower the BPM, the user's skin conductance activity decreases significantly within about 2 seconds, indicating reduced tension; however, heart rate variability does not reverse until about 15 seconds later. During this approximately 13-second misalignment period, the existing system, observing that skin conductance has reached the target level but heart rate variability has not improved, mistakenly believes that the adjustment is insufficient and continues to increase the audio parameters. When heart rate variability finally responds, the over-adjusted music parameters trigger a new round of skin conductance stress response, causing the system to fall into continuous oscillation and severely disrupting the user's relaxation experience.
[0012] To address the aforementioned issues, the core idea of this application is to extract the conduction delays of various physiological channels, which are considered noise in existing technologies, as core configuration parameters of the system architecture. Using the difference between the conduction delays of two channels as the sole physical basis, a first-in-first-out (FIFO) feature cache queue with strict physical depth is constructed in memory. During real-time control, the currently acquired electrodermal skin feature values are forcibly pushed into the queue and suspended. Historical electrodermal skin feature values popped from the tail of the queue, whose acquisition time coincides with the current electrocardiogram (ECG) feature value, point to the same historical audio stimulation cycle. By performing fusion processing on these two causally aligned feature values, the resulting joint state error value truly reflects the body's comprehensive physiological assessment of the audio modulation strategy at a given moment, thereby completely eliminating the root cause of control overshoot.
[0013] Therefore, such as Figure 1 As shown, this application provides a real-time dynamic adjustment system for music parameters based on multimodal physiological signals, including the following modules: The time difference acquisition module is used to acquire the skin conduction delay and electrocardiogram conduction delay of the target object in response to the preset test audio, and to obtain the conduction time difference based on the skin conduction delay and the electrocardiogram conduction delay; The queue construction module is used to construct a feature cache queue corresponding to the target depth based on the transmission time difference and the preset system sampling rate; The first extraction module is used to synchronously acquire the skin electrophysiological signal and electrocardiographic signal of the target object at the current moment according to the preset system sampling rate during the playback of the target audio, and extract the skin electrophysiological feature value and electrocardiographic feature value at the current moment respectively. The second extraction module is used to push the current electrodermal feature value into the feature cache queue and extract the electrodermal feature value of the target historical time, wherein the time interval between the target historical time and the current time is equal to the conduction time difference. The error value evaluation module is used to perform causal alignment fusion processing on the skin electrophysiological characteristic values of the target historical time and the electrocardiological characteristic values of the current time to obtain the joint state error value of the current time. The parameter control module is used to generate audio control parameters based on the joint state error value, update the playback parameters of the corresponding target audio according to the parameters, obtain feedback optimization parameters that characterize the music control effect, and update the preset system sampling rate according to the parameters.
[0014] The core innovation of this application lies in: Using the pre-calibrated multi-channel physiological response conduction time difference as the sole constraint, a feature cache queue with a depth strictly determined by this time difference is inserted between the fast response physiological feature extraction link and the multimodal fusion node. This actively delays the fusion timing of the skin conductance features, aligning them with the current electrocardiogram features in the same historical audio stimulation cycle in terms of physical causality, thereby completely eliminating overshoot regulation and system oscillation caused by misalignment waiting.
[0015] The specific implementation of the scheme in this application is as follows: During the system initialization phase, the time difference acquisition module sends a control command to the audio processing engine to trigger a preset test audio, records the excitation timestamp, and then monitors the instantaneous change rate of the skin electrophysiological signal and the high-frequency power fitting slope of the electrocardiographic signal. Based on preset criteria, it records the skin electrophysiological abrupt change timestamp and the electrocardiographic reversal timestamp, and then calculates the skin electrophysiological conduction delay and the electrocardiographic conduction delay, and uses the difference between the two as the conduction time difference.
[0016] The queue construction module calculates and allocates a first-in-first-out circular cache queue containing data nodes of the target depth based on the propagation time difference and the preset system sampling rate, thus completing the instantiation of the feature cache queue.
[0017] During the playback of the target audio, the first extraction module synchronously acquires the skin electrophysiological signal and the electrocardiographic signal at a preset system sampling rate. The skin electrophysiological feature values are obtained by high-pass filtering and normalization, and the electrocardiographic feature values are obtained by pulse wave period extraction and high-low frequency power ratio normalization.
[0018] The second extraction module pushes the current electrodermal feature value into the feature cache queue, and after the queue reaches the target depth, pops the target historical electrodermal feature value from the tail of the queue. This historical feature value and the current electrocardiographic feature value both point to the same historical audio stimulation cycle.
[0019] The error value assessment module calculates the difference between the skin electrophysiological characteristic value at the target historical time and the electrocardiological characteristic value at the current time and their respective target benchmark values, and then sums them with a pre-set confidence weight to obtain the joint state error value at the current time.
[0020] The parameter control module inputs the joint state error value into the proportional-integral controller (with the differential coefficient set to zero), outputs the parameter control increment value, maps it to the specific acoustic engine control command through a lookup table, and sends it to the audio processing engine to update the playback parameters of the target audio; at the same time, it calculates the variance of the joint state error value within the preset evaluation period as a feedback optimization parameter, and dynamically adjusts the preset system sampling rate and feature buffer queue depth accordingly.
[0021] Through the above scheme, this application achieves overshoot-free closed-loop control of multi-channel physiological signals to audio control parameters. Based on the characteristic buffer queue with pre-calibrated conduction time difference, the inherent misalignment in causal timing between the skin conductance and electrocardiogram channels is accurately compensated; the coordinated operation of the proportional-integral controller and the sampling rate adaptive mechanism further ensures the smoothness and stability of the system output throughout its entire life cycle.
[0022] This application further proposes the following process for obtaining the conduction delay of the skin conduction and the conduction delay of the electrocardiogram: a control command to trigger a preset test audio is sent to the audio processing engine, and the corresponding excitation timestamp is recorded. The instantaneous rate of change of the skin conduction physiological signal of the target object is obtained. When the instantaneous rate of change is greater than a preset rate of change threshold, the skin conduction mutation timestamp is recorded, and the difference between it and the excitation timestamp is used as the skin conduction delay. At the same time, the preset high-frequency band power of the electrocardiogram physiological signal of the target object is obtained, and the fitting slope of the preset high-frequency band power is calculated. When the fitting slope undergoes a sign reversal, the electrocardiogram reversal timestamp is recorded, and the difference between it and the excitation timestamp is used as the electrocardiogram conduction delay.
[0023] Skin conduction delay and electrocardiographic conduction delay are core configuration parameters of the system architecture. Both are calibrated by measuring the response time of different physiological channels to the same test stimulus, with the excitation time of the preset test audio as the zero point reference.
[0024] Specifically, the skin conduction delay reflects the time required for the sympathetic nervous system to receive auditory stimulation and for a significant change in skin resistance. Its calibration uses a differential algorithm to detect the instantaneous rate of change of the skin conduction signal. When the rate of change exceeds a preset rate of change threshold (i.e., a significant abrupt change beyond the resting noise baseline), it is recorded as the skin conduction response time. The electrocardiogram conduction delay reflects the time required for the vagus nervous system to receive auditory stimulation and for the high-frequency component of heart rate variability (HF-HRV) to reverse its trend. Its calibration is achieved by linearly fitting high-frequency power time-series data and detecting the sign reversal of the fitting slope.
[0025] Let the trigger timestamp be The timestamp of the skin conductance mutation is The ECG flip timestamp is Then the skin conduction delay With ECG conduction delay Calculated separately as follows: ; ; Because the response speed of the skin conductance signal is much faster than that of the heart rate variability signal, it must have a physical meaning. Conduction time difference This is the physical prerequisite for subsequently building the feature cache queue. This conduction delay varies from person to person and is affected by the sensor wearing position. Therefore, it needs to be recalibrated every time the system starts or the wearing status changes to ensure that the queue depth accurately matches the actual physiological response characteristics of the current user.
[0026] This application further proposes a feature cache queue construction process as follows: the difference between the ECG conduction delay and the skin conduction delay is used as the conduction time difference. The conduction time difference is compared with the preset system sampling rate. Multiply to obtain the sampling depth product, and then perform floor function on the sampling depth product to obtain the target depth. , Indicates to Round up and allocate a first-in-first-out (FIFO) structure queue in memory containing the same number of data nodes as the target depth, and use it as a feature cache queue.
[0027] The feature cache queue is the core data structure for achieving causal alignment in this application. Its design logic is as follows: if we want to achieve causal alignment at the current time... Simultaneously, the acquisition of skin charge has been delayed. The historical feature values and current ECG feature values after a certain number of seconds only need to be maintained in memory as a single database with a depth of [value missing]. A first-in-first-out queue is used for each data node. At the same time, the latest skin conductance characteristic value is pushed to the head of the queue during each sampling period, and the node popped from the tail of the queue is the one that has been delayed exactly. Each sampling period (corresponding to physical time) Skin conductance characteristics collected beforehand. Target depth. Due to conduction time difference With preset system sampling rate The product is determined by rounding up: ; The rounding up operation ensures that the queue depth in the discrete sampling domain is not less than the number of sampling points corresponding to consecutive physical time differences, avoiding timing undercompensation caused by truncation errors. The first-in-first-out (FIFO) circular buffer structure ensures that each popped node is always the earliest historical data from the current time, thus accurately reproducing the data in the continuous domain in the discrete sampling system. Time delay.
[0028] This application further proposes the following extraction process for skin electrophysiological feature values and cardiac electrophysiological feature values at the current moment: High-pass filtering is performed on the skin electrophysiological signal of the target object at the current moment to obtain transient AC components; the peak amplitude of the transient AC components within a preset time window is extracted; the peak amplitude is normalized to generate skin electrophysiological feature values; pulse wave period data representing the heartbeat interval is extracted based on the cardiac electrophysiological signal of the target object at the current moment; the ratio of high-frequency power to low-frequency power of the pulse wave period data within the preset time window is calculated; and the ratio is normalized to generate cardiac electrophysiological feature values.
[0029] Electrodermal physiological signals typically consist of a slowly drifting DC baseline component (reflecting the skin's basic conductivity) and a transient AC component superimposed on it (reflecting the immediate activation level of the sympathetic nervous system). High-pass filtering aims to remove baseline drift and retain the transient AC component, i.e., the skin conductance response (SCR), which sensitively reflects changes in emotional state. Based on this, the peak amplitude within a preset time window is extracted and normalized to obtain electrodermal physiological characteristic values normalized to the [0,1] interval. The higher the value, the stronger the current skin conductance activation. After the electrocardiographic signals are acquired by photoplethysmography (PPG) or electrocardiogram (ECG) sensors, pulse wave period data (i.e., heart rate variability time-domain sequence) is obtained by extracting adjacent heartbeat intervals (RR intervals). Within a preset time window, the high-frequency power (HF, 0.15–0.4 Hz, mainly reflecting vagal tone) and low-frequency power (LF, 0.04–0.15 Hz, reflecting mixed sympathetic and vagal regulation) of this sequence are calculated, and the ratio of these two values is normalized to obtain normalized electrocardiographic characteristic values. The trend of its numerical change can reflect the overall balance of the user's autonomic nervous system.
[0030] It is important to note that and All within system physical time They are generated synchronously, with the same sampling timestamp, but differ in their causal response to historical musical stimuli. The timing misalignment is the core issue that the subsequent second extraction module and error value evaluation module need to correct through the feature cache queue.
[0031] This application further proposes the following process for extracting the electrodermal feature values at the target historical moment: obtaining the number of data nodes currently stored in the feature cache queue, and determining whether the number of data nodes currently stored has reached the target depth. If it has not reached the target depth, the electrodermal feature values at the current moment are continuously pushed into the feature cache queue. If it has reached the target depth, the electrodermal feature values at the current moment are continuously pushed into the head of the feature cache queue, and the electrodermal feature values at the target historical moment are extracted and removed from the tail of the feature cache queue.
[0032] The feature cache queue's working mechanism embodies the core design idea of this application, trading memory space for alignment of physiological causal time. In the initial stage of system startup regulation (i.e., the queue is not full, corresponding to the previous...), (seconds), each sampling cycle only records the latest skin electrophysiological characteristic value Push it to the head of the queue without popping it out. At this point, the system forcibly skips the subsequent causal fusion steps and keeps the current audio parameters unchanged, ensuring that the system remains silent until the slow response channel (ECG) provides effective baseline feedback, and rejecting any blind adjustments.
[0033] When the number of data nodes accumulated in the queue reaches the target depth After that (i.e., entering the steady-state control period), each sampling period will update the latest... While pushing a historical node to the head of the queue, pop the earliest stored historical node from the tail of the queue according to the first-in, first-out (FIFO) principle. (Based on target depth) With conduction time difference The corresponding relationship means that the pop-up node must be at a physical time. The collected skin electrophysiological characteristics are denoted as This enables precise time-delay of skin conductance characteristics.
[0034] This application further proposes a process for obtaining a joint state error value through causal alignment fusion processing: obtaining the difference between the skin electrophysiological feature value at the target historical moment and the preset skin electrophysiological target benchmark value, and the difference between the electrocardiographic feature value at the current moment and the preset electrocardiographic target benchmark value; multiplying the skin electrophysiological feature difference value with a preset skin electrophysiological confidence weight to obtain the skin electrophysiological feature weighted error, multiplying the electrocardiographic feature difference value with a preset electrocardiographic confidence weight to obtain the electrocardiographic weighted error, and adding the skin electrophysiological feature weighted error and the electrocardiographic weighted error to obtain the joint state error value at the current moment.
[0035] Causal alignment and fusion is the theoretical core of the effectiveness of this application's scheme, and its correctness is based on the following causal logic: at the current moment... It has been delayed. The historical electrodermal characteristics that were subsequently displayed The physical time it was collected was The audio stimulus it responds to occurs in Current electrophysiological characteristics The audio stimulus that is responded to occurs in Therefore, both point to the same historical moment. The audio stimuli that occur are perfectly aligned in causality.
[0036] Based on the above alignment relationship, the two causally related physiological characteristic values are respectively compared with the preset target benchmark value (skin conductance target benchmark value). Compared with ECG target baseline value The difference is calculated, and weights reflecting the confidence levels of their respective signal engineering are used. and We perform a weighted summation to obtain the joint state error value at the current time. The calculation formula is as follows: ; in, To preset the confidence weights for skin conductance, To preset the ECG confidence weights, both are pre-configured by staff based on the engineering reliability of the two types of signals in specific user scenarios. In the formula... This represents the physical time difference between the conduction of auditory stimuli through the skin conduction and electrocardiogram channels; This indicates that the system retrieves past data through the feature cache queue. Electrodermal physiological characteristics at a given time.
[0037] The core difference between this application and the existing synchronous fusion equations lies in the application of a precise time index offset to the electrodermal characteristics. This mathematically-driven artificial time reversal eliminates the biological physical conduction phase difference within the computational domain, ensuring that the system consistently evaluates the comprehensive physiological outcome of the same stimulus source, fundamentally preventing control overshoot caused by logical inconsistencies. (Joint state error value) The symbols and amplitudes accurately reflect the degree of deviation of the current music parameters from the target relaxation state, providing a reliable and smooth control input for the downstream controller.
[0038] This application further proposes the following process for generating audio control parameters: inputting the joint state error value at the current moment into a preset proportional-integral controller model to output the parameter control increment value at the corresponding moment, wherein the differential term coefficient of the preset proportional-integral controller model is set to zero; matching the target adjustment rule in a preset mapping relationship table based on the parameter control increment value, extracting the acoustic engine control instruction corresponding to the target adjustment rule as the audio control parameter, and sending it to the audio processing engine to control the audio processing engine to update the playback parameters of the target audio.
[0039] Joint state error value The input signal, used as the input signal to the closed-loop controller, is fed into a preset proportional-integral (PI) controller model. The reason for discarding the derivative term (i.e., setting the derivative coefficient to zero) is that physiological signals themselves contain a certain amount of high-frequency noise components, and the derivative term would amplify the high-frequency changes of the signal significantly, introducing additional control jitter; the PI controller, while possessing sufficient steady-state accuracy, has good suppression capabilities for high-frequency noise, making it more suitable for audio parameter control scenarios driven by physiological signals.
[0040] After the PI controller outputs an incremental value, it matches the target adjustment rule in a preset mapping table using a lookup table. This converts the incremental value into a specific acoustic engine control command (such as changes in audio tempo BPM, volume attenuation in decibels for a specific frequency band, etc.), which is then sent to the audio processing engine via the communication bus or software interface. The engine then smoothly updates the audio playback parameters accordingly. Compared to real-time calculation of mapping functions, the lookup table method has lower computational latency and is more suitable for the deployment requirements of real-time embedded systems.
[0041] This application further proposes the following update process for the preset system sampling rate: obtaining the joint state error value and its average value of multiple consecutive target historical moments within a preset evaluation period; obtaining the variance of the joint state error value of the multiple consecutive target historical moments relative to the average value and using it as a feedback optimization parameter; when the feedback optimization parameter is greater than a preset upper limit threshold, increasing the preset system sampling rate by a preset adjustment step size based on the current preset system sampling rate to obtain the updated preset system sampling rate; When the feedback optimization parameter is less than the preset lower limit threshold, the preset system sampling rate at the current time is reduced by a preset adjustment step to obtain the updated preset system sampling rate; the updated preset system sampling rate is multiplied by the transmission time difference to recalculate the target depth, and the number of data nodes in the feature cache queue is adjusted based on the recalculated target depth.
[0042] The feedback optimization parameters characterize the stability of the music control effect by the variance of the joint state error value within a preset evaluation period: a large variance indicates that the system output is still fluctuating significantly and the control has not yet converged. In this case, the sampling rate should be increased to increase the control bandwidth and speed up the system response. A small variance indicates that the system has reached a steady state, and the sampling rate can be appropriately reduced to save computational resources. The above adaptive mechanism enables the system to automatically switch its working mode between the dynamic and stable stages of control, achieving a dynamic balance between computational efficiency and control accuracy.
[0043] It is worth noting that the preset system sampling rate Changes will synchronously trigger the target depth of the feature cache queue. Recalculation: ; Due to the conduction time difference The queue depth remains unchanged after calibration. With sampling rate It increases with the increase of, and with As the sampling rate decreases, the system adjusts the queue depth by increasing or decreasing the number of queue data nodes to ensure that the feature cache queue can accurately provide features under any sampling rate configuration. Duration-based delay compensation maintains the physical correctness of cross-channel causal alignment.
[0044] In another implementation, such as Figure 2 As shown, this application also provides a method for real-time dynamic control of music parameters based on multimodal physiological signals, including the following steps: Obtain the skin conduction delay and electrocardiogram conduction delay of the target object in response to a preset test audio, and obtain the conduction time difference based on the skin conduction delay and the electrocardiogram conduction delay; Construct a feature cache queue corresponding to the target depth based on the transmission time difference and the preset system sampling rate; During the playback of the target audio, the skin electrophysiological signal and electrocardiographic signal of the target object at the current moment are acquired synchronously according to the preset system sampling rate, and the skin electrophysiological feature value and electrocardiographic feature value at the current moment are extracted respectively. The current electrodermal feature value is pushed into the feature cache queue, and the electrodermal feature value of the target historical time is extracted. The time interval between the target historical time and the current time is equal to the conduction time difference. The skin electrophysiological characteristic values at the target historical moment and the electrocardiological characteristic values at the current moment are subjected to causal alignment and fusion processing to obtain the joint state error value at the current moment; Audio control parameters are generated based on the joint state error value, and the playback parameters of the corresponding target audio are updated accordingly to obtain feedback optimization parameters that characterize the music control effect, and the preset system sampling rate is updated accordingly.
[0045] During the system initialization phase, a control command to trigger a preset test audio is sent to the audio processing engine, and the excitation timestamp is recorded. Then, the instantaneous change rate of the skin electrophysiological signal and the high-frequency power fitting slope of the electrocardiographic signal are monitored respectively. The skin electrophysiological mutation timestamp and the electrocardiographic reversal timestamp are recorded according to the preset criteria. Then, the skin electrophysiological conduction delay and the electrocardiographic conduction delay are calculated, and the difference between the two is taken as the conduction time difference.
[0046] Based on the transmission time difference and the preset system sampling rate, calculate and allocate a first-in-first-out circular buffer queue containing the target depth number of data nodes, and complete the instantiation of the feature buffer queue.
[0047] During the playback of the target audio, the skin electrophysiological signal and the electrocardiographic signal are synchronously acquired at a preset system sampling rate. The skin electrophysiological feature values are obtained by high-pass filtering and normalization, and the electrocardiographic feature values are obtained by pulse wave period extraction and high-low frequency power ratio normalization.
[0048] The current electrodermal feature value is pushed into the feature buffer queue. After the queue reaches the target depth, the target historical electrodermal feature value is popped from the tail of the queue. This historical feature value points to the same historical audio stimulation cycle as the current electrocardiographic feature value.
[0049] The joint state error value at the current moment is obtained by subtracting the skin electrophysiological characteristic value at the target historical moment and the current electrocardiological characteristic value from their respective target baseline values, and then summing them with a pre-set confidence weight.
[0050] The joint state error value is input into the proportional-integral controller (with the differential coefficient set to zero), and the output parameter control increment value is mapped to the specific acoustic engine control command through a lookup table method. This command is then sent to the audio processing engine to update the playback parameters of the target audio. Simultaneously, the variance of the joint state error value within the preset evaluation period is calculated as a feedback optimization parameter, and the preset system sampling rate and feature buffer queue depth are dynamically adjusted accordingly.
[0051] In another embodiment, this application also provides a computer storage medium storing computer-executable instructions, which, when executed, implement the aforementioned real-time dynamic music parameter control system based on multimodal physiological signals.
[0052] The computer storage medium can be a non-volatile memory, including but not limited to flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), disk, optical disk, solid-state drive (SSD), and other forms of storage media. It can also be a composite storage device composed of a combination of the above types of storage media.
[0053] The computer-executable instructions stored therein can run as firmware on an embedded system (such as an ARM Cortex-M series microcontroller) or as a software process on a general-purpose computer operating system environment (such as Linux or Windows). The specific form of operation does not constitute a limitation on the scope of protection of this invention.
[0054] The above embodiments are only used to illustrate the technical methods of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of this application without departing from the spirit and scope of the technical methods of this application.
Claims
1. A real-time music parameter dynamic regulation system based on multi-modal physiological signals, characterized in that, Includes the following modules: The time difference acquisition module is used to acquire the skin conduction delay and electrocardiogram conduction delay of the target object in response to the preset test audio, and to obtain the conduction time difference based on the skin conduction delay and the electrocardiogram conduction delay; The queue construction module is used to construct a feature cache queue corresponding to the target depth based on the transmission time difference and the preset system sampling rate; The first extraction module is used to synchronously acquire the skin electrophysiological signal and electrocardiographic signal of the target object at the current moment according to the preset system sampling rate during the playback of the target audio, and extract the skin electrophysiological feature value and electrocardiographic feature value at the current moment respectively. The second extraction module is used to push the current electrodermal feature value into the feature cache queue and extract the electrodermal feature value of the target historical time, wherein the time interval between the target historical time and the current time is equal to the conduction time difference. The error value evaluation module is used to perform causal alignment fusion processing on the skin electrophysiological characteristic values of the target historical time and the electrocardiological characteristic values of the current time to obtain the joint state error value of the current time. The parameter control module is used to generate audio control parameters based on the joint state error value, update the playback parameters of the corresponding target audio according to the parameters, obtain feedback optimization parameters that characterize the music control effect, and update the preset system sampling rate according to the parameters.
2. The real-time dynamic music parameter control system based on multimodal physiological signals according to claim 1, characterized in that, The process of obtaining skin conduction delay and electrocardiographic conduction delay includes: Send a control command to the audio processing engine to trigger a preset test audio, and record the corresponding excitation timestamp. Obtain the instantaneous change rate of the target object's electrodermal physiological signal. When the instantaneous change rate is greater than a preset change rate threshold, record the electrodermal abrupt change timestamp, and use the difference between it and the excitation timestamp as the electrodermal conduction delay. Simultaneously, the preset high-frequency band power of the target object's electrocardiographic signal is acquired, and the fitting slope of the preset high-frequency band power is calculated. When the fitting slope undergoes sign reversal, the electrocardiographic reversal timestamp is recorded, and the difference between it and the excitation timestamp is used as the electrocardiographic conduction delay.
3. The real-time dynamic music parameter control system based on multimodal physiological signals according to claim 1, characterized in that, The process of building a feature cache queue includes: The difference between the electrocardiogram conduction delay and the skin conduction delay is taken as the conduction time difference. The conduction time difference is compared with the preset system sampling rate. Multiply to obtain the sampling depth product, and then perform floor function on the sampling depth product to obtain the target depth. , Indicates to Round up and allocate a first-in-first-out (FIFO) structure queue in memory containing the same number of data nodes as the target depth, and use it as a feature cache queue.
4. The real-time dynamic music parameter control system based on multimodal physiological signals according to claim 1, characterized in that, The process of extracting the current skin electrophysiological and cardiac electrophysiological characteristics includes: High-pass filtering is performed on the skin electrophysiological signal of the target object at the current moment to obtain transient AC component. The peak amplitude of the transient AC component within a preset time window is extracted, and the peak amplitude is normalized to generate skin electrophysiological feature value. Based on the target object's electrocardiographic signal at the current moment, pulse wave cycle data representing its heartbeat interval is extracted, and the ratio of high-frequency power to low-frequency power of the pulse wave cycle data within a preset time window is calculated. The ratio is then normalized to generate electrocardiographic feature values.
5. The real-time dynamic music parameter control system based on multimodal physiological signals according to claim 3, characterized in that, The process of extracting the electrodermal physiological characteristics of a target historical moment includes: The number of data nodes currently stored in the feature cache queue is obtained, and it is determined whether the number of data nodes currently stored has reached the target depth. If it has not reached the target depth, the current skin electrophysiological feature value is continuously pushed into the feature cache queue. When the target is reached, the current electrodermal feature value is continuously pushed into the head of the feature cache queue, and the corresponding target historical time's electrodermal feature value is extracted and removed from the tail of the feature cache queue.
6. The real-time dynamic music parameter control system based on multimodal physiological signals according to claim 1, characterized in that, The process of performing causal alignment fusion to obtain joint state error values includes: Obtain the skin electrophysiological characteristic value at the target historical moment and the skin electrophysiological characteristic value at the preset skin electrophysiological target reference value, as well as the electrocardiographic characteristic value at the current moment and the electrocardiographic characteristic value at the preset electrocardiographic target reference value; The skin conductance feature difference is multiplied by a preset skin conductance confidence weight to obtain the skin conductance weighted error. The electrocardiogram feature difference is multiplied by a preset electrocardiogram confidence weight to obtain the electrocardiogram weighted error. The skin conductance weighted error and the electrocardiogram weighted error are added together to obtain the joint state error value at the current time.
7. The real-time dynamic music parameter control system based on multimodal physiological signals according to claim 1, characterized in that, The process of generating audio control parameters includes: The joint state error value at the current moment is input into the preset proportional-integral controller model to output the parameter control increment value at the corresponding moment. The differential term coefficient of the preset proportional-integral controller model is set to zero. Based on the parameter control increment value, the target adjustment rule is matched in the preset mapping relationship table, and the acoustic engine control instruction corresponding to the target adjustment rule is extracted as the audio control parameter and sent to the audio processing engine to control the audio processing engine to update the playback parameters of the target audio.
8. The real-time dynamic music parameter control system based on multimodal physiological signals according to claim 3, characterized in that, The process of updating the preset system sampling rate includes: The joint state error value and its average value of multiple consecutive target historical moments within a preset evaluation period are obtained, and the variance of the joint state error value of the multiple consecutive target historical moments relative to the average value is obtained and used as a feedback optimization parameter. When the feedback optimization parameter is greater than the preset upper limit threshold, the preset system sampling rate at the current time is increased by a preset adjustment step size to obtain the updated preset system sampling rate. When the feedback optimization parameter is less than the preset lower limit threshold, the preset system sampling rate at the current time is decreased by a preset adjustment step size to obtain the updated preset system sampling rate. The updated preset system sampling rate is multiplied by the transmission time difference to recalculate the target depth, and the number of data nodes in the feature cache queue is adjusted based on the recalculated target depth.
9. A method for real-time dynamic control of music parameters based on multimodal physiological signals, characterized in that, Includes the following steps: Obtain the skin conduction delay and electrocardiogram conduction delay of the target object in response to a preset test audio, and obtain the conduction time difference based on the skin conduction delay and the electrocardiogram conduction delay; Construct a feature cache queue corresponding to the target depth based on the transmission time difference and the preset system sampling rate; During the playback of the target audio, the skin electrophysiological signal and electrocardiographic signal of the target object at the current moment are acquired synchronously according to the preset system sampling rate, and the skin electrophysiological feature value and electrocardiographic feature value at the current moment are extracted respectively. The current electrodermal feature value is pushed into the feature cache queue, and the electrodermal feature value of the target historical time is extracted. The time interval between the target historical time and the current time is equal to the conduction time difference. The skin electrophysiological characteristic values at the target historical moment and the electrocardiological characteristic values at the current moment are subjected to causal alignment and fusion processing to obtain the joint state error value at the current moment; Audio control parameters are generated based on the joint state error value, and the playback parameters of the corresponding target audio are updated accordingly to obtain feedback optimization parameters that characterize the music control effect. The preset system sampling rate is then updated based on these parameters.
10. A computer storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, they implement the real-time dynamic control system for music parameters based on multimodal physiological signals as described in any one of claims 1-8.