Multimodal music emotion interactive teaching system based on deep learning

CN122531638APending Publication Date: 2026-08-07XIAN AERONAUTICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN AERONAUTICAL UNIV
Filing Date
2026-05-18
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

然而,当这些现有技术应用于ASD儿童这一特殊群体时,大多致力于对当前时刻的情感状态进行分类(如“开心”或“悲伤”),这是一种静态的、滞后的判断,对于ASD儿童而言,当其情绪已被识别为“过度刺激”时,往往已接近或处于情绪崩溃的爆发点,此时再进行干预为时已晚,干预效果大打折扣,缺乏对其情感状态动态转移趋势进行预判的能力,无法在情绪恶化前实施“事前干预”

Benefits of technology

[0043] This invention utilizes an emotional state transition matrix model through a pre-judgment intervention module. Taking the previous emotional state and the current 128-dimensional joint feature vector as input, it calculates the probability of transitioning to each future state. This shifts the intervention timing from reactive remediation to proactive prevention of impending outbreaks. The pre-judgment advance constitutes a golden window for intervention, greatly improving the success rate and effectiveness of intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531638A_ABST
    Figure CN122531638A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal music emotion interactive teaching system based on deep learning, it is related to music emotion interactive teaching technical field, including: pre-judgment intervention module: for extracting previous moment emotional state based on historical database and judging whether it is calm state, the emotional state includes calm state, over-stimulation state and withdrawal state;Yes then the emotional state of previous moment and the 128-dimensional joint feature vector of current moment are input into pre-constructed emotional state transition matrix model, and the transition probability set of each emotional state is obtained in current moment. The pre-judgment intervention module of the application can calculate the probability of each state transition to the future, and the intervention opportunity is advanced from after-the-fact remedy to prevent in advance, and the pre-judgment advance constitutes the golden window of intervention, which greatly improves the success rate and effectiveness of intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of music emotional interaction teaching technology, specifically a multimodal music emotional interaction teaching system based on deep learning. Background Technology

[0002] Autism spectrum disorder (ASD) is a common neurodevelopmental disorder, with one of its core symptoms being impairments in social interaction and emotional engagement. Music therapy, as an effective adjunctive intervention, has been widely used in the rehabilitation training of children with ASD because it can bypass language barriers and directly affect emotions and perception. Traditional music teaching systems largely rely on teachers' visual observation and experience to perceive children's emotional states. This approach is highly subjective, difficult to quantify, and cannot provide real-time, precise responses.

[0003] With the development of technology, some intelligent emotion computing systems have emerged. These systems typically attempt to identify user emotions through one or more sensors (such as cameras, microphones, and wearable devices) and interact accordingly. For example, existing technologies include using cameras to capture facial expressions to identify emotions, or using heart rate bracelets to monitor physiological signals to infer excitement levels. However, when these existing technologies are applied to the specific group of children with ASD, most focus on classifying their current emotional state (such as "happy" or "sad"). This is a static and lagging judgment. For children with ASD, by the time their emotions are identified as "overstimulated," they are often close to or at the point of emotional breakdown. Intervention at this point is too late, and the intervention effect is greatly reduced. Furthermore, there is a lack of ability to predict the dynamic shifts in their emotional state, making it impossible to implement "pre-emptive intervention" before the emotions worsen. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a multimodal music emotion-interactive teaching system based on deep learning.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows:

[0006] A deep learning-based multimodal music emotional interaction teaching system includes a data acquisition module for non-contact acquisition of multimodal data on ASD children at the current moment, including physiological characteristics, micro-expression characteristics, and behavioral characteristics.

[0007] Data processing module: used to input the multimodal data of the current moment into the multimodal temporal alignment network for cross-scale spatiotemporal alignment and output a 128-dimensional joint feature vector;

[0008] Predictive intervention module: used to extract the emotional state of the previous moment based on the historical database and determine whether it is a calm state, the emotional state includes a calm state, an overstimulated state and a withdrawn state;

[0009] Then, the emotional state of the previous moment and the 128-dimensional joint feature vector of the current moment are input into the pre-constructed emotional state transition matrix model to obtain the transition probability set of each emotional state at the current moment;

[0010] Extract the transition probability values ​​from the calm state to the overstimulated state from the transition probability set. The probability value of transition from a peaceful state to a state of retreat. ,judge Does it exceed 0.7 or Does it exceed 0.6?

[0011] The intervention commands include music intervention commands executed by the directional sound field system and tactile intervention commands executed by the haptic feedback vest.

[0012] Preferably, in the prediction and intervention module, when At that time, an overstimulation intervention instruction is generated, specifically including:

[0013] Generates music intervention commands for low-frequency pulse beats of 55-65 BPM and narrow-band harmonics of 300-500 Hz;

[0014] Generate a tactile intervention command with a 3mm amplitude of gentle vibration.

[0015] Preferably, in the prediction and intervention module, when At that time, a withdrawal intervention instruction is generated, specifically including:

[0016] Music intervention commands that generate random rests, irregular rhythms, and sudden loud 82-88dB bursts lasting 0.5 seconds;

[0017] Generate tactile intervention commands for sudden strong vibrations with an amplitude of 8mm.

[0018] Preferably, the timing synchronization error between the music intervention command and the tactile intervention command is controlled within 5ms.

[0019] Preferably, in the prediction and intervention module and The threshold can be dynamically adjusted based on historical intervention effect data after extracting key probability values ​​and before determining whether the probability value exceeds the threshold. Specifically:

[0020] Record the effect data after each intervention. The effect data is determined by analyzing the new 128-dimensional joint feature vector and the emotional state classification results within a predetermined time window after the intervention.

[0021] If three consecutive interventions fail, the threshold is lowered by 0.05, i.e. and ;

[0022] If the success rate of the most recent 10 interventions is greater than 80%, increase the threshold by 0.03, i.e. and .

[0023] Preferably, the method for acquiring the multimodal data is as follows:

[0024] Physiological characteristics: Respiratory waveforms and heart rate variability spectra of children with ASD were acquired at a distance of 1.2 meters using a 60GHz millimeter-wave radar, and the data were output at a sampling rate of 10Hz.

[0025] Micro-expression features: 17 key points on the face are tracked by an infrared binocular camera with an 850nm wavelength infrared light source. The 0.5-2Hz tremor frequency of the orbicularis oculi muscle and the zygomaticus major muscle is extracted and the video stream data is output at a frame rate of 20fps.

[0026] Behavioral characteristics: The timing curve of the key touch force and the derivative of the key touch acceleration are captured in real time by a piezoelectric ceramic sensor array, and the data is output at a sampling rate of 20Hz.

[0027] Preferably, the multimodal data is input into a multimodal temporal alignment network for cross-scale spatiotemporal alignment, specifically as follows:

[0028] Long-period physiological rhythm features of respiratory waveforms were extracted using a one-dimensional convolutional gated recurrent unit network.

[0029] The tremor frequency is encoded using a bidirectional long short-term memory network based on optical flow vectors to encode its microsecond-level dynamic changes;

[0030] An attention-weighted mechanism is used to focus on key touch events when the playing intensity is applied, specifically touch events with a touch acceleration derivative greater than 0.3.

[0031] Output a 128-dimensional joint feature vector to achieve an alignment error of less than 100ms for three types of asynchronous data.

[0032] Preferably, the criteria for determining the emotional state are as follows:

[0033] In a calm state: the standard deviation of the respiratory waveform is less than 0.1, and the tremor frequency of the orbicularis oculi and zygomaticus major muscles is less than 0.8 Hz;

[0034] Overstimulation state: key acceleration derivative greater than 0.3, heart rate variability spectrum greater than 2;

[0035] Retraction state: The displacement of the orbicularis oculi muscle lasts for 5 seconds and is close to 0.

[0036] Preferably, after the predictive intervention module generates the intervention command, it immediately starts a 300ms state lock timer;

[0037] During the state lock-in period, the 128-dimensional joint feature vector of the new input is ignored, and the current emotional state judgment result is kept unchanged, avoiding high-frequency oscillations triggered by instantaneous feedback of intervention instructions or short-term stress behaviors.

[0038] Preferably, after the state lock ends, the normal monitoring and judgment process resumes;

[0039] Obtain the new 128-dimensional joint feature vector generated after executing the intervention command;

[0040] Input the current emotional state and the new 128-dimensional joint feature vector into the emotional state transition matrix model to calculate the probability of regression to a calm state.

[0041] If the calculated regression probability value is greater than 0.8, a feedback data point indicating the success of this intervention will be generated and recorded. This feedback data will be used for subsequent dynamic adjustment of the threshold.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] This invention utilizes an emotional state transition matrix model through a pre-judgment intervention module. Taking the previous emotional state and the current 128-dimensional joint feature vector as input, it calculates the probability of transitioning to each future state. This shifts the intervention timing from reactive remediation to proactive prevention of impending outbreaks. The pre-judgment advance constitutes a golden window for intervention, greatly improving the success rate and effectiveness of intervention. Attached Figure Description

[0044] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein:

[0045] Figure 1 This is a system diagram of the present invention;

[0046] Figure 2 This is a system flowchart of the present invention;

[0047] Figure 3 This is a timing diagram of the intervention in this invention. Detailed Implementation

[0048] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.

[0049] Existing technologies include methods to identify emotions by capturing facial expressions with cameras or to infer arousal levels by monitoring physiological signals with heart rate bracelets. However, when these technologies are applied to children with ASD, they mostly focus on classifying current emotional states (such as "happy" or "sad"), which is a static and lagging judgment. For children with ASD, by the time their emotions are identified as "overstimulated," they are often close to or at the point of emotional breakdown. Intervention at this point is too late, and the intervention effect is greatly reduced. Furthermore, there is a lack of ability to predict the dynamic shifts in their emotional states, making it impossible to implement "preemptive intervention" before emotions worsen. Therefore, a multimodal music-emotional interactive teaching system based on deep learning was designed. This system collects real data non-intrusively and performs high-precision cross-modal temporal alignment on the collected data. Based on the collected data, it achieves the transition from emotion recognition to emotion prediction. Using music and touch as mediums, it enables precise and predictive intervention and regulation of the emotional states of children with ASD. The invention is described in detail below:

[0050] like Figure 1-3 As shown, a multimodal music emotion-interactive teaching system based on deep learning includes:

[0051] Data acquisition module: Used for non-contact acquisition of multimodal data of children with ASD at the current moment, including physiological characteristics, micro-expression characteristics, and behavioral characteristics;

[0052] Specifically, because children with ASD generally exhibit sensory hypersensitivity and social anxiety, they are particularly sensitive to unfamiliar devices and contact sensors, such as heart rate monitors and EEG caps. Forcing them to wear these devices can trigger tension, anxiety, and even resistance in children with ASD, leading to abnormal physiological data, such as heart rate and skin conductance data, being under stress and failing to accurately reflect their natural emotional state in that situation. This results in all subsequent analyses being based on distorted data. Therefore, non-contact multimodal data collection can completely avoid the sensory discomfort, anxiety, and rejection reactions caused by contact sensors, such as heart rate monitors and EEG head-mounted devices, in children with ASD. This ensures from the source that the collected data represents the genuine and effective physiological and emotional responses of children with ASD in a natural state, laying a reliable foundation for accurate subsequent analysis. The specific collection method is as follows:

[0053] Physiological characteristics: The respiratory waveforms and heart rate variability spectrum of children with ASD were acquired at a distance of 1.2 meters using a 60GHz millimeter-wave radar and the data were output at a sampling rate of 10Hz. Respiratory rhythm and heart rate variability are important physiological indicators of emotional state. For example, when anxious, breathing becomes rapid and heart rate fluctuates abnormally. The non-contact acquisition method avoids the anxiety interference caused by contact with the device in children with ASD.

[0054] Micro-expression features: 17 key points on the face are tracked by an infrared binocular camera with an 850nm wavelength infrared light source. The 0.5-2Hz tremor frequency of the orbicularis oculi muscle and zygomaticus major muscle is extracted and the video stream data is output at a frame rate of 20fps. The emotional fluctuations of children with ASD are often manifested as subtle muscle activities, which are easily overlooked by traditional expression recognition. However, high-frequency tremors can reflect emotional changes at an earlier stage.

[0055] Behavioral characteristics: The timing curve of the touch force and the derivative of the touch acceleration of the electronic keyboard keys are captured in real time by a piezoelectric ceramic sensor array and the data is output at a sampling rate of 20Hz. Emotions directly affect the playing behavior. For example, children with ASD exert too much force when they are irritable and weaken their force when they withdraw. The derivative of touch acceleration can more sensitively characterize the emotionally driven abrupt changes in behavior.

[0056] Data processing module: used to input the multimodal data of the current moment into the multimodal temporal alignment network for cross-scale spatiotemporal alignment and output a 128-dimensional joint feature vector;

[0057] Specifically, while existing technologies mention multimodal fusion, they typically employ simple feature concatenation or weighted averaging, ignoring the significant differences in time scales between different modalities. For example, physiological signals (such as respiration and heart rate) change slowly, on the order of seconds or even minutes; while behavioral signals (such as instantaneous keystroke force or facial muscle tremors) are on the order of milliseconds. If this cross-scale asynchronous data cannot be precisely spatiotemporally aligned, direct fusion will lead to severe delays and misjudgments in emotion assessment, failing to capture the critical points of emotion transitions. By using a multimodal temporal alignment network, corresponding dedicated sub-networks are used to process data with different sampling rates and characteristics, ultimately controlling the alignment error to within 100ms. This achieves high-precision fusion of cross-scale asynchronous data, generating a 128-dimensional joint feature vector that characterizes a child's instantaneous state. This solves the information lag and distortion problems caused by traditional simple concatenation or interpolation fusion, greatly improving the temporal resolution of emotion assessment. The method of multimodal temporal alignment network for cross-scale spatiotemporal alignment of multimodal data is as follows:

[0058] One-dimensional convolutional gated recurrent unit network is used to extract the long-period physiological rhythm features of respiratory waveforms. One-dimensional convolutional gated recurrent unit network is good at capturing temporal dependencies and is suitable for processing slow-changing data at the second level.

[0059] The tremor frequency is encoded by a bidirectional long short-term memory network based on optical flow vectors to encode its microsecond-level dynamic changes. The optical flow features can capture the dynamic trajectory of muscle movement, and the bidirectional long short-term memory network can simultaneously focus on past and future tremor trends, accurately depicting instantaneous emotional fluctuations.

[0060] An attention-weighted mechanism is used to focus on key touch events when assessing performance intensity. Touch events with a derivative of touch acceleration greater than 0.3 are ignored, while irrelevant actions are ignored, highlighting emotionally driven behavioral characteristics.

[0061] The system outputs a 128-dimensional joint feature vector, achieving an alignment error of less than 100ms for the three types of asynchronous data, thus providing a unified data language for subsequent sentiment analysis.

[0062] Predictive intervention module: used to extract the emotional state of the previous moment based on the historical database and determine whether it is a calm state, the emotional state includes a calm state, an overstimulated state and a withdrawn state;

[0063] Then, the emotional state of the previous moment and the 128-dimensional joint feature vector of the current moment are input into the pre-constructed emotional state transition matrix model to obtain the transition probability set of each emotional state at the current moment;

[0064] Extract the transition probability values ​​from the calm state to the overstimulated state from the transition probability set. The probability value of transition from a peaceful state to a state of retreat. ,judge Does it exceed 0.7 or Does it exceed 0.6?

[0065] The intervention commands include music intervention commands executed by the directional sound field system and tactile intervention commands executed by the haptic feedback vest.

[0066] Specifically, most existing systems focus on classifying current emotional states, such as happiness or sadness. This is a static and lagging judgment. For children with ASD, by the time their emotions are identified as overstimulated, they are often close to or at the point of emotional breakdown. Intervention at this point is too late, and the effectiveness is greatly reduced. They lack the ability to predict the dynamic shifts in their emotional states, making it impossible to intervene before the emotions worsen. By using a predictive intervention module, an emotional state transition matrix model is employed. Taking the previous emotional state and the current 128-dimensional joint feature vector as input, the probability of transitioning to each future state is calculated. This shifts the intervention timing from reactive remediation to proactive prevention of impending outbreaks. This predictive lead time constitutes a golden window for intervention, greatly improving the success rate and effectiveness of intervention. The specific predictive intervention steps are as follows:

[0067] 1. Retrieve the emotional state from the historical database at the previous moment. The criteria for determining the emotional state are as follows:

[0068] In a calm state: the standard deviation of the respiratory waveform is less than 0.1, and the tremor frequency of the orbicularis oculi and zygomaticus major muscles is less than 0.8 Hz;

[0069] Overstimulation state: key acceleration derivative greater than 0.3, heart rate variability spectrum greater than 2;

[0070] Retraction state: The displacement of the orbicularis oculi muscle remains close to 0 for 5 seconds;

[0071] The emotional state of the previous moment is determined based on the multimodal data collected in the previous moment, and the determination result is saved to the historical database.

[0072] 2. Determine whether the emotional state extracted from the historical database was calm in the previous moment. Since the previous moment was in a state of overstimulation or withdrawal, it indicates that the emotion is abnormal, and the intervention logic will shift to how to guide it back to calm. However, the core here is early intervention. Therefore, we should focus on the case where the previous moment was calm. Moreover, a calm state is the foundation for effective teaching. Therefore, we should prioritize the probability value of transitioning from a calm state to an abnormal state. So, we should first determine whether the emotional state was calm in the previous moment. This can save unnecessary calculation of the transition probability value, improve calculation efficiency, and buy more time for early intervention.

[0073] If it is determined that the emotional state in the previous moment was not calm, it means that the previous moment was in a state of overstimulation or withdrawal, and the emotions have become abnormal and erupted. At this time, artificial intervention is needed to guide the emotions back to stability.

[0074] If the emotional state at the previous moment is determined to be calm, then the emotional state at the previous moment and the 128-dimensional joint feature vector at the current moment are input into the pre-constructed emotional state transition matrix model to obtain the transition probability value of the current moment to each emotional state.

[0075] The pre-constructed emotional state transition matrix model is as follows:

[0076] ;

[0077] i represents the emotional state at the previous moment;

[0078] j represents the emotional state that may shift at the current moment;

[0079] It represents the probability of transitioning from the emotional state of the previous moment to the emotional state that may be transitioned to in the current moment; This is the transpose of the weight vector, that is, converting the column vector into a row vector;

[0080] This is the multimodal joint feature vector, i.e., the 128-dimensional feature vector output by the data processing module;

[0081] This is the state transition bias term;

[0082] It is an exponential function;

[0083] The denominator is the sum of the exponential scores of all possible states. There are three corresponding emotional states, namely, the transition from state i to... The exponential scores in the three directions are summed.

[0084] By inputting the previous emotional state and the current 128-dimensional joint feature vector into a pre-constructed emotional state transition matrix model, we obtain:

[0085] The transition probability value from a calm state to a calm state

[0086] The probability value of transition from a calm state to a hyperstimulated state ;

[0087] The probability value of transitioning from a calm state to a retreat state ;

[0088] 3. Extract the transition probability values ​​of key emotional states requiring intervention, and determine the transition probability value from a calm state to an overstimulated state. Does it exceed 0.7, or is it the probability value of transitioning from a calm state to a retreating state? If the threshold exceeds 0.6, then based on the emotional shift state corresponding to the exceeded threshold, generate and execute the corresponding intervention instruction:

[0089] like If the system detects that a child with ASD is about to enter a state of overstimulation, such as agitation or resistance, an overstimulation intervention instruction is generated. This instruction, consisting of a low-frequency pulse beat of 55-65 BPM and a narrow-frequency harmonic of 300-500 Hz, is precisely projected onto the child's location via a directional sound field system. Simultaneously, a tactile intervention instruction with a 3mm amplitude of gentle vibration is output through a tactile feedback vest. This gentle vibration provides sensory stimulation to the child, and the coordinated execution of these two instructions regulates their emotions, thus enabling early intervention before the child with ASD enters a state of overstimulation, such as agitation or resistance.

[0090] like If a child with ASD is determined to be about to enter a state of lethargy or avoidance and withdrawal, a withdrawal intervention instruction is generated. The generated random rests and irregular rhythms, along with a sudden 82-88dB loud sound lasting 0.5 seconds, are precisely projected onto the location of the child with ASD through a directional sound field system. The generated tactile intervention instruction of sudden strong vibration with an amplitude of 8mm is output through a tactile feedback vest. The strong vibration provides sensory stimulation to the child with ASD. By coordinating the execution of the two instructions, emotions are regulated, thus achieving early intervention for the child with ASD who is about to enter a state of lethargy or avoidance and withdrawal.

[0091] In the predictive intervention module, when At that time, an overstimulation intervention instruction is generated, specifically including:

[0092] Generates music intervention commands for low-frequency pulse beats of 55-65 BPM and narrow-band harmonics of 300-500 Hz;

[0093] Generate a tactile intervention command with a 3mm amplitude of gentle vibration.

[0094] Specifically, the low-frequency pulse beat of 55-65 BPM is very close to the average heart rate of 60-100 BPM in healthy adults at rest. This physiological synchronization effect can gently guide the breathing and heart rate rhythm of ASD towards a slower and more stable state, effectively combating the rapid heartbeat and shortness of breath caused by anxiety, and providing a stable and predictable temporal framework to help children re-establish their internal order.

[0095] Furthermore, the 300-500Hz narrow harmonic band, due to the strong sound energy below 300Hz, can easily cause discomfort and resonance in the chest or abdomen, potentially exacerbating sensory burden. Sounds above 500Hz are relatively sharp and may be a stimulus rather than a soothing one for children with ASD who have hypersensitivity to sound. Therefore, the 300-500Hz frequency band sounds warm, soft, and full to the human ear, providing sufficient sensory filling without causing aggression. Its harmonic structure is simple and less likely to cause confusion in auditory processing.

[0096] Furthermore, the 3mm amplitude can be clearly perceived, but not so strong as to be surprising or uncomfortable. It is a deep, gentle pressure sensation, similar to a soothing, continuous feeling of being embraced, which helps to provide a calming effect through deep pressure sensation.

[0097] In the predictive intervention module, when At that time, a withdrawal intervention instruction is generated, specifically including:

[0098] Music intervention commands that generate random rests, irregular rhythms, and sudden loud 82-88dB bursts lasting 0.5 seconds;

[0099] Generate tactile intervention commands for sudden strong vibrations with an amplitude of 8mm.

[0100] Specifically, the irregular rhythm of random rests creates an organized irregularity, avoiding the chaos of complete disorder while breaking the inherent rhythms of ASD children in a withdrawn state. A restricted random algorithm is used to generate rhythmic patterns, with a basic cyclic unit of one or two measures in 4 / 4 time to ensure a basic musical structure. At each possible sixteenth note value starting point (i.e., within a measure, there are 16 potential positions), a rest is inserted with a 30% probability. The inserted rest values ​​range from 200ms, 400ms, to 800ms. The algorithm randomly selects from three predefined values ​​(ms). This cross-beat time value design aims to completely disrupt the predictability of the rhythm. The algorithm ensures that the total duration of continuous pauses does not exceed 1200ms to prevent children from falling back into silence. The generated rhythmic pattern has a minimum repetition cycle of no less than 8 measures to ensure that the subject cannot learn the pattern in a short period of time. This structured but random rhythm can effectively attract the auditory attention of children with ASD because their brains will subconsciously try to find the pattern, thereby passively pulling them out of their introverted and withdrawn state.

[0101] Furthermore, the sudden loud sound at 82-88 dB lasting 0.5 seconds, measured at a distance of 1 meter from the sound-generating unit of the directional sound field system, is strictly controlled below the safety line, far below the 120-140 dB that could cause hearing damage, but is sufficient to produce significant physiological arousal, such as a slight increase in heart rate and a shift in attention. It is a safe startle, and 0.5 seconds is enough for the brain to clearly perceive and process the sound event, forming a complete auditory impression, rather than a fleeting, easily ignored ticking sound. It is a brief pulse, not a continuous noise, and therefore will not evolve into a new unpleasant stimulus. Its purpose is to interrupt rather than cover up. This loud sound is not white noise, but a short sound with rich harmonic components.

[0102] The timing synchronization error between music intervention commands and tactile intervention commands is controlled within 5ms.

[0103] Specifically, the synchronization error between music intervention instructions and tactile intervention instructions is controlled within 5ms. Tactile vibrations and audio beat pulses are highly synchronized in time, resulting in a unified sensory experience. This enhances the overall integrity and predictability of the intervention, ensures the consistency of perception, and greatly improves the efficiency of soothing.

[0104] In the predictive intervention module and The threshold can be dynamically adjusted based on historical intervention effect data after extracting key probability values ​​and before determining whether the probability value exceeds the threshold. Specifically:

[0105] Record the effect data after each intervention. The effect data is judged by analyzing the new 128-dimensional joint feature vector and the emotional state classification results within the predetermined time window after the intervention.

[0106] If three consecutive interventions fail, the threshold is lowered by 0.05, i.e. and ;

[0107] If the success rate of the most recent 10 interventions is greater than 80%, increase the threshold by 0.03, i.e. and .

[0108] Specifically, existing systems employ fixed or preset intervention strategies, lacking the ability to self-optimize based on intervention effects. They cannot dynamically adjust based on each ASD child's real-time response to specific intervention methods, such as sensitivity to certain sound rhythms. This results in poor universality and reduced long-term effectiveness of intervention strategies. By introducing a dynamic threshold adjustment mechanism based on historical effects, the system is no longer rigid but can learn and adapt to each ASD child's unique response. Over time, the system's intervention for specific children becomes increasingly precise, truly realizing the concept of personalized medicine and improving long-term effectiveness.

[0109] Furthermore, after each intervention is triggered, the system observes and records the effect data within a predetermined time window to ensure objective and accurate judgment. After the intervention is implemented, the system continuously collects new multimodal data within a preset time window, and the data processing module generates a new 128-dimensional joint feature vector. Combining the new 128-dimensional joint feature vector with the emotional state, the system determines whether the intervention was successful.

[0110] If the emotional state changes from an impending state of overstimulation or withdrawal to a calm state after intervention, i.e., the new characteristics meet the criteria for a calm state: stable breathing, mild muscle tremors, etc., then the intervention is considered successful.

[0111] If the condition continues to deteriorate after intervention, such as the overstimulation characteristics not being relieved or the withdrawal state not improving, it is recorded as an intervention failure.

[0112] Furthermore, if the system detects three consecutive intervention failures, it indicates that the current threshold may be too high, causing the intervention to be triggered too late or not sensitive enough to the child's emotional changes. In this case, the threshold needs to be lowered.

[0113] The overstimulation threshold was reduced from the baseline value of 0.7 to 0.65 by 0.05, and the withdrawal threshold was reduced from the baseline value of 0.6 to 0.55 by 0.05. After the thresholds were reduced, the system became more sensitive to the probability of transitioning from a calm state to an abnormal state. Even if the probability of transition did not reach the original threshold, intervention would be triggered as long as it exceeded the new low threshold, thus intervening in the emotional change process earlier and avoiding intervention failure due to judgment delay.

[0114] Furthermore, when the system's statistics show a success rate exceeding 80% for the most recent 10 interventions, it indicates that the current threshold is well-suited to the child and the intervention is highly accurate. In this case, the threshold can be appropriately increased to reduce redundant interventions.

[0115] The overstimulation threshold was increased from the baseline value of 0.7 to 0.73 by 0.03, and the withdrawal threshold was increased from the baseline value of 0.6 to 0.63 by 0.03. After increasing the threshold, the system is more cautious in triggering interventions and only intervenes when the conversion probability is higher, thus avoiding over-intervention due to excessively low thresholds.

[0116] The specific methods for acquiring multimodal data are as follows:

[0117] Physiological characteristics: Respiratory waveforms and heart rate variability spectra of children with ASD were acquired at a distance of 1.2 meters using a 60GHz millimeter-wave radar, and the data were output at a sampling rate of 10Hz.

[0118] Micro-expression features: 17 key points on the face are tracked by an infrared binocular camera with an 850nm wavelength infrared light source. The 0.5-2Hz tremor frequency of the orbicularis oculi muscle and the zygomaticus major muscle is extracted and the video stream data is output at a frame rate of 20fps.

[0119] Behavioral characteristics: The timing curve of the key touch force and the derivative of the key touch acceleration are captured in real time by a piezoelectric ceramic sensor array, and the data is output at a sampling rate of 20Hz.

[0120] Multimodal data is input into a multimodal temporal alignment network for cross-scale spatiotemporal alignment, as detailed below:

[0121] Long-period physiological rhythm features of respiratory waveforms were extracted using a one-dimensional convolutional gated recurrent unit network.

[0122] The tremor frequency is encoded using a bidirectional long short-term memory network based on optical flow vectors to encode its microsecond-level dynamic changes;

[0123] An attention-weighted mechanism is used to focus on key touch events when the playing intensity is applied, specifically touch events with a touch acceleration derivative greater than 0.3.

[0124] Output a 128-dimensional joint feature vector to achieve an alignment error of less than 100ms for three types of asynchronous data.

[0125] The criteria for determining emotional state are as follows:

[0126] In a calm state: the standard deviation of the respiratory waveform is less than 0.1, and the tremor frequency of the orbicularis oculi and zygomaticus major muscles is less than 0.8 Hz;

[0127] Overstimulation state: key acceleration derivative greater than 0.3, heart rate variability spectrum greater than 2;

[0128] Retraction state: The displacement of the orbicularis oculi muscle lasts for 5 seconds and is close to 0.

[0129] After the predictive intervention module generates an intervention command, it immediately starts a 300ms state lock timer;

[0130] During the state lock-in period, the 128-dimensional joint feature vector of the new input is ignored, and the current emotional state judgment result is kept unchanged, avoiding high-frequency oscillations triggered by instantaneous feedback of intervention instructions or short-term stress behaviors.

[0131] Specifically, because children with ASD may experience transient stress responses to sensory stimuli, but these responses are not genuine changes in emotional state, misjudgment by the system can lead to a vicious cycle of intervention → stress → misjudgment → new intervention, i.e., high-frequency oscillations. Therefore, at the same moment the intervention prediction module generates and issues the intervention instruction, a 300ms countdown is immediately initiated. The system temporarily ignores the newly input 128-dimensional joint feature vector from the data processing module, i.e., it does not use the new features for state transition probability calculation, and forcibly maintains the current emotional state judgment result unchanged. For example, if the intervention is triggered and the system determines that excessive stimulation is imminent... If the system is activated, it maintains this determination during the lockout period and does not update to the state that the new feature might imply. After 300ms, the system resumes normal operation, receives and processes the new 128-dimensional joint feature vector, and updates the state determination based on the latest data. By briefly locking the state determination after the intervention is initiated, the system ignores the instantaneous stress signals caused by the intervention itself, avoids the system overreacting to false changes and falling into a vicious cycle of high-frequency interventions, ensures the stability of emotional state judgment and the consistency of intervention strategies, and allows the system to more accurately distinguish between stress interference and real emotional changes, thereby improving the reliability of the intervention effect.

[0132] After the status lock ends, the normal monitoring and judgment process resumes;

[0133] Obtain the new 128-dimensional joint feature vector generated after executing the intervention command;

[0134] Input the current emotional state and the new 128-dimensional joint feature vector into the emotional state transition matrix model to calculate the probability of regression to a calm state.

[0135] If the calculated regression probability value is greater than 0.8, a feedback data point indicating the success of this intervention will be generated and recorded. This feedback data will be used for subsequent dynamic adjustment of the threshold.

[0136] Specifically, the process after the state lock ends is a crucial step for the system to accurately evaluate the intervention effect and form a closed-loop optimization. The core purpose is to determine whether the intervention is truly effective and to use the results for subsequent threshold adjustments, allowing the system to continuously adapt to the emotional characteristics of children with ASD. After the 300ms state lock ends, the system will immediately resume the normal data collection → feature processing → state judgment process:

[0137] The data acquisition module resumed capturing multimodal data such as breathing, micro-expressions, and touch behavior in real time. At this time, the momentary stress response caused by the intervention command has been filtered out, such as the physiological fluctuations caused by a brief shock from vibration.

[0138] The data processing module generates a new 128-dimensional joint feature vector based on the newly collected data. These features reflect the changes in the child's actual emotional state after the intervention instructions are executed, such as whether breathing has become stable and muscle tremors have been relieved after overstimulation intervention.

[0139] Furthermore, the system inputs the current emotional state—the abnormal state before intervention, such as overstimulation or withdrawal—along with a new 128-dimensional joint feature vector into the emotional state transition matrix model, focusing on calculating the regression probability to a calm state:

[0140] If the current state is an overstimulated state, calculate the probability of the overstimulated state turning into a calm state, that is, the possibility of the emotion changing from irritability to stability.

[0141] If the current state is a state of withdrawal, calculate the probability of the withdrawal state turning into a calm state, that is, the possibility of the emotion changing from stagnation to participation.

[0142] This probability is a quantitative benchmark for evaluating the effectiveness of intervention: the higher the probability, the more effective the intervention and the more likely the child's emotions are to stabilize; the lower the probability, the less effective the intervention is and the more necessary it is to adjust the strategy.

[0143] A deep learning-based multimodal music emotion-interaction teaching method includes the following steps:

[0144] Step 1: Non-contact collection of multimodal data of the child with ASD at the current moment, including physiological characteristics, micro-expression characteristics, and behavioral characteristics;

[0145] Step 2: Input the multimodal data at the current moment into the multimodal temporal alignment network to perform cross-scale spatiotemporal alignment and output a 128-dimensional joint feature vector;

[0146] Step 3: Used to extract the emotional state of the previous moment based on the historical database and determine whether it is a calm state. The emotional state includes a calm state, an overstimulated state, and a withdrawn state.

[0147] Then, the emotional state of the previous moment and the 128-dimensional joint feature vector of the current moment are input into the pre-constructed emotional state transition matrix model to obtain the transition probability set of each emotional state at the current moment;

[0148] Extract the transition probability values ​​from the calm state to the overstimulated state from the transition probability set. The probability value of transition from a peaceful state to a state of retreat. ,judge Does it exceed 0.7 or Does it exceed 0.6?

[0149] The intervention commands include music intervention commands executed by the directional sound field system and tactile intervention commands executed by the haptic feedback vest.

[0150] Real-time example 1:

[0151] Experimental preparation and implementation process: In order to verify the effectiveness, advancement and inventiveness of the system described in this invention, this embodiment designed a clinical controlled trial to compare with the prior art.

[0152] Participants: Sixty children aged 6-12 years with ASD diagnosed by a professional institution and of similar functional levels were recruited and randomly assigned to an experimental group (n=30) and a control group (n=30). All children participated with informed consent, and were led by the same experienced music therapist in a 4-week group music teaching course, twice a week for 30 minutes each time. The course content was interactive electronic keyboard playing, including rhythmic changes that could easily induce emotional fluctuations and social interaction.

[0153] Test environment and equipment:

[0154] Experimental group: Using the complete system described in this invention, a 60GHz millimeter-wave radar (model: TIAWR1843) and an 850nm infrared binocular camera (model: Intel RealSense D455) were installed on the top of the classroom, 2.5 meters above the ground, to monitor the entire classroom without being noticed. A piezoelectric ceramic sensor array (sensitivity: 10pC / N) was integrated under the electronic keyboard keys. A directional sound field system (model: USCH HyperSonicSound) was deployed around the classroom. Children wore haptic feedback vests with built-in linear resonant actuators (LRA). All data was processed in real time by an edge computing workstation (NVIDIA Jetson AGX Orin).

[0155] Control group: Using existing technology, children wore a chest-worn heart rate monitor (PolarH10) to collect physiological data, and consumer-grade emotion recognition glasses (JinsMeme) to collect eye movement and facial electromyography (EMG) data. Behavioral data was captured using a regular camera. The system used a post-processing feature stitching and fusion algorithm for emotion recognition (current state classification). When overstimulation or withdrawal was detected, the teacher manually selected and played pre-stored soothing or rousing music, without tactile feedback.

[0156] Implementation details and optimizations:

[0157] 1. Data Collection and Authenticity Assurance: After the experiment began, the data of 12 children in the control group were marked as invalid because they strongly resisted the heart rate monitor and glasses, pulled on the equipment, or showed obvious anxiety. However, all children in the experimental group did not show any stress response to the non-contact monitoring device, and the data validity rate reached 100%. This initially proved the fundamental advantage of non-contact data collection in solving the problem of data distortion in ASD data sources.

[0158] 2. Data Processing and Predictive Triggering: In a typical teaching session, when the pace of the lesson accelerated, the experimental group system processed data in real time using the MTAN network (pre-trained with 10,000 sets of ASD data). For example, the system detected that a child's respiratory waveform standard deviation rose to 0.15 (resting <0.1), the zygomaticus major muscle tremor frequency reached 1.6Hz (resting <0.8Hz), and the derivative of key-touch acceleration showed three peaks greater than 0.4 within 1 second. The MTAN network completed alignment and fusion within 80ms, generating a 128-dimensional joint vector, from which the emotional state transition matrix was calculated. (Exceeding the dynamic threshold of 0.71). The system issues an overstimulation warning 3.5 seconds in advance.

[0159] 3. Precise Intervention and Closed-Loop Feedback: The system immediately generates intervention instructions: a 60 BPM beat, 420Hz harmonics, and synchronous slow vibration with an amplitude of 3mm (phase difference optimized and controlled within 3ms by hardware). The sound field system precisely focuses the sound on the child to avoid disturbing others. The tactile vest synchronously provides stable rhythmic vibrations. After intervention, the system initiates a 300ms state lock, ignoring the child's brief startled reaction. After the lock ends, the system calculates the probability of regression. If the value is greater than 0.8, the intervention is considered successful, and this result is recorded for subsequent threshold optimization (the success rate is improved, and the threshold is then fine-tuned to 0.72).

[0160] 4. Control group comparison: At the same time, the control group system was greatly affected by motion interference in the heart rate band signal and the emotional glasses were frequently touched by the child. Its simple fusion algorithm did not alarm to indicate overstimulation until the child began to show the behavior of throwing the piano keys (2 seconds later). The delay in manual intervention by the teacher led to a full-blown emotional outburst and interruption of teaching.

[0161] Throughout the entire test period, the system recorded data for every single event.

[0162] Comparison table of test data

[0163] The table below shows the average values ​​of key performance data collected during the trial period.

[0164] Effective data collection rate 100% 60% +66.7% (Completely eliminates equipment rejection) Emotion recognition accuracy 91.5% 68.0% +34.6% (High-fidelity data + Advanced fusion) Average forecast lead time +3.2 seconds -1.5 seconds (lag) +4.7 seconds (from post-event remediation to pre-event prevention) Overstimulation intervention success rate 89.7% 45.2% (Teacher's manual input) +98.5% (Precise parametric intervention) Success rate of intervention for withdrawal status 82.4% 38.1% (Teacher's manual input) +116.3% (Strong arousal stimulation is effective) Average time to calm down 2.4 minutes 6.8 minutes -64.7% (Significantly improved teaching efficiency) Number of interruptions in a single teaching session 0.3 times 2.1 times -85.7% (Ensuring smooth teaching flow) System false alarm rate 5.8% 15.3% -62.1% (Dynamic threshold optimization played a crucial role) Cross-modal synchronization error < 5ms N / A (This feature is not available) Achieve μs-level synchronization and a unified user experience Child participation (teacher assessment) 4.6 / 5.0 2.8 / 5.0 +64.3% (Stress-free interaction increases interest)

[0165] Compared with existing technologies:

[0166] 1. The existing technology (control group) with an effective acquisition rate of 60% exposes the fatal flaw of contact-based devices, as the device itself is a source of interference. This system achieves 100% effective acquisition through non-contact multimodal sensor fusion, which means that all subsequent analyses are based on real and reliable data.

[0167] 2. Existing technologies, with a 1.5-second lag, are essentially "hindsight biases," only offering remedies after an emotional outburst. Our system, however, provides a revolutionary 3.2-second lead time. This is thanks to the precise instantaneous state representation provided by MTAN's high-precision alignment and the trend prediction capabilities of the emotional state transition matrix. This allows the system to intervene before the emotional crisis reaches its breaking point, nipping problems in the bud, resulting in significant technological effectiveness.

[0168] 3. The nearly 90% intervention success rate of this system contrasts sharply with the less than 50% success rate of existing technologies (even manual intervention by teachers). This is not because this system is smarter than the therapist, but because the series of parameterized instructions it provides—55-65 BPM, 300-500 Hz, 3mm / 8mm amplitude, and <5ms synchronization—are standardized prescriptions based on neuroscience principles, capable of accurately regulating the target physiological state. In contrast, teacher intervention is more subjective, delayed, and unable to provide tactilely synchronized stimulation.

[0169] 4. The system's low false alarm rate of 5.8% and dynamically changing thresholds (e.g., adjusted from 0.70 to 0.72) demonstrate the effectiveness of its feedback-based dynamic threshold adjustment mechanism. The system is no longer rigid but learns from historical interventions, continuously fine-tuning its predictive sensitivity and becoming increasingly knowledgeable about the child at hand. This self-evolutionary capability is a significant advancement not found in existing static systems.

[0170] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.

Claims

1. A multimodal music emotional interaction teaching system based on deep learning, characterized in that, include: Data acquisition module: Used for non-contact acquisition of multimodal data of children with ASD at the current moment, including physiological characteristics, micro-expression characteristics, and behavioral characteristics; Data processing module: used to input the multimodal data of the current moment into the multimodal temporal alignment network for cross-scale spatiotemporal alignment and output a 128-dimensional joint feature vector; Predictive intervention module: used to extract the emotional state of the previous moment based on the historical database and determine whether it is a calm state, the emotional state includes a calm state, an overstimulated state and a withdrawn state; Then, the emotional state of the previous moment and the 128-dimensional joint feature vector of the current moment are input into the pre-constructed emotional state transition matrix model to obtain the transition probability set of each emotional state at the current moment; Extract the transition probability values ​​from the calm state to the overstimulated state from the transition probability set. The probability value of transition from a peaceful state to a state of retreat. ,judge Does it exceed 0.7 or Does it exceed 0.6? The intervention commands include music intervention commands executed by the directional sound field system and tactile intervention commands executed by the haptic feedback vest.

2. The deep learning-based multimodal music emotion interaction teaching system according to claim 1, characterized in that: In the prediction and intervention module, when At that time, an overstimulation intervention instruction is generated, specifically including: Generates music intervention commands for low-frequency pulse beats of 55-65 BPM and narrow-band harmonics of 300-500 Hz; Generate a tactile intervention command with a 3mm amplitude of gentle vibration.

3. The deep learning-based multimodal music emotion interaction teaching system according to claim 2, characterized in that: In the prediction and intervention module, when At that time, a withdrawal intervention instruction is generated, specifically including: Music intervention commands that generate random rests, irregular rhythms, and sudden loud 82-88dB bursts lasting 0.5 seconds; Generate tactile intervention commands for sudden strong vibrations with an amplitude of 8mm.

4. The deep learning-based multimodal music emotion interaction teaching system according to claim 3, characterized in that: The timing synchronization error between the music intervention command and the tactile intervention command is controlled within 5ms.

5. The deep learning-based multimodal music emotion interaction teaching system according to claim 3, characterized in that: The prediction and intervention module and The threshold can be dynamically adjusted based on historical intervention effect data after extracting key probability values ​​and before determining whether the probability value exceeds the threshold. Specifically: Record the effect data after each intervention. The effect data is determined by analyzing the new 128-dimensional joint feature vector and the emotional state classification results within a predetermined time window after the intervention. If three consecutive interventions fail, the threshold is lowered by 0.05, i.e. and ; If the success rate of the most recent 10 interventions is greater than 80%, increase the threshold by 0.03, i.e. and .

6. The deep learning-based multimodal music emotion interaction teaching system according to claim 1, characterized in that: The specific methods for acquiring the multimodal data are as follows: Physiological characteristics: Respiratory waveforms and heart rate variability spectra of children with ASD were acquired at a distance of 1.2 meters using a 60GHz millimeter-wave radar, and the data were output at a sampling rate of 10Hz. Micro-expression features: 17 key points on the face are tracked by an infrared binocular camera with an 850nm wavelength infrared light source. The 0.5-2Hz tremor frequency of the orbicularis oculi muscle and the zygomaticus major muscle is extracted and the video stream data is output at a frame rate of 20fps. Behavioral characteristics: The timing curve of the key touch force and the derivative of the key touch acceleration are captured in real time by a piezoelectric ceramic sensor array, and the data is output at a sampling rate of 20Hz.

7. The deep learning-based multimodal music emotion interaction teaching system according to claim 6, characterized in that: The multimodal data is input into a multimodal temporal alignment network for cross-scale spatiotemporal alignment, as detailed below: Long-period physiological rhythm features of respiratory waveforms were extracted using a one-dimensional convolutional gated recurrent unit network. The tremor frequency is encoded using a bidirectional long short-term memory network based on optical flow vectors to encode its microsecond-level dynamic changes; An attention-weighted mechanism is used to focus on key touch events when the playing intensity is applied, specifically touch events with a touch acceleration derivative greater than 0.

3. Output a 128-dimensional joint feature vector to achieve an alignment error of less than 100ms for three types of asynchronous data.

8. The deep learning-based multimodal music emotion interaction teaching system according to claim 6, characterized in that: The criteria for determining the emotional state are as follows: In a calm state: the standard deviation of the respiratory waveform is less than 0.1, and the tremor frequency of the orbicularis oculi and zygomaticus major muscles is less than 0.8 Hz; Overstimulation state: key acceleration derivative greater than 0.3, heart rate variability spectrum greater than 2; Retraction state: The displacement of the orbicularis oculi muscle lasts for 5 seconds and is close to 0.

9. The deep learning-based multimodal music emotion interaction teaching system according to claim 1, characterized in that: After the predictive intervention module generates the intervention command, it immediately starts a 300ms state lock timer. During the state lock-in period, the 128-dimensional joint feature vector of the new input is ignored, and the current emotional state judgment result is kept unchanged, avoiding high-frequency oscillations triggered by instantaneous feedback of intervention instructions or short-term stress behaviors.

10. The deep learning-based multimodal music emotion interaction teaching system according to claim 9, characterized in that: After the state lock ends, the normal monitoring and judgment process resumes; Obtain the new 128-dimensional joint feature vector generated after executing the intervention command; Input the current emotional state and the new 128-dimensional joint feature vector into the emotional state transition matrix model to calculate the probability of regression to a calm state. If the calculated regression probability value is greater than 0.8, a feedback data point indicating the success of this intervention will be generated and recorded. This feedback data will be used for subsequent dynamic adjustment of the threshold.