Multi-modal sensor signal analysis method and system for sleep emotional state recognition
By analyzing the timing interference of multimodal signal transmission and performing dynamic calibration, the problems of signal timing deviation and misjudgment in sleep emotional state recognition are solved, achieving high-precision sleep emotional state recognition and personalized control, and improving the effectiveness and security of intelligent feedback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHWEST MEDICAL UNIV
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies fail to effectively address real-time fluctuations in multimodal data synchronization and differences in physiological signal acquisition in dynamic scenarios during sleep emotion state recognition. This leads to signal timing deviations and misjudgments, affecting the accuracy of emotion-sleep state recognition and the effectiveness of intelligent feedback regulation.
By analyzing the timing interference of multimodal signal transmission, it is determined whether to activate fixed clock calibration alignment or dynamic resynchronization of multimodal signals, and to verify the timing logic of sleep physiological signals to ensure the timing consistency and physiological logic rationality of multimodal signals. Based on this, feature extraction and fusion are performed, combined with personalized regulation and safety protection, to achieve accurate identification and intelligent regulation of sleep emotional state.
It improves the accuracy of sleep mood state recognition and the safety of intelligent regulation, reduces the risk of misjudgment and over-regulation, ensures the temporal consistency of multimodal signals and the rationality of physiological logic, and realizes rapid and accurate sleep mood state recognition and personalized regulation.
Smart Images

Figure CN121506511B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sleep emotion data processing technology, and in particular to a multimodal sensor signal analysis method and system for sleep emotion state recognition. Background Technology
[0002] Sleep quality is closely related to emotional health. Nighttime mood fluctuations not only affect sleep structure but are also an important early manifestation of mental disorders (such as anxiety and depression). People's sleep processes are often disturbed by unconscious emotional fluctuations caused by environmental or psychological stress, yet they cannot record their emotional state through voluntary feedback. However, home sleep aids and mood relaxation pods are available in homes, transportation rest areas, or ICUs (Intensive Care Units). The intelligent comfort system for ICU caregivers, designed for scenarios like the ICU care area, possesses technology for non-intrusive, continuous, and accurate identification of emotional states. This technology holds significant scientific value for early warning of emotional disorders, personalized assessment and intervention of sleep quality, and exploring the neural mechanisms of sleep-emotional regulation. The existing process for achieving this sleep-emotional state identification mainly involves: First, in the data acquisition phase, a multi-sensor array deployed on a non-invasive headband, smart mattress, or patch simultaneously collects multimodal raw signals during sleep, including EEG, ECG, skin conductance, body surface data, and environmental audio. These raw signals are then transmitted to a sleep-emotional state data processing center. Next, the raw signals are denoised (e.g., using filters to remove power frequency interference), segmented (by fixed time windows or based on events), and timestamped to generate a multimodal preprocessed signal. Then… Subsequently, in the feature extraction and fusion stage, features are extracted from the multimodal preprocessed signals based on the multimodal Transformer network and cross-correlation modeling is performed (i.e., the correlation weights between different modal signal features are dynamically calculated and fused through the attention mechanism in the network), outputting a unified multimodal fused emotional state feature vector. Then, in the emotional state modeling and recognition stage, the multimodal fused emotional state feature vector is input into a multi-classification model (such as a support vector machine, random forest, or deep neural network classifier) built based on the multimodal fused emotional state feature vector to identify specific emotional and sleep states. Afterward, in the intelligent feedback regulation process, based on the identified emotional and sleep states, environmental stimuli (such as sound, light, temperature, and touch) are adjusted in real time through regulation units (such as visual regulation units and auditory regulation units) to form a closed-loop neural feedback regulation.
[0003] For example, the Chinese invention patent with publication number CN119538051A discloses a multi-view fusion emotion recognition method under sleep deprivation conditions, which includes: constructing and training a recognition model containing a multi-view fusion Transformer network in an offline stage, and using the trained recognition model to perform multi-view fusion on the input signal in an online stage.
[0004] The above-mentioned technology has at least the following technical problems:
[0005] Existing technologies for sleep mood state recognition complete signal processing and recognition by training specific fusion models offline (such as multi-view fusion Transformer networks), using fixed clock calibration to achieve data synchronization, or reducing multimodal data dependence. However, they do not consider the inherent differences in the acquisition responses of different physiological signals under dynamic scenarios such as limb micro-movements during sleep, as well as the weak electromagnetic interference brought by the acquisition equipment. Furthermore, they lack the ability to dynamically calibrate against such real-time interference, making it difficult to cope with real-time fluctuations in multimodal data synchronization.
[0006] When collecting multimodal physiological signals related to sleep emotional state and stress level during sleep in dynamic scenarios such as turning over and micro-limb movements during sleep, the inherent differences in the response speed of physiological signals, such as EEG (Electroencephalogram), HRV (Heart Rate Variability), PPG (Photoplethysmography), GSR (Galvanic Skin Response), and respiration, can lead to problems. Furthermore, weak electromagnetic interference in the acquisition devices (such as dry electrode EEG sensors, flexible fabric sensing modules, and photoplethysmography sensors) can affect signal transmission timing. Current technologies rely solely on fixed clock calibration for multimodal data synchronization in dynamic scenarios, lacking a dynamic compensation mechanism for real-time interference. This can result in millisecond-level alignment deviations in the time axis of different modal data. Such deviations disrupt the inherent physiological timing logic of each signal reflecting the same stress event (e.g., the rapid rise in skin conductance during a stress event should slightly precede the acceleration of heart rate), thus causing cross-modal phases during feature fusion. Distortion in the modeling of emotional states leads to fusion bias in the multimodal fusion of emotional state feature vectors. Consequently, when identifying emotion-sleep states based on these feature vectors, misjudgments occur (e.g., a brief increase in GSR or HRV fluctuations caused by turning over is misjudged as an anxiety attack or enhanced stress response; brief awakenings during light sleep are misjudged as drowsiness subsidence or stress relief). Ultimately, during intelligent feedback regulation, the feedback strategy becomes disconnected from the actual emotion-sleep state and stress state (e.g., non-stressful signal fluctuations are mistakenly identified as stress responses, leading to excessive activation of low-frequency soothing sounds or inappropriate increases in body surface temperature). This results in low accuracy in emotion-sleep state identification and intelligent feedback regulation. Summary of the Invention
[0007] This invention provides a multimodal sensor signal analysis method and system for sleep mood state recognition. The technical solution provided by this application is as follows:
[0008] Firstly, a multimodal sensor signal analysis method for sleep emotion state recognition is provided. The specific implementation of this method is as follows: Multimodal signal transmission timing interference analysis is performed, and it is determined whether fixed clock calibration alignment and dynamic resynchronization of multimodal signals should be initiated. If initiated, after the initiation is complete, sleep physiological signal timing logic verification is performed, and based on the obtained sleep physiological signal timing logic verification results, it is determined whether to perform multimodal sleep emotion data feature extraction and fusion. If not initiated, multimodal sleep emotion data feature extraction and fusion are performed, and sleep-emotion state recognition is performed based on the obtained multimodal sleep emotion data feature extraction and fusion results, and sleep-emotion state recognition results are obtained. Based on the sleep-emotion state recognition results, sleep emotion collaborative intelligent regulation is performed, and sleep emotion responsive safety protection is implemented during the sleep emotion collaborative intelligent regulation process.
[0009] Secondly, a multimodal sensor signal analysis system for sleep emotion state recognition is provided. This system applies a multimodal sensor signal analysis method for sleep emotion state recognition, including: a multimodal signal timing calibration and verification module, a multimodal feature fusion and state recognition module, and a sleep collaborative regulation and safety protection module. The multimodal signal timing calibration and verification module performs multimodal signal transmission timing interference analysis and determines whether to initiate fixed clock calibration alignment and dynamic resynchronization of multimodal signals. If initiated, it performs sleep physiological signal timing logic verification after activation and determines whether to perform multimodal sleep emotion data feature extraction and fusion based on the obtained sleep physiological signal timing logic verification results. The multimodal feature fusion and state recognition module performs multimodal sleep emotion data feature extraction and fusion if not initiated, and performs sleep-emotion state recognition based on the obtained multimodal sleep emotion data feature extraction and fusion results, obtaining the sleep-emotion state recognition result. The sleep collaborative regulation and safety protection module performs sleep emotion collaborative intelligent regulation based on the sleep-emotion state recognition results, and provides sleep emotion responsive safety protection during the sleep emotion collaborative intelligent regulation process.
[0010] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0011] 1. By performing multimodal signal transmission timing interference analysis, and based on the obtained results, it is determined whether to initiate fixed clock calibration alignment and dynamic resynchronization of multimodal signals. If initiated, after the activation, sleep physiological signal timing logic verification is performed. Based on the obtained results, it is determined whether to perform multimodal sleep emotion data feature extraction and fusion. This helps to identify and correct timing deviations and interference problems in multimodal signal transmission from the source, ensuring consistency and physiological logic rationality of each modality signal in the time dimension. This provides high-quality, timing-consistent basic data for subsequent feature extraction and fusion, reducing the risk of misjudgment of sleep-emotion state caused by timing disorder of sleep emotion multimodal signals, and improving the accuracy of state recognition. If not initiated, multimodal... This method involves extracting and fusing features from multimodal sleep emotion data, and then identifying sleep-emotion states based on the obtained feature extraction and fusion results. This helps to eliminate redundant calibration steps and improve the efficiency of data processing and control response in scenarios with good multimodal signal transmission quality and no significant temporal interference, enabling rapid identification and real-time intervention of sleep-emotion states. Based on the sleep-emotion state identification results, collaborative intelligent control of sleep emotions is implemented. During the collaborative intelligent control process, sleep emotion responsive safety protection is provided, which helps to monitor changes in the physiological state of sleep emotions in real time during the control process, while taking into account the timeliness of control and the continuity of sleep. By triggering protective actions in a targeted manner, potential risks caused by improper control parameters are reduced, thereby improving the safety of intelligent control.
[0012] 2. Targeting the standard deviation of data packet timing and the coefficient of variation of data packet timing as jitter indicators for multimodal signal transmission helps to more accurately and comprehensively quantify the timing stability of multimodal signal transmission, providing a more reliable basis for the selection of subsequent calibration strategies. This overcomes the shortcomings of existing technologies that only use a single timing indicator to evaluate transmission jitter, and cannot adapt to signal transmission scenarios with different transmission rates and average arrival times. The system determines whether the multimodal signal transmission jitter indicator exceeds a preset jitter threshold. If so, dynamic resynchronization of the multimodal signal is initiated; otherwise, fixed clock calibration is initiated. Alignment helps to adaptively match the corresponding calibration strategy according to the severity of transmission jitter. When the data timing jitter is small, fixed clock calibration alignment with lower computational cost and shorter time consumption is used to quickly correct the overall clock offset, balancing calibration efficiency and sleep continuity. When the data timing jitter is large, dynamic resynchronization of multi-modal signals with higher accuracy is used to accurately compensate for instantaneous timing interference and ensure the timing consistency of multi-modal signals. This helps to solve the problem that the overall data timing offset caused by the use of fixed calibration strategies in existing technologies cannot compensate for timing interference related to physical events.
[0013] 3. When a single, one-time trigger protection cannot adapt to the physiological tolerance differences of users of different ages and health conditions, it is easy to cause overprotection that interferes with sleep or underprotection that leads to risks. It is necessary to match the intensity of protection with the degree of physiological abnormality. Therefore, the second embodiment of sleep emotion response safety protection is adopted. Obtaining multimodal signal deviation parameters helps to accurately reflect the degree of physiological abnormality of different users, and provides a quantitative and personalized judgment basis for matching the intensity of differentiated protection. This helps to solve the problem of insufficient adaptability caused by the general physiological threshold used in the existing technology to determine the protection action. Based on the multimodal signal deviation parameters, the current safety level is determined, which helps to match the corresponding progressive protection strategy based on the safety level. It can realize layered defense of mild optimization, moderate pause and severe cut-off, which can not only ensure physiological safety, but also minimize the interference of protection actions on sleep quality. This helps to solve the problem of overprotection that interferes with sleep and underprotection caused by the inability of the intensity of protection actions to match the degree of risk in the existing technology. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart of a multimodal sensor signal analysis method for sleep mood state recognition provided in an embodiment of the present invention;
[0016] Figure 2 This is a flowchart summarizing the multimodal sensor signal analysis method for sleep mood state recognition provided in this embodiment of the invention.
[0017] Figure 3 This is a multimodal signal dynamic resynchronization logic diagram of the multimodal sensor signal analysis method for sleep mood state recognition provided in this embodiment of the invention;
[0018] Figure 4 This is a schematic diagram of the structure of a multimodal sensor signal analysis system for sleep mood state recognition provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0020] Before providing a detailed explanation of the embodiments of this application, the application scenarios of these embodiments will be described first.
[0021] Example 1 provides a multimodal sensor signal analysis method for sleep mood state recognition. For example... Figure 1 The flowchart shown is for a multimodal sensor signal analysis method for sleep mood state recognition. The processing flow of this method may include the following steps: multimodal signal transmission timing interference analysis. During the acquisition of multimodal physiological signals for sleep mood states, multimodal signal transmission timing interference analysis is performed to evaluate the timing consistency and the degree of electromagnetic interference during multimodal signal transmission. It is then determined whether to initiate fixed clock calibration alignment to correct the deviation between the sensor's local clock and the system's synchronous clock, and whether to initiate dynamic resynchronization of the multimodal signals to compensate for timing errors introduced by electromagnetic interference and response differences. If initiated, after the activation is complete, a process is performed to verify the rationality of the multimodal signal physiological timing after timing alignment. The system verifies the temporal logic of sleep physiological signals and, based on the verification results, determines whether to perform multimodal sleep emotion data feature extraction and fusion for extracting key information related to sleep emotions. Through multimodal signal transmission temporal interference analysis, it helps control the temporal quality of multimodal physiological signals from the data source. On the one hand, it can promptly identify timing jitter caused by electromagnetic interference during signal transmission, ensuring the temporal consistency and physiological logic rationality of each modality signal, providing reliable basic data without temporal deviation for subsequent sleep emotion state data feature extraction. On the other hand, it can determine whether to initiate calibration, skipping redundant calibration steps when the signal transmission status is good, saving computational resources.
[0022] If not initiated, multimodal sleep emotion data feature extraction and fusion are performed. Based on the obtained multimodal sleep emotion data feature extraction and fusion results, sleep-emotion state recognition is performed to determine the sleep and emotion state categories, and the sleep-emotion state recognition results are obtained. Through multimodal sleep emotion data feature extraction and fusion and sleep-emotion state recognition, it is helpful to integrate physiological information related to sleep state from multimodal signals such as EEG, PPG, and GSR (such as brainwave rhythm in EEG signals and heart rate variability derived from PPG signals), reducing the limitations of single-modal signals being easily affected by individual physiological differences and environmental interference. Through the fusion of multimodal features, a multidimensional characterization of sleep state is achieved. The generated sleep-emotion state recognition results can accurately reflect the current sleep state, providing an accurate decision-making basis for subsequent personalized regulation.
[0023] This system utilizes sleep-emotion state recognition results to implement personalized regulation tailored to different sleep-emotion states. During this process, a sleep-emotion responsive safety protection mechanism is implemented to safeguard physiological homeostasis during sleep and reduce excessive stress responses triggered by multisensory stimulation. Through this system, customized multisensory interventions based on sleep-emotion state recognition results can be achieved. This includes adjusting lighting, sound, temperature, and humidity to improve sleep quality, while simultaneously monitoring physiological homeostasis changes in real time to avoid excessive stress responses caused by multisensory stimulation, thus balancing the intervention effect with the natural continuity of sleep.
[0024] It should be understood that the multimodal sensor signal analysis method for sleep emotion state recognition provided in this application relies on the establishment and maintenance of core data infrastructure for its effective implementation and continuous optimization. This knowledge base provides decision-making basis and learning foundation for the entire analysis process. Its knowledge system integrates two key sources: one is standardized parameters derived from prior knowledge in the fields of sleep medicine, neuroscience, and physiology, such as preset jitter threshold, preset time deviation threshold, and preset amplitude threshold; the other is empirical data sets derived from large-scale experiments and real-world monitoring, including typical multimodal signals with accurately labeled emotion states, feature vector templates of different individuals in various emotion-sleep states, and records of the effectiveness of closed-loop regulation strategies. These historical data provide a solid foundation for the calibration of theoretical parameters and personalized training of models.
[0025] At the engineering implementation level, the knowledge base adopts a layered and integrated architecture for organization and management. Structured calibration parameters, physiological logic rules, and feature template metadata are stored in a relational database to ensure the rigor of the definitions and the efficiency of related queries. Meanwhile, massive amounts of raw signal fragments, continuous feature sequences, and regulatory logs organized by time series are stored in a time series database to support efficient retrieval and retrospective analysis of physiological patterns over long periods. In addition, the knowledge base is designed to have incremental learning capabilities. The system can automatically update and optimize feature templates and regulatory parameters incrementally based on data from new users and closed-loop feedback results, thereby achieving continuous adaptation to individual differences and iterative improvement of overall model performance.
[0026] As described above, by employing multimodal signal transmission temporal interference analysis, multimodal sleep emotion data feature extraction and fusion, sleep-emotion state recognition, sleep-emotion collaborative intelligent regulation, and sleep emotion responsive safety protection, a closed-loop intelligent sleep regulation solution covering multimodal physiological signal quality control, from precise sleep state perception to personalized sleep regulation, and finally to sleep safety protection, can be constructed. This solution addresses the problems of large signal timing deviations, low state recognition accuracy, insufficient regulation adaptability, and lack of safety protection in traditional sleep regulation solutions, thereby improving the accuracy and effectiveness of sleep regulation. Specifically, multimodal signal transmission temporal interference analysis is responsible for screening and calibrating the timing quality of the collected raw multimodal physiological signals, directly outputting timing-qualified multimodal signals to multimodal sleep emotion data feature extraction and... The fusion of multimodal sleep emotion data features transforms fragmented single-modal physiological signals into a multimodal fused emotion state feature vector that integrates multidimensional sleep-related information. This multimodal fused emotion state feature vector is the core input for sleep-emotion state recognition. The completeness and accuracy of the features directly determine the reliability of the state recognition results. Sleep-emotion state recognition completes the sleep state determination based on the multimodal fused emotion state feature vector. Its output sleep-emotion state result is the sole execution basis of the sleep emotion collaborative intelligent regulation module. The accuracy of state recognition directly affects the adaptability of the regulation strategy. The sleep emotion collaborative intelligent regulation executes a multi-sensory regulation strategy. The sleep emotion responsive safety protection monitors the physiological impact of the regulation actions to avoid safety risks caused by over-regulation or inappropriate regulation.
[0027] like Figure 2 The diagram shown is a flowchart summarizing the multimodal sensor signal analysis method for sleep mood state recognition provided in this embodiment of the invention. Figure 2It can be seen that: Multimodal signal transmission timing interference analysis is performed, and multimodal signal transmission jitter index is obtained. It is determined whether the multimodal signal transmission jitter index exceeds a preset jitter threshold. If it does, dynamic resynchronization of the multimodal signal is initiated; otherwise, fixed clock calibration and alignment are initiated. After fixed clock calibration and alignment, timestamp deviation is obtained, and it is determined whether the timestamp deviation is less than a preset time deviation threshold. If it is, multimodal sleep emotion data feature extraction and fusion are performed; otherwise, dynamic resynchronization of the multimodal signal is performed. After dynamic resynchronization of the multimodal signal, sleep physiological signal timing logic verification is performed, and sleep... The system evaluates sleep physiological signal indicators and determines whether all indicators meet the preset physiological time sequence logic requirements. If so, it performs multimodal sleep emotion data feature extraction and fusion; otherwise, it sends a time sequence resynchronization failure prompt. After the multimodal sleep emotion data feature extraction and fusion is completed, sleep-emotion state recognition is performed. After the sleep-emotion state recognition is completed, sleep emotion collaborative intelligent regulation is performed, and sleep emotion physiological signals are acquired. The system determines whether the sleep emotion physiological signals meet the initial regulation target conditions. If so, it gradually reduces the feedback intensity to the point of shutdown and terminates the current regulation; otherwise, it sends a poor regulation effect prompt.
[0028] Preferably, the specific process of multimodal signal transmission timing interference analysis is as follows: The interval sequence of data packets (such as EEG data packets, HRV data packets, GSR data packets, etc.) arriving at the sleep emotional state data processing center from each modal signal sensing channel is monitored to obtain a multimodal signal transmission jitter index reflecting the timing stability of data packet transmission. The multimodal signal sensing channel is a dedicated data transmission path for transmitting sleep emotional physiological signals. One end connects to the corresponding type of multimodal signal sensor, and the other end connects to the sleep emotional state data processing center. It is responsible for transmitting the data packets collected by the sensors in a directional and continuous manner, such as through the EEG sensing channel, HRV sensing channel, etc. The system includes multiple sensing channels such as the GSR sensing channel. It determines whether the multimodal signal transmission jitter index exceeds a preset jitter threshold. If so, it determines that the sensing channel is experiencing interference (such as weak electromagnetic interference or signal attenuation along the transmission path), and initiates dynamic resynchronization of the multimodal signal to achieve more accurate, event-driven multimodal signal timing alignment, providing a consistent foundation for subsequent feature extraction and multimodal fusion. Conversely, it initiates fixed clock calibration alignment to ensure that the data from each sensing channel are aligned based on a unified clock reference when the multimodal signal transmission jitter is small. The preset jitter threshold is represented by the average value of the multimodal signal transmission jitter index over a historical time period.
[0029] Specifically, the formula for the multimodal signal transmission jitter index is as follows:
[0030]
[0031] Wherein, TJ is the multimodal signal transmission jitter index, σ is the data packet timing standard deviation, CV is the data packet timing coefficient of variation, σ0 represents the preset timing threshold, which is represented by the average of the data packet timing standard deviation over a historical period. The data packet timing standard deviation is represented by the arithmetic standard deviation of the difference between the actual arrival time of the data packet at the sleep mood state data processing center through the multimodal signal sensing channel and the preset arrival time of the data packet, and is used to reflect the absolute dispersion of data packet timing fluctuations. The data packet timing coefficient of variation is represented by the ratio of the data packet timing standard deviation to the average arrival time of the data packet in the multimodal signal sensing channel, and is used to reflect the relative dispersion of data packet timing fluctuations. a represents the influence value of the data packet timing standard deviation, which is used to reflect the degree of influence of the data packet timing standard deviation on the multimodal signal transmission jitter index, and b represents the influence value of the data packet timing coefficient of variation, which is used to reflect the degree of influence of the data packet timing coefficient of variation on the multimodal signal transmission jitter index.
[0032] It should be understood that in the analysis of timing interference in multimodal signal transmission, a set of influencing parameters based on the quantitative evaluation and optimization of historical transmission link performance are involved, including the influence value of the standard deviation of data packet timing and the influence value of the coefficient of variation of data packet timing. The generation and calibration of these two influencing parameters follow a data-driven empirical process. First, it is necessary to accumulate a multimodal transmission timing dataset covering typical sleep scenarios, device states, and interference types. Based on this dataset, the algorithm performs two core analyses in parallel: one is to calculate the standard deviation of the data packet arrival interval sequence in stable and interfered states for each sensing channel to quantify its absolute dispersion; the other is to calculate the coefficient of variation of the sequence (the ratio of standard deviation to mean) to assess its relative volatility. Through statistical induction of a large number of samples, the contribution patterns of the standard deviation and coefficient of variation to the final perceived transmission jitter can be determined, and the initial values of the influencing parameters are assigned by the statistical characteristics of this contribution pattern (such as median or mean).
[0033] To ensure the robustness of the influencing parameters, the initial values of the influencing parameters will be verified and adjusted in various scenarios. For example, in a bedroom environment where the wireless signal strength varies, if the impact of the data packet timing variation coefficient is more significant, the weight of the data packet timing variation coefficient may be increased accordingly. In cases near fixed interference sources, the role of the data packet timing standard deviation may be enhanced. All verified and effective parameter combinations and their applicable scenarios will be stored in the system configuration library to form a parameter knowledge system that can be dynamically matched or fine-tuned based on real-time environmental assessment results. This will ensure that the transmission jitter index can reliably reflect the timing stability of the channel under different actual conditions.
[0034] As described above, multimodal signal transmission timing interference analysis can systematically identify timing anomalies in multimodal sleep emotional physiological signals during acquisition and transmission, caused by factors such as the offset between the sensor's local clock and the system's synchronous clock, environmental electromagnetic radiation interference, transmission bandwidth fluctuations, and link delay jitter. It can accurately quantify the timing dispersion, consistency level, and interference impact level of each modality's signal transmission, providing a quantitative and traceable basis for selecting subsequent timing calibration strategies. This establishes a standardized timing quality benchmark for the subsequent processing of multimodal sleep emotional physiological signals, helping to control the effectiveness and reliability of multimodal physiological signals from the data source and ensuring that the input data for subsequent feature extraction and state recognition processes has strict timing consistency and physiological logical rationality.
[0035] Preferably, the specific process for fixed clock calibration and alignment is as follows: A1, extract the system synchronization clock signal from the sleep mood state data processing center as the unified time reference for the multimodal signal sensing channels; A2, read the original acquisition timestamp of the data packet corresponding to each multimodal signal sensing channel (generated by the local clock of each sensor) as the local clock of the multimodal signal sensor; A3, calculate the clock deviation between the local clock of the multimodal signal sensor and the system synchronization clock; the clock deviation is represented by the result of the difference operation between the system synchronization clock of the sleep mood state data processing center and the local clock of the multimodal signal sensor. A positive clock deviation indicates that the local clock of the multimodal signal sensor lags behind the system synchronization clock, a negative clock deviation indicates that the local clock of the multimodal signal sensor leads the system synchronization clock, and a clock deviation of 0 indicates that the local clock of the multimodal signal sensor is ahead of the system synchronization clock. The timing is fully synchronized; A4, the timestamps of all data packets in the multimodal signal sensor channel are corrected by clock deviation in one go to obtain the corrected timestamp; the one-time clock correction means superimposing the clock deviation on the original timestamp of each data packet in the multimodal signal sensor channel; after the one-time clock correction is completed, the multimodal signal sensor channel data is interpolated to a unified time axis based on the one-time corrected timestamp and an interpolation algorithm (such as linear interpolation, cubic spline interpolation, etc.) to ensure that no data is lost; A5, after the fixed clock calibration and alignment is completed, the timestamp deviation of the aligned multimodal signal sensor channel data is verified. If the timestamp deviation is less than the preset timestamp deviation threshold, multimodal sleep emotion data feature extraction and fusion are performed; otherwise, multimodal signal dynamic resynchronization is performed. The preset timestamp deviation threshold is represented by the average value of timestamp deviations over historical time periods.
[0036] As described above, aligning with a fixed clock helps to correct the systematic offset between the local clock of the multimodal sleep emotion physiological signal sensor and the system synchronization clock from the root, unifies the timing measurement benchmark of each sensing channel, reduces long-term and regular timing misalignment caused by inconsistent clock benchmarks in the acquisition and transmission links of multimodal signals, reduces the risk of feature mismatch and physiological logic contradiction caused by timing benchmark chaos in the subsequent feature extraction and fusion stages of multimodal signals, and reduces the computational resource consumption caused by redundant processing of timing abnormal data in the system. This provides reliable and qualified timing data for the entire link of sleep-emotion state recognition and sleep multi-sensory collaborative regulation.
[0037] like Figure 3 The diagram shown is a multimodal signal dynamic resynchronization logic diagram of the multimodal sensor signal analysis method for sleep mood state recognition provided in this embodiment of the invention. Figure 3 It can be seen that: multimodal signal dynamic resynchronization is performed, and the number of effective peak occurrences is obtained. It is determined whether the number of effective peak occurrences is greater than the preset number of action judgments. If so, multimodal signal timing calibration is performed; otherwise, sleep physiological signal timing logic verification is performed. Multimodal signal timing calibration is performed and event-physiological timing deviation is obtained. It is determined whether the event-physiological timing deviation is greater than the preset positive deviation judgment threshold. If so, positive shift correction is performed; otherwise, it is determined whether the event-physiological timing deviation is less than the preset negative deviation judgment threshold. If so, reverse shift correction is performed; otherwise, sleep physiological signal timing logic verification is performed.
[0038] Preferably, the specific process of dynamic resynchronization of multimodal signals is as follows: Since high-sampling-rate accelerometers collect physical and mechanical signals of limb movements during sleep, they do not rely on the potential conduction or acquisition mechanism of electrophysiological signals. They are minimally affected by environmental electromagnetic interference (such as electromagnetic radiation from medical equipment or electromagnetic interference from household appliances). Furthermore, their high sampling rate characteristic can accurately capture instantaneous mechanical changes in limb movements. The signal peak directly corresponds to the occurrence time of sleep physical events such as turning over and limb movements, exhibiting strong temporal correlation and no physiological response delay. Therefore, the high-sampling-rate accelerometer signal is used as a temporal reference for sleep physical events unaffected by electromagnetic interference. Feature points of sleep physical events monitored by the high-sampling-rate accelerometer (such as the acceleration peak point corresponding to turning over) are extracted, and the timestamps of these feature points are recorded. The specific extraction process involves obtaining the number of effective peak occurrences used to assess the frequency of spontaneous limb movement sleep physical events such as turning over and limb movements during sleep. The number of effective peak occurrences represents the signal... Within the dynamic synchronization period, the number of times the acceleration signal amplitude exceeds a preset amplitude threshold is counted. The acceleration signal amplitude is represented by the vector amplitude of the original output components of the X, Y, and Z axes of the triaxial accelerometer. The preset amplitude threshold is represented by the average value of the acceleration signal amplitude over a historical time period. The system determines whether the number of valid peak occurrences exceeds the preset number of action judgments. If so, the sleep physical event feature point corresponding to the acceleration signal amplitude exceeding the preset amplitude threshold is marked as the sleep physical event baseline feature point. The sleep physical event timestamp and the corresponding preset physiological signal response duration of the baseline feature point are obtained. Multimodal signal timing calibration is performed based on the sleep physical event timestamp and the preset physiological signal response duration. Otherwise, sleep physiological signal timing logic verification is performed. The preset number of action judgments is represented by the average value of the number of valid peak occurrences over a historical time period. The preset physiological signal response duration is pre-set by a preset person to cover the complete physiological response process triggered by the sleep physical event.
[0039] Specifically, multimodal signal timing calibration is performed based on sleep physical event timestamps and preset physiological signal response durations. The specific process is as follows: B1, extract the timestamps corresponding to the physiological signal associated feature points in the multimodal signal sensor channel data that are common to the sleep physical event timestamps. This represents a continuous signal segment with a total duration equal to the preset physiological signal response duration, centered on the sleep physical event timestamps. Within the continuous signal segment, physiological signal associated feature points (such as α / β wave peaks in EEG, R wave peaks in PPG, conductance abrupt changes in GSR, and inspiratory peak points in respiratory signals) are extracted using multimodal signal feature recognition algorithms (such as peak detection algorithms and abrupt change point recognition algorithms). For example, for periodic rhythmic peaks like α / β wave peaks in EEG signals, a peak detection algorithm is used. The application of the peak detection algorithm is to traverse the continuous EEG signal segment and select local maxima points where the EEG amplitude is greater than the preset EEG amplitude threshold and meets the periodic range of the α wave (8-13Hz) and β wave (14-30Hz) characteristic frequency bands. These are used as EEG signal peaks. The alpha / beta peaks of the signal are defined as follows: EEG amplitude is represented by the absolute difference between the EEG signal amplitude at a point monitored by an EEG sensor and the baseline amplitude of the corresponding continuous EEG signal segment; the preset EEG amplitude threshold is represented by the average EEG amplitude over a historical time period; and the baseline amplitude of the EEG signal segment is represented by the average EEG amplitude of the continuous EEG signal segment. A system synchronization timestamp is marked for each physiological signal-related feature point, serving as the temporal identifier for the sleep data point to be calibrated. The sleep data point to be calibrated represents a feature point data unit extracted from multimodal signal sensing channels such as EEG, PPG / HRV, GSR, and respiration within the physiological response cycle triggered by a sleep physical event, and directly physiologically related to that sleep physical event. This data unit has an initial system synchronization timestamp, but this timestamp may be affected by weak electromagnetic interference leading to signal transmission delays, sensor local clock deviations, etc., resulting in signal timing errors. This data unit is the core object for subsequent multimodal signal timing alignment calibration based on the sleep physical event timing benchmark.
[0040] B2, calculate the event-physiological time series deviation between the sleep data point to be calibrated and the corresponding sleep physical event feature point; the event-physiological time series deviation is the difference between the timestamp of the sleep physical event and the timestamp of the corresponding sleep data point to be calibrated, used to characterize the time series offset of the physiological signal associated feature point relative to the sleep physical event feature point; the positive or negative selection interpolation algorithm based on the event-physiological time series deviation is as follows: determine whether the event-physiological time series deviation is greater than the preset positive deviation judgment threshold. If so, it means that the sleep data point to be calibrated of the physiological signal lags behind the sleep physical event feature point of the physical event. Therefore, based on the forward linear interpolation method, the sleep physical event time series deviation is used as the basis for the calculation. Using the time stamp as a baseline, a preset number of data samples before and after the sleep data point to be calibrated are forward-shifted for correction. After the forward shift correction is completed, sleep physiological signal timing logic verification is performed; otherwise, event-physiological timing inverse discrimination is performed. The preset positive deviation judgment threshold represents a time sequence deviation threshold greater than 0 used to determine whether the sleep emotional physiological signal time stamp lags behind the sleep physical event time stamp. It is represented by the average of event-physiological timing deviations greater than 0 over a historical period. The preset positive deviation judgment threshold is set in advance by preset personnel. Forward shift correction means shifting the timestamp of the sleep data point to be calibrated in the positive direction of the time axis (i.e., the time of the sleep physical event). The shift is performed using the absolute value of the event-physiological timeline deviation as the shift amount to ensure the temporal matching between the sleep data points to be calibrated and the corresponding sleep physical event feature points, while also ensuring the temporal continuity of the corrected data samples. The event-physiological timeline reverse discrimination determines whether the event-physiological timeline deviation is less than a preset negative deviation threshold. If so, it indicates that the sleep data points to be calibrated for the physiological signal precede the sleep physical event feature points. Therefore, based on backward linear interpolation, using the sleep physical event timestamp as a reference, the timestamps of a preset number of data samples before and after the sleep data points to be calibrated are reverse-shifted and corrected. The reverse shift correction ends. Then, the timing logic verification of sleep physiological signals is performed; otherwise, the timing logic verification of sleep physiological signals is performed. The preset negative deviation judgment threshold is a timing deviation threshold less than 0 used to determine whether the timestamp of sleep emotional physiological signals precedes the timestamp of sleep physical events. It is represented by the average value of event-physiological timing deviations less than 0 over a historical period. The reverse translation correction means that the timestamp of the sleep data point to be calibrated is translated in the negative direction of the time axis (i.e., the direction of sleep physical event timestamps) by using the absolute value of the event-physiological timing deviation as the translation amount. This is used to compensate for the advance deviation of physiological signals relative to physical events and maintain the timing integrity of the data sequence.
[0041] Specifically, the temporal logic verification of sleep physiological signals follows this process: Step 1: Screening a combination of sleep physiological signal indicators reflecting emotional-stress states. This combination must cover all modalities, including but not limited to: GSR rise rate, HRV characteristic band power ratio change rate, and EEG characteristic band power ratio change rate. The GSR rise rate is represented by the change in GSR signal amplitude monitored by a skin conductance sensor during the signal dynamic synchronization period, used to quantify the dynamic response intensity of skin conductance. The HRV characteristic band power ratio change rate is represented by the change in the ratio of low-frequency band power to high-frequency band power of HRV monitored by an electrocardiogram sensor during the signal dynamic synchronization period, used to quantify the dynamic balance of autonomic nervous activity. The EEG characteristic band power ratio change rate is represented by the change in the ratio of specific characteristic band power (such as alpha wave, beta wave) to other characteristic band power (such as theta wave) monitored by an electroencephalogram sensor during the signal dynamic synchronization period, used to quantify the dynamic trend of cerebral cortex activity.
[0042] Step two involves performing a temporal correlation verification process. Specifically, this involves obtaining the peak times of the sleep physiological signal indicators in the combination of sleep physiological signal indicators and the time interval between these peak times. Each sleep physiological signal indicator is then evaluated to determine if its temporal relationship satisfies the preset physiological temporal logic. The temporal relationship represents the quantifiable sequence and time interval between different physiological responses when sleep emotions trigger a stress event. For example, when anxiety triggers a stress response, the peak times of the rise in GSR (Geophysiological Skin Response) conductance and the peak times of the LF / HF (Low Frequency / High Frequency) ratio of HRV (Heart Rate Variability) are considered. If all sleep physiological signal indicators meet the preset physiological temporal logic requirements, the dynamic resynchronization is deemed effective, and the corresponding multimodal sleep emotion data is transferred. Based on the time-series qualified multimodal sleep emotion data, multimodal sleep emotion data feature extraction and fusion are performed. If any sleep physiological signal indicator does not meet the preset physiological time sequence logic requirements, a time sequence resynchronization failure prompt is sent. The preset physiological time sequence logic is set in advance by preset personnel. For example, in a stress or anxiety state, the peak time of the conductance rise of GSR should lead the peak time of HRV LF / HF ratio, and the lead time should reach the preset GSR lead time. In a relaxed state, the rise time of EEG alpha power ratio should be synchronized with the fall time of respiratory rate, and the EEG time deviation should not be greater than the preset EEG time deviation threshold. The preset GSR lead time and the preset EEG time deviation threshold are both set in advance by preset personnel.
[0043] As described above, dynamic resynchronization of multimodal signals and temporal logic verification of sleep physiological signals help to accurately compensate for the instantaneous timing deviations introduced by factors such as electromagnetic interference, sensor response differences, and instantaneous link delays during the transmission of multimodal sleep emotional physiological signals. At the same time, the inherent temporal correlation logic of sleep physiological signals verifies the physiological rationality of the timing alignment of each modality signal, reducing the probability of problems such as multimodal signal feature misalignment and physiological correlation logic contradictions caused by timing inaccuracies. It also reduces the risk of invalid or erroneous features in the subsequent multimodal sleep emotional data feature extraction stage, improves the reliability of the timing identification of sleep data points to be calibrated, and realizes closed-loop control from timing error compensation to physiological logic verification, providing high-quality data support for intelligent sleep regulation solutions that combines timing accuracy and physiological rationality.
[0044] Preferably, feature extraction and fusion of multimodal sleep emotion data based on time-series qualified multimodal sleep emotion data are performed. The specific process is as follows: C1, each time-series qualified sleep emotion modality signal is input into an independent lightweight encoder network to generate a multimodal primary feature embedding sequence. The specific process is as follows: the time-series qualified sleep emotion modality signal includes time-series consistent EEG signals, PPG signals, GSR signals, respiratory signals, and body movement signals. Among them, the respiratory signal is monitored by a respiratory impedance sensor, and the body movement signal is monitored by a triaxial accelerometer. The EEG signal is input into a one-dimensional convolutional neural network to extract local features in the time and frequency domains, and the output is an EEG feature embedding sequence. This sequence reflects the sensory perception of the cerebral cortex. The study investigates neural activity states related to wakefulness, relaxation, and drowsiness. PPG signals are input into a recurrent neural network to extract their rhythmic and morphological features, outputting an HRV feature embedding sequence that reflects the balance between the sympathetic and parasympathetic nervous systems. GSR signals are input into a one-dimensional convolutional neural network to extract their phase and tension components, outputting a GSR feature embedding sequence that reflects the excitation level of the sympathetic nervous system. Respiratory and body movement signals are concatenated and input into a fully connected network to extract their rhythmic and energy features, outputting a respiratory-body movement feature embedding sequence that reflects the stability of respiratory rhythm and emotional synchronization. Based on preset concatenation rules, the multimodal primary feature embedding sequences are concatenated to form a hybrid multimodal sequence.
[0045] C2 inputs the mixed multimodal sequence into the multimodal Transformer encoder, performs cross-modal feature fusion, and outputs a unified multimodal fused emotional state feature vector; C3 obtains sleep emotion intensity parameters for quantifying the intensity of sleep emotion and stress-related states based on the multimodal fused emotional state feature vector and the sigmoid function, including the physiological arousal intensity quantification value for quantifying the intensity of anxiety state, the physiological homeostasis intensity quantification value for quantifying the intensity of relaxation state, and the sleep emotion intensity quantification value for quantifying the intensity of drowsiness state; sleep-emotion state recognition is performed based on the multimodal fused emotional state feature vector and the sleep emotion intensity parameters.
[0046] Specifically, the sleep-emotion state recognition process is as follows: Multimodal fused emotion state feature vectors and sleep emotion intensity parameters are input into a pre-defined classification model (such as support vector machine, random forest, deep neural network classifier, etc.). The output is a sleep-emotion state category (e.g., high-arousal sleep-emotion state, low-arousal sleep-emotion state, etc.) and the corresponding sleep-emotion state confidence score (e.g., confidence score for high-arousal sleep-emotion state, confidence score for low-arousal sleep-emotion state, etc.). The specific training process of the classification model is as follows: Multimodal sleep emotion sample data labeled with sleep-emotion state category tags are collected (e.g., alpha / beta wave power ratio features collected and extracted by EEG sensors, and conductance data collected by skin conductance sensors). The data (such as rise rate features and HRV LF / HF ratio features acquired by photoplethysmography sensors) are divided into training and testing sets according to a preset ratio. Steps C1-C3 are performed on the sample data in the training set to obtain the corresponding multimodal fusion emotional state feature vector and sleep emotional intensity parameters. These two are then concatenated as the input features of the classification model. The labeled sleep-emotional state category is used as the target label. The classification model is iteratively trained using a preset loss function (such as cross-entropy loss function). The model parameters are continuously adjusted to minimize the prediction error. The recognition accuracy and generalization ability of the model are verified through the test set. When the model performance meets the preset evaluation threshold, training is stopped and the final model parameters are saved to form a classification model that can be directly used for online recognition.
[0047] As described above, multimodal sleep emotion data feature extraction and fusion, along with sleep-emotion state recognition, helps to transform fragmented sleep-related information scattered across different modal physiological signals such as EEG, GSR, HRV, and respiration into a structured fusion feature system with cross-modal complementarity. This breaks through the blind spots in state characterization caused by individual physiological specificity and environmental interference in single-modal signals. It can uncover the intrinsic relationship between sleep state and physiological response from the synergistic relationship of multi-dimensional physiological feedback, capture subtle physiological fluctuations and state transition features that are easily overlooked by a single modality during sleep, and improve the accuracy of sleep-emotion state recognition.
[0048] The sleep-emotion state collaborative intelligent regulation is based on sleep-emotion state categories and sleep-emotion state confidence scores. The specific process is as follows: Pairwise comparisons are performed on the sleep-emotion state confidence scores corresponding to each sleep-emotion state category. The sleep-emotion state category with the highest confidence score is marked as the dominant sleep-emotion state category. If two or more sleep-emotion state confidence scores are equal and represent the current maximum value, the following rules apply: Retrieve the respiratory-motor feature embedding sequences corresponding to the sleep-emotion state categories with equal confidence scores. (The respiratory-motor feature embedding sequences reflect the stability of respiratory rhythm and the coordination of limb movements during sleep. These sequences have the highest correlation with sleep states, are least affected by environmental interference, and can accurately distinguish the physiological differences between different sleep states.) To reduce the judgment error of weakly associated features, the matching degree of emotion-sleep state categories is obtained. The sleep-emotion state category with the higher matching degree is marked as the dominant sleep-emotion state category. The matching degree of emotion-sleep state categories is represented by the cosine similarity between the respiratory and body movement feature embedding sequence and the preset respiratory and body movement sequence, where the preset respiratory and body movement sequence is set in advance by preset personnel. Based on the dominant sleep-emotion state category, a preset multi-sensory feedback strategy combination is obtained (feedback types include sound feedback, light feedback, temperature feedback, and tactile feedback. The strategy combination is preset based on the neurophysiological regulation mechanism. For example, if the dominant state is anxiety, a soothing combination of low-frequency rhythmic sound + warm light + gentle skin heating + low-frequency tactile vibration is matched. The core function is to activate the vagus nerve and inhibit the excitation of the sympathetic nerve.
[0049] Specifically, the process involves regulating and monitoring the physiological responses to sleep-emotional states. The process is as follows: It determines whether the physiological signals of sleep-emotional states (such as EEG, PPG / HRV, GSR, and respiratory signals) meet the initial regulation target conditions (e.g., the quantified value of physiological arousal intensity is less than the preset physiological arousal threshold, the quantified value of physiological homeostasis intensity is greater than the preset physiological state stability threshold, and the quantified value of sleep-emotional intensity is greater than the preset sleep intensity threshold). If so, the feedback intensity is gradually reduced to 0 (e.g., by reducing the feedback intensity of the sleep-aid ambient light to the sleep-emotional physiological state through the visual regulation unit until the sleep-aid ambient light is turned off, or by reducing the volume of the sleep-aid white noise to the sleep-emotional physiological state until the sleep-aid white noise is turned off through the auditory regulation unit), terminating the current regulation to avoid continuous stimulation affecting sleep quality. The preset physiological state stability threshold is represented by the average of the quantified values of physiological homeostasis intensity over a historical period, the preset physiological arousal threshold by the average of the quantified values of physiological arousal intensity over a historical period, and the preset sleep intensity threshold by the average of the quantified values of sleep-emotional intensity over a historical period. Conversely, if the signals do not meet the target conditions, a warning of poor regulation effect is sent, and the current sleep-emotional coordinated intelligent regulation parameters (such as the current brightness and color temperature of the sleep-aid ambient light in the visual regulation unit) are updated. Temperature and light adjustment rate, current volume and frequency range of the sleep-aiding white noise from the auditory control unit, and physiological response data (such as the quantized value of the β / α ratio of EEG, the quantized value of the LF / HF ratio of HRV, and the steady-state conductance value of GSR, etc.) are fed back to the sleep emotional state data processing center. The initial control target conditions are preset by the pre-set personnel as the benchmark for judging the control effect. The benchmark is determined based on the sleep-emotional state recognition results, and specific quantitative feature values of the corresponding matching sleep emotional physiological signals (such as the GSR steady-state conductance value corresponding to anxiety-type hyperarousal sleep-emotional state, the β wave / The alpha wave ratio, corresponding to the LF / HF ratio of HRV in hyperactive high-arousal sleep-emotional states, and the alpha / theta wave ratio of EEG in sleep-emotional low-arousal sleep-emotional states, are used for determination (e.g., if the dominant state is anxiety, the initial control target conditions are that the GSR steady-state conductance is less than the preset GSR control target threshold, and the EEG β / α ratio is less than the preset EEG control target threshold, etc.). The preset GSR control target threshold is represented by the average value of GSR steady-state conductance over a historical period, and the preset EEG control target threshold is represented by the average value of the EEG β / α ratio over a historical period.
[0050] As described above, by using intelligent regulation of sleep and emotions, a multi-sensory sleep guidance system can be constructed. This system overcomes the limitations of single-sensory regulation, which has a limited dimension and is easily affected by individual sensory sensitivity. It can match differentiated regulation combinations and intensity thresholds for different levels of sleep initiation, maintenance, and transition states, reduce the interference of uniform regulation on the natural sleep rhythm, reduce the probability of micro-arousals caused by regulatory discomfort during sleep, prolong the duration of stable sleep, and improve the fit of sleep regulation.
[0051] Preferably, the sleep emotion-responsive safety protection process is as follows: Monitoring sleep emotion physiological multimodal signals (such as EEG, HRV, GSR, skin temperature, etc.), and obtaining sleep emotion physiological state monitoring parameters such as heart rate, HRV, GSR change rate, skin temperature, and the β-wave to α-wave ratio of EEG based on feature fusion algorithms; comparing each sleep emotion physiological state monitoring parameter with preset safety thresholds (such as preset heart rate threshold, preset HRV threshold, preset GSR change rate threshold, preset skin temperature); if any sleep emotion physiological state monitoring parameter triggers protection logic (e.g., if HRV exceeds the preset HRV range, it triggers the suspension of photo-acoustic modulation; if skin temperature exceeds the preset skin temperature range, it triggers the cessation of temperature control heating, etc.), and immediately... This involves executing the corresponding control action and sending a notification that the safety protection has been activated. Simultaneously, the control execution counter is incremented to obtain the number of control executions. Based on this number, a sleep-emotion-responsive safety protection mechanism is established. The preset HRV range and preset skin temperature range are both pre-set by designated personnel and include both endpoints. The increment operation represents adding 1 to the control execution counter. The sleep-emotion-responsive safety protection mechanism determines whether the number of control executions exceeds a preset control execution threshold. If so, the sleep-emotion-responsive safety protection method is triggered; otherwise, sleep-emotion-responsive physiological multimodal signals are continuously monitored. The preset control execution threshold is also pre-set by designated personnel.
[0052] As described above, sleep emotion-responsive safety protection helps to capture abnormal fluctuations in multimodal physiological signals such as EEG, GSR, and HRV during the intelligent regulation of sleep emotion in real time, promptly identify potential disturbances of regulatory stimuli to physiological homeostasis, reduce excessive physiological stress caused by excessive intensity of multisensory regulatory stimuli and deviation in regulatory strategy adaptation, improve the controllability of intelligent sleep regulation, and achieve a dynamic balance of sleep regulation intervention effects.
[0053] Example 2 provides an alternative to sleep-emotion-responsive safety protection. When single-trigger protection cannot adapt to the physiological tolerance differences of users of different ages and health conditions, it is prone to overprotection interfering with sleep or underprotection causing risks. Therefore, it is necessary to match differentiated protection strength according to the degree of physiological abnormality risk. The specific process is as follows: Obtain multimodal signal deviation parameters to quantify the degree of abnormal risk of sleep-emotion physiological signals; the multimodal signal deviation parameters include HRV deviation amplitude and GSR rate of change deviation amplitude. The time-domain characteristic values of HRV monitored by the ECG sensor are... The absolute difference between the standard deviation of adjacent normal RR intervals (e.g., the standard deviation of the RR interval) and the preset HRV time-domain feature reference value is calculated, and then the ratio of the absolute difference is calculated to obtain the HRV deviation magnitude. The preset HRV time-domain feature reference value is represented by the average value of HRV time-domain feature values over a historical time period. The absolute difference between the GSR change rate monitored by the skin conductance response sensor and the preset GSR change rate threshold is calculated, and then the ratio of the absolute difference is calculated to obtain the GSR change rate deviation magnitude. The preset GSR change rate threshold is represented by the average value of GSR change rate over a historical time period.
[0054] The current safety level is determined based on the deviation parameters of multimodal signals. The specific process is as follows: First, it is determined whether the deviation parameter of any modal signal is less than a preset first-level deviation threshold. If so, an output limiting control strategy is executed (e.g., reducing the intensity of acoustic / optical feedback to the lower limit of the preset safety range and sending a prompt indicating slight physiological fluctuations and adjusted output). The preset first-level deviation threshold is represented by the average value of multimodal signal deviation parameters over a historical period, reflecting the normal fluctuation range of physiological signals during sleep. Physiological fluctuations within this range are normal physiological changes during sleep. Conversely, if the deviation parameter is less than the preset first-level deviation threshold and less than the preset second-level deviation threshold, a temporary pause in output is executed, and a prompt indicating abnormal physiological parameters and paused control is sent (e.g., pausing acoustic / optical / tactile feedback, shutting down the temperature control module and starting heat dissipation, and sending a prompt indicating abnormal physiological parameters and paused control). The preset second-level deviation threshold is set in advance by designated personnel. Otherwise, a system-wide safety disconnection control strategy is executed (e.g., cutting off the energy output of all feedback modules, stopping all control actions, saving current physiological data and logs, and activating a buzzer alarm).
[0055] As described above, the sleep emotion-responsive safety protection provided in Example 2 effectively compensates for the shortcomings of single-trigger protection, which cannot adapt to the physiological tolerance differences of users of different ages and health conditions. It solves the problems of overprotection interfering with sleep rhythm or underprotection causing health risks, and achieves precise matching between risk level and protection intensity. It achieves a dynamic balance between safety protection and sleep continuity during sleep regulation. It can provide adaptive protection for different physiological tolerance characteristics and take corresponding protective actions at different risk levels, effectively improving the safety and dynamic adaptability of intelligent sleep regulation.
[0056] like Figure 4 The diagram shown is a structural schematic of a multimodal sensor signal analysis system for sleep emotional state recognition provided in an embodiment of the present invention. The system includes: a multimodal signal timing calibration and verification module, a multimodal feature fusion and state recognition module, and a sleep collaborative regulation and security protection module.
[0057] The multimodal signal timing calibration and verification module is used to analyze the timing interference of multimodal signal transmission and determine whether to start fixed clock calibration alignment and dynamic resynchronization of multimodal signals. If started, after the start is completed, sleep physiological signal timing logic verification is performed, and based on the obtained sleep physiological signal timing logic verification results, it is determined whether to perform multimodal sleep emotion data feature extraction and fusion. Through multimodal signal timing calibration and verification, it is helpful to screen out non-physiologically related signal segments caused by timing disorder, and avoid subsequent modules from constructing incorrect feature associations based on signals without physiological correlation. At the same time, the timing calibration can be dynamically adjusted according to the physiological rhythm fluctuations during the user's sleep process, so that the calibrated sleep emotion state multimodal signal meets both clock synchronization requirements and fits the sleep physiological rhythm characteristics.
[0058] The multimodal feature fusion and state recognition module is used, if not activated, to perform multimodal sleep emotion data feature extraction and fusion, and to perform sleep-emotion state recognition based on the acquired multimodal sleep emotion data feature extraction and fusion results, and to obtain the sleep-emotion state recognition results. Through multimodal feature fusion and state recognition, it helps to integrate complementary physiological information from different modal signals to construct multidimensional physiological fusion features of sleep emotion states, improve the robustness of sleep emotion state recognition, and even if a certain modal signal is affected by environmental interference and data is missing or biased, the features of other modalities can be completed, providing a more forward-looking decision-making basis for subsequent collaborative intelligent regulation of sleep emotions.
[0059] The sleep-emotional collaborative regulation and safety protection module is used for intelligent sleep-emotional collaborative regulation based on sleep-emotional state recognition results. During the intelligent sleep-emotional collaborative regulation process, sleep-emotional responsive safety protection is implemented. Through sleep-emotional collaborative regulation and safety protection, dynamic linkage can be achieved during sleep regulation. The timing and intensity of intervention can be adjusted based on the dynamic trend of sleep-emotional state recognition, avoiding strong stimulation intervention during the transition phase of sleep state. At the same time, the monitoring data of safety protection is transformed into the basis for optimizing regulation strategies, ensuring the adaptability of sleep-emotional collaborative regulation.
[0060] As described above, the multimodal signal timing calibration and verification module, the multimodal feature fusion and state recognition module, and the sleep collaborative regulation and safety protection module help to build a personalized intelligent sleep emotion management system adapted to sleep physiological rhythms. This enables end-to-end personalized adaptation from sleep emotion multimodal signal processing to sleep emotion state recognition and then to sleep collaborative regulation and protection, improving the accuracy and safety of sleep emotion state management. Among them, the multimodal signal timing calibration and verification module defines the boundary of effective data for the multimodal feature fusion and state recognition module. Only multimodal signals that have passed the timing logic verification can enter the feature fusion stage. Its calibration accuracy directly determines the basic data quality of feature fusion. The multimodal feature fusion and state recognition module provides dynamic decision references for the sleep collaborative regulation and safety protection module. The accuracy of recognition directly determines the adaptability of the regulation strategy. The three modules are progressively integrated, forming a closed-loop optimization mechanism between modules.
[0061] In summary, by performing multimodal signal transmission timing interference analysis and based on the obtained results, it is determined whether to initiate fixed clock calibration alignment and dynamic resynchronization of multimodal signals. If initiated, after the activation is completed, sleep physiological signal timing logic verification is performed. Based on the obtained results of the sleep physiological signal timing logic verification, it is determined whether to perform multimodal sleep emotion data feature extraction and fusion. This helps to identify and correct timing deviations and interference problems in multimodal signal transmission from the source, ensuring that each modality signal maintains consistency and physiological logic rationality in the time dimension. This provides high-quality, timing-consistent basic data for subsequent feature extraction and fusion, reduces the risk of misjudgment of sleep-emotion states due to timing disorder of sleep emotion multimodal signals, and improves the accuracy of state recognition. If not initiated, multimodal... This method involves extracting and fusing features from multimodal sleep emotion data, and then identifying sleep-emotion states based on the obtained feature extraction and fusion results. This helps to eliminate redundant calibration steps and improve the efficiency of data processing and control response in scenarios with good multimodal signal transmission quality and no significant temporal interference, enabling rapid identification and real-time intervention of sleep-emotion states. Based on the sleep-emotion state identification results, collaborative intelligent control of sleep emotions is implemented. During the collaborative intelligent control process, sleep emotion responsive safety protection is provided, which helps to monitor changes in the physiological state of sleep emotions in real time during the control process, while taking into account the timeliness of control and the continuity of sleep. By triggering protective actions in a targeted manner, potential risks caused by improper control parameters are reduced, thereby improving the safety of intelligent control.
[0062] The above-disclosed embodiments are merely some examples of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A multimodal sensor signal analysis method for sleep mood state recognition, characterized in that, The method includes: Perform multimodal signal transmission timing interference analysis and determine whether to start fixed clock calibration alignment and multimodal signal dynamic resynchronization. If started, after the start is completed, perform sleep physiological signal timing logic verification and, based on the obtained sleep physiological signal timing logic verification results, determine whether to perform multimodal sleep emotion data feature extraction and fusion. The specific process of the multimodal signal transmission timing interference analysis is as follows: The interval sequence of data packets from the multimodal signal sensing channel arriving at the sleep mood state data processing center is monitored to obtain multimodal signal transmission jitter index; Determine whether the jitter index of multimodal signal transmission is greater than the preset jitter threshold. If so, start dynamic resynchronization of multimodal signal; otherwise, start fixed clock calibration alignment. The specific process of fixed clock calibration and alignment is as follows: A1. Extract the system synchronization clock signal from the sleep mood state data processing center and use it as a unified time reference for the multimodal signal sensing channel. A2, read the original acquisition timestamp of the data packet corresponding to each multimodal signal sensing channel, and use it as the local clock of the multimodal signal sensor; A3, calculate the clock deviation between the local clock of the multimodal signal sensor and the system synchronization clock; A4, perform a one-time clock correction on the timestamps of all data packets in the multimodal signal sensor channel according to the clock deviation, and obtain the one-time corrected timestamp; The clock one-time correction refers to superimposing the clock offset into the original timestamp of each data packet within the multimodal signal sensor channel; After the clock correction is completed, the multimodal signal sensor channel data are interpolated to a unified time axis based on the correction timestamp and an interpolation algorithm. A5. After the fixed clock calibration and alignment are completed, the timestamp deviation of the multimodal signal sensor channel data after alignment is verified. If the timestamp deviation is less than the preset timestamp deviation threshold, multimodal sleep emotion data feature extraction and fusion are performed; otherwise, multimodal signal dynamic resynchronization is performed. The specific process of dynamic resynchronization of the multimodal signal is as follows: Extract sleep physical event feature points and record the timestamps of these feature points. The specific extraction process is as follows: Obtain the number of valid peak occurrences; The number of effective peak occurrences represents the number of times the acceleration signal amplitude exceeds a preset amplitude threshold within the signal dynamic synchronization time period; If the number of valid peak occurrences is greater than the preset number of action judgments, the sleep physical event feature points corresponding to acceleration signal amplitudes greater than the preset amplitude threshold are marked as sleep physical event baseline feature points. The sleep physical event timestamp and the corresponding preset physiological signal response duration of the sleep physical event baseline feature points are obtained. Multimodal signal timing calibration is performed based on the sleep physical event timestamp and the preset physiological signal response duration. Otherwise, sleep physiological signal timing logic verification is performed. The process of performing multimodal signal timing calibration based on sleep physical event timestamps and preset physiological signal response durations is as follows: B1 extracts the timestamps corresponding to the physiological signal associated feature points in the multimodal signal sensor channel data that are common to the timestamps of sleep physical events. This means that a continuous signal segment with a total duration equal to the preset physiological signal response duration is extracted, centered on the timestamps of sleep physical events. Within the continuous signal segment, physiological signal associated feature points are extracted using a multimodal signal feature recognition algorithm, and the system synchronization timestamp corresponding to each physiological signal associated feature point is marked as a time sequence identifier for the sleep data points to be calibrated. B2, calculate the event-physiological time series deviation between the sleep data points to be calibrated and the corresponding sleep physical event characteristic points; The positive / negative selection interpolation algorithm based on event-physiological time series bias is as follows: If the event-physiological time sequence deviation is greater than the preset positive deviation judgment threshold, then based on the forward linear interpolation method, the timestamps of the data samples before and after the sleep physical event timestamps are forward shifted and corrected with a preset number of timestamps before and after the sleep data point to be calibrated. After the forward shift correction is completed, the sleep physiological signal time sequence logic verification is performed. Otherwise, the event-physiological time sequence reverse judgment is performed. The positive translation correction means that the timestamp of the sleep data point to be calibrated is translated in the positive direction of the time axis by using the absolute value of the event-physiological time sequence deviation as the translation amount; The event-physiological time series reverse discrimination means judging whether the event-physiological time series deviation is less than the preset negative deviation judgment threshold. If it is, then based on the backward linear interpolation method, the timestamps of a preset number of data samples before and after the sleep physical event timestamp are reverse shifted and corrected with the sleep physical event timestamp as the benchmark. After the reverse shift correction is completed, the sleep physiological signal time series logic verification is performed. Otherwise, the sleep physiological signal time series logic verification is performed. The reverse translation correction means that the timestamp of the sleep data point to be calibrated is translated in the negative direction of the time axis by using the absolute value of the event-physiological time sequence deviation as the translation amount; If not started, multimodal sleep emotion data feature extraction and fusion will be performed, and sleep-emotion state recognition will be performed based on the obtained multimodal sleep emotion data feature extraction and fusion results, and the sleep-emotion state recognition results will be obtained. Based on the sleep-emotion state recognition results, sleep-emotion collaborative intelligent regulation is carried out, and sleep-emotion responsive safety protection is implemented during the sleep-emotion collaborative intelligent regulation process.
2. The multimodal sensor signal analysis method for sleep mood state recognition as described in claim 1, characterized in that, The specific process for verifying the timing logic of sleep physiological signals is as follows: Step 1: Screen for combinations of sleep physiological signal indicators that reflect emotional-stress states; Step two, perform timing correlation verification, the specific process is as follows: Obtain the peak time of the sleep physiological signal index in the combination of sleep physiological signal indexes and the time interval between the peak time of each sleep physiological signal index, and determine whether the temporal relationship of each sleep physiological signal index satisfies the preset physiological temporal logic. The temporal relationship represents the quantifiable sequence and time interval between different physiological responses when sleep-related emotions trigger stressful events. If all sleep physiological signal indicators meet the preset physiological time sequence logic requirements, then the dynamic resynchronization is deemed effective, the corresponding multimodal sleep emotion data is marked as time-series qualified multimodal sleep emotion data, and multimodal sleep emotion data feature extraction and fusion are performed based on the time-series qualified multimodal sleep emotion data. If any sleep physiological signal indicator fails to meet the preset physiological timing logic requirements, a timing resynchronization failure prompt will be sent.
3. The multimodal sensor signal analysis method for sleep mood state recognition as described in claim 2, characterized in that, The specific process of extracting and fusing multimodal sleep emotion data features based on qualified time-series multimodal sleep emotion data is as follows: C1: Input the qualified sleep emotion modality signals from each time series into an independent lightweight encoder network to generate a multimodal primary feature embedding sequence. The specific process is as follows: The time-compliant sleep mood modality signals include time-compliant EEG signals, PPG signals, GSR signals, respiratory signals, and body movement signals. The EEG signal is input into a one-dimensional convolutional neural network to extract local features in the time and frequency domains, and the output is an EEG feature embedding sequence. The PPG signal is input into a recurrent neural network to extract its rhythm and morphological features, and the output is an HRV feature embedding sequence. The GSR signal is input into a one-dimensional convolutional neural network to extract its phase and tension component features, and the output is a GSR feature embedding sequence. After concatenating respiratory and body movement signals, the signals are input into a fully connected network to extract their rhythm and energy features, and the output is a respiratory and body movement feature embedding sequence. The concatenation of respiratory signals and body movement signals means that the preprocessed features of respiratory signals and preprocessed features of body movement signals corresponding to the same sampling time are horizontally merged in a preset concatenation order in the feature dimension to form a joint time-series signal sequence that integrates respiratory rhythm information and body movement behavior information. Based on preset splicing rules, the multimodal primary feature embedding sequence is spliced to form a hybrid multimodal sequence; C2 inputs the mixed multimodal sequence into the multimodal Transformer encoder, performs cross-modal feature fusion, and outputs a unified multimodal fused emotion state feature vector; C3 obtains sleep emotion intensity parameters based on multimodal fusion of emotion state feature vectors and sigmoid function, including physiological arousal intensity quantification value, physiological homeostasis intensity quantification value and sleep emotion intensity quantification value; Sleep-emotion state recognition is performed based on multimodal fusion of emotional state feature vectors and sleep emotional intensity parameters.
4. The multimodal sensor signal analysis method for sleep mood state recognition as described in claim 3, characterized in that, The sleep-emotion state recognition process is as follows: The multimodal fusion of emotional state feature vectors and sleep emotional intensity parameters is input into a preset classification model to output the sleep-emotional state category and the corresponding sleep-emotional state confidence score. The sleep-emotion collaborative intelligent regulation is based on sleep-emotion state categories and sleep-emotion state confidence levels. The specific process is as follows: Pairwise comparisons were made of the confidence scores of sleep-emotion state categories, and the sleep-emotion state category with the highest confidence score was marked as the dominant sleep-emotion state category. Based on the dominant sleep-emotional state category, obtain a preset combination of multi-sensory feedback strategies; The specific process for implementing regulation and monitoring sleep-emotional physiological responses is as follows: Determine whether the physiological signals of sleep-related emotions meet the initial regulatory target conditions. If so, gradually reduce the feedback intensity to 0 and terminate the current regulation. Conversely, if the effect of regulation is not good, a prompt will be sent, and the current sleep emotion collaborative intelligent regulation parameters and physiological response data will be fed back to the sleep emotion state data processing center.
5. The multimodal sensor signal analysis method for sleep mood state recognition as described in claim 1, characterized in that, The sleep-emotion-responsive safety protection process is as follows: Monitoring sleep-emotional physiological multimodal signals and obtaining sleep-emotional physiological state monitoring parameters based on feature fusion algorithms; The system compares the monitoring parameters of various sleep emotion and physiological states with the preset safety thresholds in real time. If any monitoring parameter of sleep emotion and physiological state triggers the protection logic, the corresponding control action is immediately executed, and a prompt that the safety protection has been activated is sent. At the same time, the value of the control execution counter is accumulated to obtain the number of control executions. Based on the number of control executions, sleep emotion-responsive safety protection is judged. The sleep emotion-responsive safety protection discrimination means determining whether the number of control executions exceeds the preset control execution threshold. If so, the sleep emotion-responsive safety protection method is triggered; otherwise, the sleep emotion physiological multimodal signal is continuously monitored.
6. The multimodal sensor signal analysis method for sleep mood state recognition as described in claim 1, characterized in that, The specific process of the sleep emotion-responsive safety protection is as follows: Obtain the deviation parameters of the multimodal signal; The current security level is determined based on the deviation parameters of the multimodal signal. The specific process is as follows: Determine whether the deviation parameter of any multimodal signal is less than the preset first-level deviation threshold. If so, execute the output limiting control strategy. Conversely, it continues to determine whether the deviation parameter of any multimodal signal is not less than the preset first-level deviation threshold and less than the preset second-level deviation threshold. If so, it will temporarily pause the output and send a physiological parameter abnormality and control pause prompt. Conversely, the system-wide safety disconnection control strategy will be implemented.
7. A multimodal sensor signal analysis system for sleep emotion state recognition, employing the multimodal sensor signal analysis method for sleep emotion state recognition as described in any one of claims 1-6, characterized in that, include: Multimodal signal timing calibration and verification module, multimodal feature fusion and state recognition module, and sleep collaborative regulation and security protection module: The multimodal signal timing calibration and verification module is used to perform multimodal signal transmission timing interference analysis and determine whether to start fixed clock calibration alignment and multimodal signal dynamic resynchronization. If started, after the start is completed, sleep physiological signal timing logic verification is performed, and based on the obtained sleep physiological signal timing logic verification results, it is determined whether to perform multimodal sleep emotion data feature extraction and fusion. The multimodal feature fusion and state recognition module is used to perform multimodal sleep emotion data feature extraction and fusion if not activated, and to perform sleep-emotion state recognition based on the obtained multimodal sleep emotion data feature extraction and fusion results, and to obtain the sleep-emotion state recognition results. The sleep-emotion collaborative regulation and safety protection module is used to perform sleep-emotion collaborative intelligent regulation based on sleep-emotion state recognition results, and to perform sleep-emotion responsive safety protection during the sleep-emotion collaborative intelligent regulation process.
Citation Information
Patent Citations
Multi-view fusion emotion recognition method under sleep deprivation condition
CN119538051A
Facial expression-based emotion real-time identification and long-term monitoring method
CN121527823A