AI sleep coach system based on dream metaphor and multi-mode interaction

By using an AI sleep coaching system based on dream metaphors and multimodal interaction, and leveraging the physiological data sensors and voice acquisition module within a smart pillow, combined with various modular technologies, the system addresses the issues of limited functionality and user cognitive biases in existing sleep aids. It enables personalized sleep intervention and cognitive correction, thereby improving sleep quality and user compliance.

CN121944332APending Publication Date: 2026-05-01ZHEJIANG YOUYOU MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG YOUYOU MEDICAL TECHNOLOGY CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing sleep aids are limited in function, lack interactivity and personalized guidance, fail to deeply explore the underlying psychological causes of insomnia, have rigid intervention methods, cannot correct users' cognitive biases, and have poor user compliance.

Method used

An AI sleep coaching system based on dream metaphors and multimodal interaction is adopted. Through physiological data sensors and voice acquisition modules in a smart pillow, combined with technologies such as metaphor-pathology mapping, dynamic questionnaire adjustment, multimodal biofeedback, hierarchical reinforcement learning decision-making, and voice synthesis synchronous control, personalized sleep intervention and cognitive bias correction can be achieved.

Benefits of technology

It enables accurate understanding of user status, provides personalized interactive guidance, improves the effectiveness of sleep intervention and user compliance, shortens sleep onset time, corrects users' cognitive biases, and improves sleep quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121944332A_ABST
    Figure CN121944332A_ABST
Patent Text Reader

Abstract

The invention relates to an AI sleep coach system based on a dream metaphor and multi-mode interaction in the technical field of intelligent sleep aiding. Comprising a metaphor-pathology mapping module, a dynamic questionnaire adjustment module, a multi-mode biofeedback emotional state recognition module, a reinforcement learning decision module, a speech synthesis synchronous control module, an intervention effect evaluation module and a cognitive deviation isomerization correction module. According to the method, the sleep state of the user is accurately and initially evaluated through the dream metaphor questionnaire, and a personalized sleep intervention scheme with emotion intelligence is constructed in combination with multi-mode biological signal perception and reinforcement learning decision-making mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

An AI sleep coaching system based on dream metaphors and multimodal interaction Technical Field

[0001] This invention relates to the field of intelligent sleep aid technology, and in particular to an AI sleep coaching system based on dream metaphors and multimodal interaction. Background Technology

[0002] With the fast pace and increasing pressure of modern life, more and more people are facing declining sleep quality. To improve this situation, various sleep aids have emerged on the market, such as light therapy devices, white noise machines, and aromatherapy diffusers. In today's world, high-quality sleep is crucial for maintaining physical and mental health. However, due to changes in lifestyle and the influence of the external environment, many people struggle to obtain sufficient quality sleep, and existing sleep aids have the following shortcomings:

[0003] Some sleep aids (such as white noise devices and ordinary sleep pillows) have limited functionality and lack interactivity and personalized guidance. While some smart sleep aids can monitor sleep, their analysis results are superficial, their intervention methods are rigid, and they cannot address the user's deeper psychological and cognitive issues.

[0004] For example, the questionnaires are static: the questions and options in the existing questionnaires used to develop sleep plans are fixed, which makes it impossible to deeply explore the potential psychological causes of users' insomnia and the accuracy of plan matching is low.

[0005] One-way intervention: Most products only monitor and intervene in one direction (such as playing music), lacking a feedback loop based on the user's real-time physiological state, so the intervention effect cannot be guaranteed and user compliance is poor.

[0006] Cognitive bias neglect: There is a lack of effective means to correct common cognitive biases among insomnia patients, such as "subjective insomnia" (i.e., actually sleeping but feeling like they haven't slept), making it difficult to establish positive sleep beliefs.

[0007] Therefore, there is an urgent need in this field for an intelligent solution that can accurately understand the user's state, provide personalized interactive guidance, and positively change the user's sleep cognition. Summary of the Invention

[0008] To address the aforementioned technical problems, this invention proposes an AI sleep coaching system based on dream metaphor-driven and multimodal interaction. The technical solution of this invention is implemented as follows:

[0009] This invention discloses an AI sleep coaching system based on dream metaphor and multimodal interaction. The system uses an intelligent sleep aid device, which includes a smart pillow, a voice acquisition module (such as a microphone), an audio playback module (such as a speaker), and a main control chip. The smart pillow integrates a physiological data sensor component and an airbag component. The physiological data sensor component is communicatively connected to the main control chip. The main control chip is equipped with a wireless interface for communication with the cloud.

[0010] The sleep coaching system includes a metaphor-pathology mapping module, a dynamic questionnaire adjustment module, a multimodal biofeedback emotional state recognition module, a reinforcement learning decision-making module, a speech synthesis synchronization control module, an intervention effect evaluation module, and a cognitive bias differential structure correction module.

[0011] The multimodal biofeedback emotion state recognition module integrates the user's voiceprint features, breathing features, and stress features to construct a three-dimensional emotion feature vector, thereby recognizing the user's emotional state and generating the user's emotional label, physiological urgency parameters, and sleep stages;

[0012] The dynamic questionnaire adjustment module evaluates the matching degree between the user's selected intention and the real-time physiological data based on the user's initial questionnaire and the user's real-time physiological data, and calculates the confidence degree of the matching degree; when the confidence degree is lower than the threshold, a second metaphorical supplementary questionnaire is triggered.

[0013] The metaphor-pathology mapping module associates the causes of insomnia, physiological characteristics, intervention plans, and user-selected images, and outputs the intervention plan.

[0014] The reinforcement learning decision-making module is used to execute and optimize the intervention plan;

[0015] The reinforcement learning decision-making module includes a macro-strategy layer and a micro-execution layer;

[0016] The macro-strategy layer selects intervention targets based on the user's sleep stages; the micro-execution layer executes intervention actions based on the intervention targets.

[0017] The reinforcement learning decision-making module includes a reward function; the reward function is used to optimize the intervention plan of the reinforcement learning decision-making module.

[0018] The speech synthesis synchronization control module collects the user's speech information through the speech acquisition module and receives the emotional tags and physiological urgency parameters from the multimodal biofeedback emotion state recognition module, and dynamically adjusts the speech guidance speed, pause duration and harmonic richness of the audio playback module.

[0019] The intervention effect evaluation module is used to calculate the DTW distance between the user's physiological response delay curve and the expected response template after the intervention plan is implemented, quantify the intervention effect, and transmit the intervention effect to the reinforcement learning decision module.

[0020] The cognitive bias differential structure correction module generates a multimodal comparison report based on the user's voice information from the speech synthesis synchronization control module and the sleep data from the multimodal biofeedback emotion state recognition module, thereby correcting the user's cognition. The user's voice data in this module is not from a questionnaire survey, but rather an independent data source. User voice information is collected in real-time by the voice acquisition module. The cognitive bias differential structure correction module matches the user's voice information with set keywords. When the user's voice information triggers a keyword (such as "I'm not asleep"), it compares the voice information with the user's objective sleep data to determine whether cognitive bias correction is necessary.

[0021] In this invention, sleep staging refers to the system dividing the sleep process into different stages based on the user's physiological signals (such as electrocardiogram and respiration), including wakefulness, light sleep, deep sleep, and REM (rapid eye movement) sleep. Sleep staging is generated in real-time by a multimodal biofeedback emotional state recognition module and is directly related to questionnaire surveys and the causes of insomnia.

[0022] The initial questionnaire is used to assess users’ subjective sleep problems (such as the causes of insomnia), while sleep staging provides objective physiological data on users. The two are combined to verify the accuracy of the questionnaire results (for example, if a user subjectively complains of “insomnia” but the objective staging shows that the deep sleep period is normal, then cognitive bias correction is triggered).

[0023] Sleep stage data can reflect the physiological characteristics of insomnia (e.g., prolonged light sleep may be related to anxiety) and be associated with imagery through the metaphor-pathology mapping module (e.g., the imagery of "desert" may map to respiratory disorders).

[0024] In this invention, the reinforcement learning decision-making module executes the intervention plan in a hierarchical manner. The macro level outputs the intervention target, which is the strategy level, and the micro level executes the intervention action, which is the execution level.

[0025] In this invention, the emotional label is indirectly related to the imagery selection in the initial questionnaire or metaphorical questionnaire (e.g., the user's selection of the "deep sea" imagery may reflect anxiety, and the intention will affect the emotional label).

[0026] The physiological urgency parameter directly depends on the real-time data of the user's emotional state identified by the multimodal biofeedback emotion state recognition module. This parameter is used to dynamically adjust the speech synthesis (e.g., the speech rate increases when the urgency is high).

[0027] Furthermore, the metaphor-pathology mapping module is constructed based on a graph neural network; the metaphor-pathology mapping module includes a dynamic knowledge base; the dynamic knowledge base includes intentional symbols, physiological indicators, psychological indicators, and relational edges;

[0028] The relation edges are used to map the dynamic knowledge base to the intervention plan.

[0029] Furthermore, the dynamic questionnaire adjustment module continuously evaluates the matching degree between user selection intentions and real-time physiological data through a hidden Markov model.

[0030] In this invention, the matching degree refers to the degree of consistency between the user's selected image and the user's real-time physiological data, and is a raw value (range 0-1) calculated using a Hidden Markov Model. The confidence degree is a reliability assessment of the matching degree and is a decaying weighted value.

[0031] The formula for calculating confidence level is as follows:

[0032] C(t)=α*C(t-1)+β*(1-|D_expected-D_actual|);

[0033] In the above formula, D is the physiological data bias, α is the decay factor of historical confidence (default is 0.8), and β is the weighting factor of current bias (default is 0.2); α+β=1 is satisfied to ensure smooth update of confidence; t is time, D_expected is the expected physiological value of image mapping; D_actual is the actual physiological value measured by the physiological sensor component;

[0034] In the above formula, the confidence level C(t) decays over time and is affected by physiological bias. The matching degree is an instantaneous assessment, while the confidence level is a cumulative assessment. When the confidence level falls below a threshold (e.g., 0.6), a second questionnaire is triggered.

[0035] The physiological values ​​include HRV (heart rate variability, in milliseconds), respiratory rate (breaths / min), and center of pressure shift velocity (mm / s).

[0036] This invention uses BCG sensing technology to capture minute mechanical vibrations caused by heart contractions and blood flow, with signals including information such as heart rate (HR), respiratory rate (BR), and cardiac output. Its advantage lies in achieving continuous monitoring without direct skin contact, making it suitable for embedded sensor devices such as mattresses and chairs.

[0037] Furthermore, the multimodal biofeedback emotion state recognition module includes a support vector machine classifier.

[0038] Furthermore, the micro-execution layer uses the Q-learning algorithm to select the specific actions of the intervention method;

[0039] The reward function is used to optimize action selection at the micro-execution layer;

[0040] Wherein, the reward function R(s,a) = W1ΔHRV + W2 user operation entropy H + W3 * intervention plan completion rate;

[0041] User operation entropy H = -Σp_i log p_i;

[0042] In the above formulas, △HRV represents the change in HRV, i.e., the difference between the HRV after intervention and the HRV before intervention; a positive value indicates improvement. W1, W2, and W3 are all weighting coefficients, and W1 + W2 + W3 = 1. p_i is the probability of the user skipping / repeating the operation. Through experimental calibration (default W1 = 0.4, W2 = 0.3, W3 = 0.3), W1, W2, and W3 represent the importance attached to physiological improvement, user compliance, and completion, respectively.

[0043] Furthermore, the speech synthesis synchronization control module uses a variant WaveNet architecture to perform emotion rendering on the synthesized speech and operates synchronously with the airbags in the airbag assembly.

[0044] The airbag assembly, which is installed inside the smart pillow, is used to implement physical intervention. The airbag assembly includes multiple modes, each corresponding to a different intervention action.

[0045] For example, Mode 1: Slow inflation and deflation (frequency 0.5Hz) is used for relaxation.

[0046] Mode 2: Medium fluctuations (1Hz frequency), used to promote sleep.

[0047] Mode 3: Rapid fluctuations (frequency 2Hz), used for REM intervention.

[0048] For example, the airbag voice synchronization process is as follows: The voice synthesis synchronization control module ensures that the audio playback and airbag inflation / deflation are synchronized through a hardware timer, with an error of <50ms. For example, the airbag inflates when the voice guides "inhale".

[0049] Furthermore, the reinforcement learning decision-making module includes a teacher model and a student model;

[0050] The teacher model is deployed in the cloud and consists of an 8-layer Transformer architecture.

[0051] The student model is deployed within the ESP32-S3 chip and consists of a 4-layer CNN.

[0052] The distillation damage function from the teacher model to the student model is: L = 0.7 * KL_div(teacher, student) + 0.3 * Cross_Entropy;

[0053] In the above formula, L represents the total loss; KL_div represents the difference in output distribution between the teacher model and the student model; and Cross_Entropy is the cross-entropy loss, which measures the error between the student model's predicted label and the true label. The weights 0.7 and 0.3 in the formula emphasize distribution matching.

[0054] Furthermore, the intervention effect evaluation module includes a response template library and a DTW distance calculation unit;

[0055] The system collects ideal response data to construct a response template library; the template for the ideal response data is a 30-second physiological indicator sequence.

[0056] The DTW distance calculation unit implements a lightweight DTW algorithm in a low-power embedded model to calculate the cumulative distance between real-time data and the template: D_t=Σ|template_i-data_i|+min(D_{t-1},D_{t-1}^vertical,D_{t-1}^horizontal);

[0057] The low-power embedded model is the deployment form of the reinforcement learning module implemented through knowledge distillation; when the cumulative distance value is less than the set intervention threshold θ, the effectiveness of the intervention plan is evaluated.

[0058] Furthermore, it also includes a federated learning evolution module;

[0059] The federated learning evolution module includes local model training and cloud parameter aggregation;

[0060] The smart sleep aid device saves the user's recent interaction data and performs local model training daily;

[0061] Local model training uses differential privacy to add Gaussian noise Δf = max|f(D) - f(D')|, where the noise is ~N(0,Δf / ε);

[0062] In the above formula, Δf is the sensitivity of function f, max|f(D)-f(D')| represents the maximum difference between adjacent datasets; the noise ~N(0,Δf / ε) is a Gaussian distribution with a mean of 0 and a variance controlled by the privacy budget ε.

[0063] The cloud aggregates locally trained parameters via the SecureAggregation protocol;

[0064] θ_global=Σ(θ_local_i*n_i) / Σn_i;

[0065] Where n_i is the amount of data in the local terminal device.

[0066] In this invention, the data from the physiological sensor component is sent to the ESP32-S3 chip via the I2C interface.

[0067] The ESP32-S3 chip communicates with the cloud (for federated learning) via a wireless communication interface (such as Wi-Fi).

[0068] Voice synthesis and airbag control are synchronized via GPIO and PWM interfaces.

[0069] The advantages of this invention are as follows:

[0070] A dynamic knowledge base based on metaphor-pathology mapping enables precise transformation from subjective imagery to objective intervention strategies, thereby improving mapping accuracy.

[0071] The dynamic questionnaire adjustment mechanism based on confidence decay solves the problem of single assessment bias and enables continuous self-correction of assessment results.

[0072] Multimodal biosignal fusion for emotion state recognition improves the accuracy of emotion recognition for users, significantly outperforming single voiceprint analysis or physiological signal analysis.

[0073] By using a hierarchical reinforcement learning decision model (HRL-DM), the poor adaptability of fixed strategies is addressed, and the acceptance of intervention programs is improved.

[0074] By using parametric emotional speech synthesis and synchronous control to achieve multimodal collaboration between voice intervention and physical intervention, the time it takes for users to fall asleep is shortened;

[0075] Achieving second-level evaluation of intervention effects through cross-modal alignment-based intervention effect assessment provides high-quality reward signals for reinforcement learning;

[0076] By using the user cognitive bias differential structure correction technology, the user's subjective insomnia fallacy is effectively broken, and the effectiveness of user cognitive correction is significantly improved.

[0077] Achieve low-power operation (power consumption < 1.2W) of complex AI algorithms on the terminal through low-power embedded model distillation and deployment, supporting real-time decision-making;

[0078] By employing a privacy-preserving federated learning evolution mechanism, the model can continuously evolve while protecting user privacy, thus solving the data silo problem. Attached Figure Description

[0079] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only one embodiment of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0080] Identical parts are indicated by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "up," and "down" used in the following description refer to directions in the accompanying drawings, while the terms "bottom surface," "top surface," "inner," and "outer" refer to directions toward or away from the geometric center of a specific part, respectively.

[0081] Figure 1 is a structural diagram of the pillow used in an embodiment of the present invention;

[0082] Figure 2 is a diagram of the internal structure of the pillow shown in Figure 1;

[0083] Figure 3 is a diagram of the air passage structure of the pillow shown in Figure 1.

[0084] The symbols in the above figures have the following meanings:

[0085] 1. Pillow insert;

[0086] 2. Speaker;

[0087] 3. Microphone;

[0088] 4. Heating pad;

[0089] 5. Pressure relief unit;

[0090] 6. Air passage;

[0091] 7. Solenoid valve;

[0092] 8. Air pump;

[0093] 9. Motherboard;

[0094] 10. Airbag assembly;

[0095] 11. Array pressure sensor. Detailed Implementation

[0096] The technical solutions of the present invention will now be clearly and completely described with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0097] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used in the detailed description is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0098] In the description of specific embodiments of the present invention, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present invention, "multiple" means two or more, unless otherwise explicitly defined.

[0099] In this invention, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this invention can be combined with other embodiments.

[0100] In the description of the embodiments of this invention, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this invention, the character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0101] The embodiments of the present invention will be described in more detail below through examples. It should be noted that the embodiments of the present invention are not limited to these examples.

[0102] In one specific embodiment, an AI sleep coaching system based on dream metaphor and multimodal interaction uses a smart sleep aid device as shown in Figures 1-3. The smart sleep aid device includes a smart pillow, a voice acquisition module (microphone 3), an audio playback module (speaker 2), and a main control chip. The smart pillow integrates a physiological data sensor component and an airbag component 10. The physiological data sensor component is communicatively connected to the main control chip. The main control chip is equipped with a wireless interface for communication with the cloud.

[0103] In this embodiment, the smart pillow includes a pillow core 1; a speaker 2 and a microphone 3 are provided on the pillow core 1; a heating pad 4, an airbag assembly 10 and an array pressure sensor 11 are provided in the area of ​​the pillow core 1 near the neck.

[0104] The pillow core 1 is equipped with an airbag assembly 10 and an air pump 8 connected by an air passage 6; the air passage 6 is equipped with a pressure relief unit 5 and a solenoid valve 7. The main control chip controls the on / off switch of the air pump 8 by controlling the solenoid valve 7. The main control chip is located on the motherboard 9.

[0105] In this embodiment, the main control chip is the ESP32-S3 chip.

[0106] In this embodiment, the heating pad is controlled by a main control chip for heating.

[0107] The system in this embodiment includes a dynamic knowledge base of metaphor-pathology mapping, a dynamic questionnaire adjustment model, an emotion state recognition model based on multimodal biological signal fusion, a hierarchical reinforcement learning decision model, a parameterized emotional speech synthesis and synchronization control module, an intervention effect evaluation module based on cross-modal alignment, a user cognitive bias deconstruction correction module, a low-power embedded model distillation and deployment module, and a privacy-preserving federated learning evolution module.

[0108] The dynamic knowledge base for metaphor-pathology mapping is constructed using a metaphor mapping model based on graph neural networks (GNNs). This dynamic knowledge base associates user-selected images (such as "desert") with causes of insomnia (Sjögren's syndrome), physiological characteristics (respiratory disorder > 0.35), and intervention plans.

[0109] The dynamic knowledge base for metaphor-pathology mapping contains over 200 nodes (image symbols, physiological indicators, and psychological labels) and 500+ relational edges, supporting real-time reasoning mapping.

[0110] The construction process of the dynamic knowledge base of metaphor-pathology mapping is as follows:

[0111] Construct a heterogeneous knowledge graph. The nodes of the heterogeneous knowledge graph include:

[0112] Image node: Attributes include visual feature vector and acoustic feature vector;

[0113] Physiological nodes: Attributes include HRV threshold range and respiratory disturbance index;

[0114] Intervention node: Attributes associated with CBT-I scheme ID and execution parameters.

[0115] The weights of the relationship edges are obtained through training with clinical data (e.g., "deep sea-anxiety", weight = 0.93, etc.);

[0116] The dynamic knowledge base for metaphor-pathology mapping employs a subgraph matching-based reasoning algorithm. When a user selects an image, a subgraph is extracted and its cosine similarity to the pathological pattern is calculated. The subgraph is a subnetwork extracted from the knowledge graph that is related to the user's image (e.g., if the user selects "deep sea", a subgraph containing the "deep sea" node and its associated edges is extracted).

[0117] The mapping results are output in JSON format: {"mapping scheme":"CBT-I_03","confidence level":0.89,"expected physiological indicators":{"HRV":>50ms,"respiratory rate":14-16 breaths / min}}.

[0118] The pathological pattern refers to the typical physiological-psychological combination of insomnia (such as "anxiety-type insomnia" corresponding to decreased HRV and rapid breathing), which is constructed in the metaphor-pathology mapping module through clinical data.

[0119] The dynamic questionnaire adjustment model includes a dynamic questionnaire adjustment mechanism based on confidence decay. After the user completes the initial questionnaire, the system continuously evaluates the matching degree between the user's choices and real-time physiological data through a Hidden Markov Model (HMM). When the confidence level falls below a threshold, a secondary metaphorical supplementary questionnaire is triggered.

[0120] The confidence level calculation formula is: C(t)=α*C(t-1)+β*(1-|D_expected-D_actual|), where D is the physiological data bias, and α and β are decay factors.

[0121] Where D_expected is the expected physiological value of the image mapping, and D_actual is the actual value measured by the sensor.

[0122] When C(t) < 0.6, a second questionnaire process is triggered.

[0123] The strategy for supplementing the questionnaire with secondary metaphors is as follows:

[0124] The system uses a multi-armed gambling machine algorithm to select alternative intentions orthogonal to the user's current intention from a dynamic knowledge base of metaphor-pathology mapping (e.g., if the user initially selects "deep sea", the alternatives are "desert" or "starry sky").

[0125] The device plays a voice question through the audio playback module in the smart sleep aid device, or connects to a wireless mobile device via a wireless module to display the question on the screen of the wireless mobile device, and then collects the user's voice input or other forms of information input (such as text).

[0126] A multimodal biosignal fusion-based emotion state recognition model includes a feature extraction unit and a classifier;

[0127] The feature extraction unit is used to extract the user's voiceprint features, breathing features, and stress features, and to construct a three-dimensional feature vector of emotions (anxiety / calm / resistance).

[0128] Voiceprint features include 1-12 dimensions of MFCC (Mel frequency cepstral coefficients), fundamental frequency perturbation (Jitter), and amplitude perturbation (Shimmer);

[0129] Respiratory characteristics include respiratory waveform harmonic distortion rate (THD) and inspiratory / expiratory time ratio;

[0130] Pressure characteristics include the velocity of the pillow's pressure center offset (mm / s).

[0131] The classifier is a support vector machine (SVM) classifier, which enables real-time emotion recognition on embedded devices with a latency of <80ms.

[0132] In this embodiment, a lightweight SVM classifier (with RBF kernel function and support vectors compressed to 200) is deployed on the ESP32-S3 chip; the classifier performs classification every 30 seconds and outputs emotion labels and probability values ​​based on the three-dimensional feature vector of emotion.

[0133] The Hierarchical Reinforcement Learning Decision Model (HRL-DM) consists of two layers of decision structure: a macro-policy layer and a micro-execution layer.

[0134] Among them, the macro strategy layer selects intervention targets based on the user's sleep stage.

[0135] This embodiment divides the user's sleep into the waking period, light sleep period, deep sleep period, and REM sleep period.

[0136] Intervention goals include promoting sleep onset, maintaining sleep, intervening in snoring, morning wakefulness, and shortening sleep latency.

[0137] The micro-execution layer uses the Q-learning algorithm to select specific actions (such as "playback scheme A + airbag mode 3").

[0138] In this embodiment, the micro-execution layer selects specific intervention instructions (plans) from a dynamic knowledge base of metaphor-pathology mapping based on the user's emotional state and physiological indicator deviation, such as "play audio_023" or "activate airbag mode 3". The physiological indicator deviation data refers to the user's real-time physiological data (such as respiratory rate) and emotional state data (from the multimodal biofeedback emotional state recognition module), which are used as state input for Q-learning.

[0139] The hierarchical reinforcement learning decision model (HRL-DM) also includes a reward function R = 0.4*(ΔHRV / HRV_max) + 0.3*(1-user operation entropy) + 0.3*scheme completion rate.

[0140] User operation entropy calculation formula: H=-Σp_i log p_i (p_i is the probability of skipping / repeating operations, etc.).

[0141] The parameterized emotion-based speech synthesis and synchronization control module receives emotion tags (such as "mild") and the user's physiological urgency parameters from the speech synthesis system, and dynamically adjusts the speech rate, pause duration, and harmonic richness based on an emotion parameter mapping table. The emotion parameter mapping table is shown below.

[0142] Emotional tags, speech rate scaling factor, fundamental frequency offset (Hz), harmonic richness

[0143] Mild 0.8x -20 +15%

[0144] Ethereal 1.0x +5 +30%

[0145] Stable 0.7x -30 -10%

[0146] The parameterized emotion speech synthesis and synchronization control module adopts a variant WaveNet architecture, supports real-time rendering of 22 emotion parameters, and is strictly synchronized with the airbag movement rhythm (the expected error is approximately <50ms).

[0147] In this embodiment, the audio playback device is an integrated speaker 2, which is installed inside the smart pillow and connected to the ESP32-S3 chip via an I2S interface.

[0148] Audio and airbag control are synchronized at the hardware level: interrupts are triggered by the hardware timer of the ESP32-S3 chip to ensure that the error between audio playback and airbag inflation / deflation is less than 50ms.

[0149] In this embodiment, the airbag is driven by an air pump 8 and a solenoid valve 7, and controlled by the GPIO of the ESP32-S3 chip. A synchronization mechanism uses a hardware timer to trigger an interrupt, ensuring that audio frame playback is synchronized with the airbag's movement.

[0150] The intervention effect assessment module based on cross-modal alignment includes a library of expected response templates; the templates are 30-second sequences of physiological indicators (such as the standard HRV rise curve, with the "successful relaxation template" corresponding to a 0.2-second HRV rise). 2 (Hz and above).

[0151] The intervention effect evaluation module based on cross-modal alignment calculates the DTW distance between the physiological response delay curve and the expected response template based on the feedback after the system executes the intervention plan, thus quantifying the intervention effect. The specific algorithm is as follows:

[0152] A lightweight DTW algorithm is implemented in the embedded system (referring to an embedded system with a MUC, deployed within the ESP32-S3 chip) to calculate the cumulative distance between real-time data and the template: D_t=Σ|template_i-data_i|+min(D_{t-1},D_{t-1}^vertical,D_{t-1}^horizontal).

[0153] A distance value less than the threshold θ (calibrated experimentally) is considered an effective intervention.

[0154] When the user's subjective complaint (voice: "I wasn't asleep") detected by the system does not match the objective data (actual 3 hours of sleep), the user cognitive bias differential structure correction module generates a multimodal comparison report. The generation process is as follows:

[0155] The visualization component uses the AntV G2 embedded chart engine to draw pie charts of sleep structure and time series line charts; the visualization component displays the sleep structure chart and highlights the difference between "actual deep sleep duration" and "self-reported sleep duration";

[0156] The voice report plays a data-driven voice explanation by splicing pre-recorded audio segments (such as "You actually" + "Deep sleep" + "72 minutes").

[0157] The user cognitive bias structure correction module is set with trigger conditions, such as preset keywords (e.g., "not asleep" / "insomnia") and the user's objective sleep duration > 2 hours.

[0158] Low-power embedded model distillation and deployment compresses large cloud-based decision models into lightweight models of <500KB using knowledge distillation technology and deploys them on the ESP32-S3 chip. It adopts a teacher-student architecture, with the teacher model being a 128-layer Transformer in the cloud and the student model being a 4-layer CNN (the compressed lightweight model size is 480KB).

[0159] Distillation loss function: L = 0.7 * KL_div(teacher, student) + 0.3 * Cross_Entropy.

[0160] This embodiment, through model lightweighting, not only improves the predicted data but also optimizes fuzzy processing.

[0161] This embodiment utilizes operator optimization, accelerates convolution calculation through vector instructions from the ESP32-S3 chip, and optimizes fuzzy processing using the distillation loss function described above, thereby improving decision accuracy and greatly reducing accuracy loss.

[0162] This embodiment incorporates a privacy-preserving federated learning evolution module. The system updates the metaphor mapping model through a federated learning framework: each terminal device trains the model parameters locally, and only the encrypted parameter increments are uploaded to the cloud for aggregation. Differential privacy technology is used to add Gaussian noise to ensure that the user's original data never leaves the terminal.

[0163] Local training process:

[0164] Each terminal device saves the most recent 100 interaction data points and trains a local model every morning at midnight;

[0165] Differential privacy is used to add Gaussian noise: Δf = max|f(D) - f(D')|, noise ~ N(0, Δf / ε)

[0166] Parameter aggregation:

[0167] The cloud aggregates parameters via the Secure Aggregation protocol: θ_global = Σ(θ_local_i*n_i) / Σn_i;

[0168] Where n_i represents the amount of data for each device, and homomorphic encryption is used during transmission.

[0169] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An AI sleep coaching system based on dream metaphor and multimodal interaction, using an intelligent sleep aid device, the intelligent sleep aid device comprising an intelligent pillow, a voice acquisition module, an audio playback module, and a main control chip; the intelligent pillow integrates a physiological data sensor component and an airbag component; the physiological data sensor component is communicatively connected to the main control chip; the main control chip is provided with a wireless interface for communication with the cloud; characterized in that, The sleep coaching system includes a metaphor-pathology mapping module, a dynamic questionnaire adjustment module, a multimodal biofeedback emotional state recognition module, a reinforcement learning decision-making module, a speech synthesis synchronization control module, an intervention effect evaluation module, and a cognitive bias differential structure correction module. The multimodal biofeedback emotional state recognition module integrates the user's voiceprint features, breathing features, and stress features to construct a three-dimensional emotional feature vector, thereby identifying the user's emotional state and generating the user's emotional label, physiological urgency parameters, and sleep stage. The dynamic questionnaire adjustment module evaluates the matching degree between the user's selected intentions and the real-time physiological data based on the user's initial questionnaire and real-time physiological data, and calculates the confidence level of the matching degree. When the confidence level is lower than the threshold, a second metaphor supplementary questionnaire is triggered; the metaphor-pathology mapping module associates the causes of insomnia, physiological characteristics, intervention plans with the images selected by the user, and outputs the intervention plan; The reinforcement learning decision-making module is used to execute and optimize the intervention plan; The reinforcement learning decision-making module includes a macro-strategy layer and a micro-execution layer. The macro-strategy layer selects intervention targets based on the user's sleep stages. The micro-execution layer executes intervention actions based on the intervention targets. The reinforcement learning decision-making module includes a reward function, which is used to optimize the intervention plan. The speech synthesis synchronization control module collects the user's speech information through the speech acquisition module and receives emotional tags and physiological urgency parameters from the multimodal biofeedback emotion state recognition module. It dynamically adjusts the speech guidance speed, pause duration, and harmonic richness of the audio playback module. The intervention effect evaluation module calculates the DTW distance between the user's physiological response delay curve and the expected response template after the intervention plan is executed, quantifies the intervention effect, and transmits the intervention effect to the reinforcement learning decision-making module. The cognitive bias differential structure correction module generates a multimodal comparison report based on the user's speech information in the speech synthesis synchronization control module and the sleep data from the multimodal biofeedback emotion state recognition module, thereby correcting the user's cognition.

2. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, The metaphor-pathology mapping module is constructed based on a graph neural network; the metaphor-pathology mapping module includes a dynamic knowledge base; the dynamic knowledge base includes intentional symbols, physiological indicators, psychological indicators, and relational edges; the relational edges are used to realize the mapping between the dynamic knowledge base and the intervention plan.

3. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, The dynamic questionnaire adjustment module continuously evaluates the matching degree between users' selection intentions and real-time physiological data through a hidden Markov model.

4. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 3, characterized in that, The confidence level is calculated using the following formula: C(t)=α*C(t-1)+β*(1-|D_expected-D_actual|); where D is the physiological data bias, α is the decay factor of historical confidence, β is the current bias weighting factor, t is the time, D_expected is the expected physiological value of the image mapping, and D_actual is the actual physiological value measured by the physiological sensor component; the physiological value includes HRV, respiratory rate, and pressure center shift velocity.

5. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, The multimodal biofeedback emotion state recognition module includes a support vector machine classifier.

6. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, The micro-execution layer uses the Q-learning algorithm to select the specific action of the intervention method; where the reward function R(s,a)=W1ΔHRV+W2User operation entropyH+W3*intervention plan completion rate; User operation entropyH=-Σp_ilogp_i; In the above formulas, ΔHRV is the change in HRV; W1, W2, and W3 are all weight coefficients, W1+W2+W3=1; p_i is the probability of the user skipping / repeating the operation.

7. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, The speech synthesis synchronization control module uses a variant WaveNet architecture to render the synthesized speech with emotion and operates synchronously with the airbags in the airbag assembly.

8. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, The reinforcement learning decision module includes a teacher model and a student model; the teacher model is deployed in the cloud and is an 8-layer Transformer; the student model is deployed in the ESP32-S3 chip and is a 4-layer CNN; the distillation loss function from the teacher model to the student model is: L = 0.7 * KL_div(teacher, student) + 0.3 * Cross_Entropy; in the above formula, L is the total loss value; KL_div is the difference in output distribution between the teacher model and the student model; Cross_Entropy is the cross-entropy loss.

9. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, The intervention effect evaluation module includes a response template library and a DTW distance calculation unit. The system collects ideal response data to construct the response template library. The template for the ideal response data is a 30-second physiological indicator sequence. The DTW distance calculation unit implements a lightweight DTW algorithm in a low-power embedded model to calculate the cumulative distance between real-time data and the template: D_t=Σ|template_i-data_i|+min(D_{t-1},D_{t-1}^vertical,D_{t-1}^horizontal). The low-power embedded model is the deployment form implemented by the reinforcement learning module through knowledge distillation. When the cumulative distance value is less than the set intervention threshold θ, the intervention plan is evaluated as effective.

10. The AI ​​sleep coaching system based on dream metaphor and multimodal interaction according to claim 1, characterized in that, It also includes a federated learning evolution module; the federated learning evolution module includes local model training and cloud parameter aggregation; the sleep aid device saves the user's recent interaction data and performs local model training daily; the local model training uses differential privacy to add Gaussian noise Δf=max|f(D)-f(D')|, noise~N(0,Δf / ε); in the above formula, Δf is the sensitivity of function f, max|f(D)-f(D')| represents the maximum difference between adjacent datasets; noise~N(0,Δf / ε) is a Gaussian distribution with a mean of 0 and a variance controlled by the privacy budget ε; the cloud aggregates the parameters of the local model training through the SecureAggregation protocol; θ_global=Σ(θ_local_i*n_i) / Σn_i; where n_i is the amount of data in the local terminal device.