Sleep-aiding interaction content generation method and system embedded with meditation and meditation guidance
By obtaining user physiological and behavioral data, performing multimodal signal processing and reinforcement learning, and generating personalized mindfulness meditation sleep aid interactive content, it solves the problem of lack of real-time monitoring and feedback adjustment in existing sleep aid products, and improves user experience and effect.
Patent Information
- Application Number
- CN202510704271.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-26
AI Technical Summary
Existing sleep aid products lack systematic integration into psychological intervention methods such as mindfulness meditation, which makes it difficult to monitor and feedback in real time to adjust user status, resulting in insufficient user experience.
By obtaining user physiological and behavioral data, multimodal signal processing is performed, sleep-aided interactive content containing mindfulness elements is generated, and reinforcement learning strategy optimization is carried out based on user interaction data, and content parameters and mindfulness elements combinations are adjusted in real time.
It realizes accurate perception of user status and adaptive content adjustment, improves the immersion, personalization and continuous effectiveness of mindfulness training, and is suitable for sleep aid and psychological adjustment for a wider population.
Smart Images

Figure CN120544800A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of physical and mental health and artificial intelligence content generation technology, and in particular to a method and system for generating sleep-aiding interactive content embedded with mindfulness meditation guidance. Background Art
[0002] Mood disorders such as insomnia and anxiety are widespread worldwide, severely impacting people's physical and mental health and quality of life. Mindfulness meditation has gained widespread recognition in psychotherapy, stress management, and sleep support. By guiding individuals to focus on their present state of mind and body and non-judgmentally accept their experiences, it can significantly reduce anxiety and promote relaxation.
[0003] Nowadays, various types of bedtime audio, stories and light interactive content are widely promoted on smart terminals to help users fall asleep.
[0004] However, existing sleep aid products usually only provide passive content playback, lack the systematic integration of psychological intervention methods (such as mindfulness meditation) into the plot or interaction, and find it difficult to monitor and provide feedback to users in real time. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this disclosure provides a method and system for generating interactive sleep-aiding content embedded with mindfulness meditation guidance. This addresses the problem that existing sleep-aiding products typically only offer passive content playback, lack systematic integration of psychological intervention methods (such as mindfulness meditation) into plots or interactions, and struggle to monitor and provide feedback on the user's real-time status.
[0006] According to a first aspect of the present disclosure, a method for generating sleep-aiding interactive content embedded with mindfulness meditation guidance is provided, comprising: obtaining physiological data and user behavior data of a user, performing multimodal signal processing and feature extraction on the user physiological data and user behavior data to obtain a user state feature vector; Performing real-time analysis based on the user state feature vector and a preset state assessment model to determine the quantitative indicators of the user's current physical and mental state; Dynamically generate content based on the quantitative indicators of physical and mental states and a preset mindfulness element template library to obtain sleep-aiding interactive content that includes naturally embedded mindfulness meditation elements; Real-time user interaction data on the sleep-aid interactive content is obtained, reinforcement learning strategy optimization is performed based on the real-time user interaction data on the sleep-aid interactive content, an optimal content adjustment strategy is determined, and playback parameters and / or mindfulness element combinations of the sleep-aid interactive content are adjusted in real time based on the optimal content adjustment strategy.
[0007] According to a second aspect of the present disclosure, a sleep-aiding interactive content generation system embedded with mindfulness meditation guidance is provided, which is used to execute the method according to the first aspect, including: a data acquisition and processing module, used to obtain user physiological data and user behavior data, and perform multimodal signal processing and feature extraction on the user physiological data and user behavior data to obtain a user state feature vector; A user state evaluation module is used to perform real-time analysis based on the user state feature vector and a preset state evaluation model to determine the quantitative indicators of the user's current physical and mental state; A sleep-aiding interactive content generation module is used to dynamically generate content based on the quantitative indicators of physical and mental states and a preset mindfulness element template library to obtain sleep-aiding interactive content containing naturally embedded mindfulness meditation elements; The interactive feedback collection module is used to obtain real-time interaction data of users on the sleep-aid interactive content, optimize the reinforcement learning strategy based on the real-time interaction data of users on the sleep-aid interactive content, determine the optimal content adjustment strategy, and adjust the playback parameters and / or mindfulness element combination of the sleep-aid interactive content in real time according to the optimal content adjustment strategy.
[0008] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: a memory and a processor. The memory stores a computer program. When the processor executes the program, the method described above is implemented.
[0009] In the sleep-aiding interactive content generation method and system embedded with mindfulness meditation guidance provided above, the disclosed embodiments achieve accurate perception of user status and adaptive adjustment of content, significantly improving the immersion, personalization, and sustained effectiveness of mindfulness training, lowering the participation threshold, and being suitable for the sleep-aiding and psychological adjustment needs of a wider range of people. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0011] Figure 1 A flowchart of a method for generating sleep-aiding interactive content embedded with mindfulness meditation guidance according to an embodiment of the present disclosure is shown; Figure 2 A flowchart of a method for generating sleep-aiding interactive content embedded with mindfulness meditation guidance according to an embodiment of the present disclosure is shown; Figure 3A schematic block diagram of a sleep-aiding interactive content generation system embedded with mindfulness meditation guidance according to an embodiment of the present disclosure is shown; Figure 4 A block diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0012] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0013] Those skilled in the art will understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and do not represent any specific technical meaning, nor do they represent the necessary logical order between them. It should also be understood that in the embodiments of the present disclosure, "multiple" may refer to two or more, and "at least one" may refer to one, two or more. It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly defined or given a contrary revelation in the context. In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present disclosure generally indicates that the associated objects before and after are in an "or" relationship. It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and the same or similar aspects thereof can be referenced to each other. For the sake of brevity, they will not be described one by one.
[0014] At the same time, it should be understood that for ease of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present disclosure and its application or use. Technologies, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and equipment should be considered part of the specification. It should be noted that similar numbers and letters represent similar items in the following figures, so once an item is defined in one figure, it does not need to be further discussed in subsequent figures.
[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0016] Figure 1 This is a flow chart of a method for generating interactive sleep-aiding content embedded with mindfulness meditation guidance, provided by an embodiment of the present disclosure. The method of this embodiment of the present disclosure aims to achieve accurate detection of large and small objects in an image.
[0017] S101 , obtaining physiological data and user behavior data of a user, performing multimodal signal processing and feature extraction on the physiological data and user behavior data of the user, and obtaining a user state feature vector.
[0018] Physiological data can be raw or processed data collected by smart terminals or lightweight sensors that reflect the user's current physiological state. It can include respiratory rate, respiratory depth, inhalation-exhalation ratio (obtained by analyzing respiratory audio with a microphone), heart rate, heart rate variability indicators (SDNN, LF / HF, etc.), facial expressions (using camera image recognition to extract key points and classify expressions), eye movement information (such as gaze point distribution, blinking frequency, pupil changes), EEG data (such as connecting an EEG headband), physical activity intensity, and posture change frequency (through an accelerometer).
[0019] User behavior data can be observable behavioral characteristics generated by users during the content interaction process, which are used to reflect their interaction intentions, participation and feedback status. Specifically, it includes interaction operation records (whether to continue, pause, skip, adjust content), interaction response rhythm (such as voice response rhythm, click frequency), breathing synchronization feedback actions (such as exhaling and inhaling according to prompts), behavioral signals in facial expressions and body movement reactions, length of stay, activity level, and whether they are immersed.
[0020] The user state feature vector is a multidimensional feature vector representing the user's current comprehensive physical and mental state, formed by multimodal signal processing and feature extraction of the user's physiological and behavioral data. The user state feature vector incorporates the following: respiratory rate parameters, respiratory depth parameters (extracted from respiratory signals), autonomic nervous system activity indicators (derived from HRV analysis), expression classification results (emotion recognition output by the CNN model), behavioral characteristics such as action intensity and interaction frequency. It may also include numerical dimensions such as heart rate, SDNN, LF / HF ratio, and movement amplitude.
[0021] Continuous user physiological and behavioral data can be collected through hardware such as microphones, cameras, photoplethysmography (PPG) sensors, and triaxial accelerometers integrated into smart terminals or external wearable devices. The respiratory audio signal collected by the microphone is processed through a Butterworth bandpass filter (passband range 0.1–0.6 Hz). The short-term energy envelope analysis algorithm is then used to extract the user's respiratory rate and depth. The inspiratory-expiratory ratio is calculated based on the duration ratio of the rising and falling segments within the respiratory signal cycle. The raw pulse wave signal obtained by the PPG sensor is first filtered through a 0.5–5 Hz IIR bandpass filter to remove baseline drift and high-frequency noise. The heartbeat peaks are then extracted using the threshold- and derivative-based Pan–Tompkins R-wave detection algorithm, and the sequence of RR intervals between adjacent peaks is calculated. The sequence is subjected to time-domain statistical analysis (such as SDNN, RMSSD) and wavelet transform (using Daubechies 4th-order mother wavelet) to extract heart rate variability (HRV) features that reflect the activity state of the autonomic nervous system, including parameters such as low frequency (LF), high frequency (HF), and LF / HF ratio. The facial image sequence captured by the camera is used to extract facial feature areas using the Dlib68-point facial landmark detection algorithm, and emotion classification is performed using a convolutional neural network (CNN) model to output facial expression emotion state classification results (such as calm, nervous, anxious, and focused). The motion signal collected by the three-axis accelerometer is used to quantify the user's physical activity intensity by calculating the signal root mean square (RMS), amplitude integral (AI), and acceleration change rate. If the device supports EEG acquisition, the user's EEG signal can also be subjected to a fast Fourier transform (FFT) to extract the frequency domain energy of alpha and beta waves for analyzing the user's cognitive concentration trend. After completing the aforementioned data processing, characteristic parameters across multiple modalities are obtained, including but not limited to: respiratory rate, respiratory depth, I / O ratio, HRV metrics (such as SDNN and LF / HF), expression classification results, physical activity intensity, and optional EEG frequency band energy metrics. These characteristic parameters are normalized (e.g., using Z-score standardization or Min-Max scaling) and concatenated in a fixed order to form a unified multidimensional numerical vector, the user state feature vector. This vector serves as input to subsequent state assessment models, enabling real-time inference of quantitative indicators of physical and mental states, such as relaxation, anxiety, and concentration.
[0022] Based on the above technical solution, optionally, the user's physiological data and user behavior data are obtained, including: Acquiring breathing audio data, performing bandpass filtering on the breathing audio data to obtain a processed audio signal, and determining the user's breathing frequency data based on the audio signal; Acquire a facial image sequence, perform key point detection and expression classification on the image sequence to obtain a classification result, and determine the user's facial expression data based on the classification result; Acquiring a motion signal, performing activity intensity analysis on the motion signal to obtain an analysis result, and determining the user's physical activity data based on the analysis result; Acquiring original pulse wave signal data, performing bandpass filtering on the original pulse wave signal data, extracting the pulse wave peak position, and calculating a continuous pulse interval sequence based on the pulse wave peak position; Performing time-domain statistical analysis and frequency-domain wavelet transform processing on the continuous pulse interval sequence to obtain heart rate variability data representing the activity state of the user's autonomic nervous system; Determining user physiological data based on the respiratory rate data and the heart rate variability data; User behavior data is determined based on the facial expression data and the body activity data.
[0023] In this solution, the breathing audio data can be the original breathing sound collected by a microphone; The audio signal may be a valid frequency band signal obtained by filtering the breathing audio; The respiratory rate data may be a respiratory rate parameter extracted from an audio signal; The facial image sequence may be a continuously captured facial image frame of the user; The classification result can be the expression label obtained by analyzing the facial image through the emotion recognition model; Facial expression data can be structured emotional information extracted from classification results for behavioral analysis; The motion signal may be motion data collected by an acceleration sensor; The analysis result can be the intermediate value of the exercise intensity analysis process; The physical activity data may be an assessed current physical activity level of the user; The raw pulse wave signal data may be a cardiovascular pulsation signal collected by a PPG sensor; The pulse wave peak position can be the peak time point of each heartbeat in the PPG signal; The continuous pulse interval sequence may be a sequence of time intervals between a plurality of adjacent pulse wave peaks; The heart rate variability data may be a physiological characteristic value reflecting the state of the autonomic nervous system calculated based on the sequence.
[0024] Acquiring breathing audio data involves collecting the user's breathing sound signals in a natural state in real time through the microphone of the terminal device. The system performs Butterworth bandpass filtering on the breathing audio data (typical passband range is 0.1–0.6 Hz), filters out background noise and non-target frequency band interference, and obtains an enhanced audio signal. Next, the short-time energy detection method is combined with zero-crossing rate analysis to extract the breathing cycle, and the number of breathing cycles per unit time is calculated to determine the user's breathing frequency data. A second-order IIR bandpass filter (passband 0.1–0.6 Hz) can also be applied to the breathing sound signal collected by the microphone. The formula is: ; Among them, x[n] is the breathing audio data at the current moment.
[0025] y[n] is the respiratory rate data at the current moment.
[0026] x[n-1], x[n-2] are the breathing audio data of the previous moment and the previous two moments.
[0027] y[n-1], y[n-2] are the respiratory rate data of the previous moment and the previous two moments.
[0028] b0, b1, b2 are the filter coefficients for the forward path.
[0029] a1 and a2 are the filter coefficients of the feedback path (note the negative signs before a1 and a2 in the formula). These coefficients determine the frequency response characteristics of the filter (such as center frequency and bandwidth).
[0030] At the same time, the system continuously collects a sequence of facial images from the user through the camera. These images are then processed using the Dlib or MediaPipe keypoint detection algorithms to locate facial features. These images are then fed into a pre-trained convolutional neural network (CNN) expression recognition model, which classifies and identifies the expressions in the image sequence, outputting classification results with probabilistic labels such as "calm," "nervous," and "focused." Based on these classification results and multi-frame consistency checks, the system determines the user's facial expression data, which is then used to further assess the user's emotional state and engagement.
[0031] To identify the user's physical movements, the system also collects motion signals using a three-axis accelerometer. After preprocessing, these signals use amplitude integration (AI) and root mean square (RMS) algorithms to calculate the overall activity energy level. Frequency analysis identifies the rhythm of movement, ultimately yielding an analysis result—a quantified activity intensity index. Based on this analysis, the system evaluates the frequency and amplitude of the user's movements within the current time window and determines their physical activity data, such as "stationary," "light activity," or "high-frequency movement."
[0032] The system also acquires raw pulse wave signal data from a wearable PPG sensor or a camera-based heart rate estimation device. This signal is then processed through a 0.5–5 Hz bandpass filter to retain the primary cardiac components. Using an R-wave peak detection algorithm based on threshold and derivative variations (e.g., a modified version of the Pan–Tompkins algorithm), the system extracts the pulse wave peak position for each heartbeat. Based on the time intervals between adjacent peaks, the system calculates a sequence of continuous pulse intervals (i.e., RR intervals).
[0033] Subsequently, the system performs time-domain statistical analysis (such as SDNN, RMSSD) and frequency-domain wavelet transform processing (using Daubechies mother wavelet) on the continuous pulse interval sequence to extract representative indicators of autonomic nervous system activity, including low-frequency power (LF), high-frequency power (HF), and LF / HF ratio, etc., to comprehensively form the user's current heart rate variability data. Based on the heart rate variability data and the aforementioned respiratory rate data, the system determines the user's physiological data to reflect their current state of relaxation and sympathetic / parasympathetic activity level. The pulse wave signal can also be obtained through the PPG sensor, and after 0.5-5Hz bandpass filtering, the RR interval sequence (the interval between adjacent heartbeat peaks) is extracted. The SDNN (time domain feature) is calculated using the following formula:
[0034] Where SDNN is the standard deviation of normal beat-to-beat intervals, reflecting the overall HRV level; RRi is the i-th RR interval (or NN interval, i.e., the time interval between the peaks of the R waves of two consecutive normal heartbeats, usually in milliseconds); RR is the average of all RRi intervals, RR=N1∑i=1NRRi; and N is the total number of RR intervals.
[0035] At the same time, the system integrates the aforementioned facial expression data and body activity data in a structured manner, and outputs user behavior data representing the user's current external behavior participation characteristics for subsequent status analysis and content adjustment reference.
[0036] In this solution, through the multi-dimensional acquisition and precise feature extraction of user physiological data and behavioral data, we can fully perceive the user's current physical and mental state, realize personalized and highly adaptable sleep-aid content generation and adjustment, effectively improve the targetedness and intervention effect of mindfulness guidance, enhance the user's immersion and continuous participation, and thus significantly improve the actual effectiveness of sleep-aid intervention and the quality of user experience.
[0037] On the basis of the above technical solution, optionally, multimodal signal processing and feature extraction are performed on the user physiological data and the user behavior data to obtain a user state feature vector, including: Performing short-time Fourier transform processing on the respiratory frequency data to obtain respiratory time-frequency characteristics, and determining a respiratory frequency parameter and a respiratory depth parameter according to the respiratory time-frequency characteristics; Inputting the facial expression data into a pre-trained expression recognition model for convolutional neural network processing to obtain a user emotional state classification result; performing wavelet transform analysis on the heart rate variability data, calculating time-frequency domain characteristic parameters, and determining an autonomic nervous system activity index based on the time-frequency domain characteristic parameters; The respiratory frequency parameter, the respiratory depth parameter, the user emotional state classification result and the autonomic nervous system activity index are subjected to feature fusion to obtain a user state feature vector.
[0038] In this scheme, the respiratory time-frequency feature can refer to a two-dimensional feature map obtained by performing short-time Fourier transform (STFT) on the respiratory signal, which is used to reflect the dynamic characteristics of the respiratory frequency changing over time, including information such as the main frequency band energy distribution and frequency drift.
[0039] The respiratory frequency parameter may be a frequency value extracted from the main energy concentration frequency band in the respiratory time-frequency feature, in units of times / minute, and is used to represent the user's current breathing rhythm.
[0040] The respiratory depth parameter may be a parameter obtained by calculating the maximum amplitude, peak-to-valley difference, or signal energy of the respiratory signal amplitude envelope, and is used to measure the amplitude and depth of breathing.
[0041] The pre-trained expression recognition model can be an image classification model based on a convolutional neural network (CNN), which is pre-trained on a large-scale facial expression dataset and is used to perform emotion recognition on an input facial image sequence.
[0042] The user emotional state classification result may refer to the classification result output by the expression recognition model for the current image sequence, representing the user's emotional state category, such as "relaxed", "tense", "focused", etc., and is accompanied by a probability score.
[0043] The time-frequency domain characteristic parameters can be parameters such as energy and frequency band ratio extracted in different frequency intervals (such as LF and HF) after performing wavelet multi-scale decomposition on the heart rate variability data, which are used to describe the heart rate regulation characteristics.
[0044] Autonomic nervous system activity indicators can be calculated from parameters such as the LF / HF ratio and SDNN in heart rate variability data. They represent the relative activity of the sympathetic and parasympathetic nervous systems. LF / HF is the ratio of low-frequency power (LF) to high-frequency power (HF) and is a core indicator in frequency-domain analysis of heart rate variability. SDNN, the standard deviation of adjacent normal RR intervals (NN), is one of the most commonly used indicators in time-domain analysis of HRV.
[0045] The system uses the Short-Time Fourier Transform (STFT) method to process the acquired respiratory frequency data. Specifically, the respiratory audio signal is first divided into several overlapping time windows, and a Fourier transform is performed within each time window to obtain a two-dimensional spectrum of frequency changes over time, namely the respiratory time-frequency characteristics. From this time-frequency spectrum, the system identifies the center position of the main frequency energy as the respiratory frequency parameter, and at the same time extracts the time-frequency energy intensity envelope or amplitude boundary to calculate the amplitude average or envelope area of each breath, thereby obtaining the respiratory depth parameter, which reflects the depth and rhythmicity of the user's breathing.
[0046] Next, to identify the user's facial expression state, the system inputs the collected facial expression data into an expression recognition model pre-trained on a large-scale facial expression dataset. This model, based on a Convolutional Neural Network (CNN) architecture, incorporates multiple layers of convolution, pooling, and a fully connected structure, capable of identifying key muscle changes in facial images. The model's output is the user's emotional state classification result, representing the emotion category label corresponding to the user's current facial expression, such as "relaxed," "anxious," "nervous," or "focused," along with a corresponding probability confidence value.
[0047] To analyze physiological heart rate characteristics, the system performs wavelet transform analysis on heart rate variability data (i.e., a series of continuous RR intervals) calculated from pulse wave signals. Using discrete wavelet transform (DWT) or continuous wavelet transform (CWT), combined with multi-scale decomposition methods, key frequency domain energy is extracted within the low-frequency (LF) and high-frequency (HF) regions, generating a set of time-frequency domain characteristic parameters representing sympathetic and parasympathetic nerve activity, such as LF power, HF power, LF / HF ratio, and SDNN. Based on these parameters, the system infers indicators of the user's current autonomic nervous system activity, thereby determining their level of physiological tension and relaxation.
[0048] Finally, the system normalizes and features the multiple core indicators extracted above, including respiratory rate parameters, respiratory depth parameters, user emotional state classification results, and autonomic nervous system activity indicators, and constructs them into a structured multidimensional vector, namely the user state feature vector.
[0049] The training process of the pre-trained expression recognition model includes: The pre-trained expression recognition model is based on a convolutional neural network (CNN) architecture and trained using supervised learning. During training, the system first collects a dataset of facial images containing various emotion categories (such as happiness, tension, concentration, and anxiety). This dataset is manually annotated or calibrated with psychological assessment labels to ensure that the images correspond to real-world emotional states. The image samples are then standardized, such as grayscale conversion, alignment, and normalization, and image enhancement techniques (such as rotation, cropping, and brightness perturbation) are used to increase sample diversity. The training network consists of multiple convolutional layers, Reluctant Unit (ReLU) activation functions, pooling layers, fully connected layers, and a Softmax output layer. The model weights are iteratively updated using backpropagation and gradient descent algorithms (such as Adam) by minimizing the cross-entropy loss function. After training, the model's accuracy is evaluated on an independent validation set, and the set of parameters with the best generalization performance is fixed as the final pre-trained model weights, which are used to classify user facial images during the online presence recognition phase.
[0050] In this solution, by extracting and fusing features from multiple sources of data such as breathing, facial expressions, and heart rate, and constructing a user state feature vector, we can more comprehensively and accurately identify the user's physical and mental state, achieve more personalized and dynamically responsive mindfulness sleep-aiding content generation, and effectively improve the relaxation guidance effect and user experience.
[0051] On the basis of the above technical solution, optionally, wavelet transform analysis is performed on the heart rate variability data to calculate time-frequency domain characteristic parameters, including: determining a continuous interval sequence based on the heart rate variability data, constructing a cumulative time axis based on the continuous interval sequence, and determining a time variable based on the cumulative time axis; A time domain signal is constructed according to a continuous interval sequence and a time variable, and the time domain signal and the time variable are input into a preset wavelet transform formula to obtain time-frequency domain representation data, and the time-frequency domain characteristic parameters are determined according to the time-frequency domain representation data.
[0052] In this solution, the continuous interval sequence can be the time interval sequence between two adjacent pulse wave peaks in the user's heartbeat signal, usually expressed as , ,..., Indicates. , T is the timestamp corresponding to each peak. This is the basic data for heart rate variability (HRV) analysis, reflecting the time-varying nature of heartbeat intervals.
[0053] The cumulative time axis can refer to a set of gradually accumulated time point sequences constructed based on a continuous interval sequence, i.e., t0=0, , ,..., This set of time points represents the cumulative time of each heartbeat from the start time, and is the horizontal axis reference required when constructing the time domain signal.
[0054] The time variable can refer to a sequence of sampling time points used to construct a heartbeat interval, usually composed of each This time variable is the "time input" required for wavelet transform and together with the interval value constitutes a set of two-dimensional time series sample pairs.
[0055] The time domain signal can be a sequence of consecutive intervals Mapping to time variables The signal sequence formed after the , This set of non-uniformly sampled time series data is used for subsequent continuous wavelet transform (CWT) to extract the local frequency features of heart rate changes.
[0056] The time-frequency domain representation data may refer to a two-dimensional transformation coefficient matrix calculated at multiple scales (frequencies) and time windows after inputting the above-mentioned time domain signal into a wavelet transform formula (such as CWT), which represents the energy distribution of a certain frequency component at a certain moment and is used to accurately describe the spectral structure of heart rate changes.
[0057] Time-frequency domain feature parameters can be quantitative statistical features extracted from time-frequency domain representation data, often including low-frequency power (LF), high-frequency power (HF), LF / HF ratio, energy center frequency, main frequency band bandwidth, etc. These parameters are used to characterize the relative activity of the sympathetic and parasympathetic nervous systems and are key physiological characteristics in the assessment of heart rate variability.
[0058] To achieve high-precision assessment of the user's autonomic nervous system status, the system first processes the collected pulse wave signal to obtain basic data for heart rate variability analysis. By applying a peak detection algorithm (such as a peak location method based on first-order derivative and threshold) to the filtered pulse wave signal, the system extracts the pulse wave peak position corresponding to each heartbeat and records its timestamp sequence. , ,..., The calculation formula based on the time difference between adjacent pulse wave peaks is: , we can get a set of continuous interval sequences { , ,..., }, this sequence reflects the dynamic changes between the user's heartbeats and is the basis for heart rate variability (HRV) analysis.
[0059] Next, the system constructs the corresponding cumulative time axis based on the above continuous interval sequence. Starting from the initial time point t0=0, each RR interval is accumulated in turn to form (k=1, 2, ..., n), thus obtaining a set of time point sequences representing the cumulative occurrence of each heartbeat { , ,..., The cumulative time axis is used to represent the distribution of RR intervals in the time domain and is the time reference for constructing subsequent time domain functions.
[0060] Based on the above accumulated time points, the system determines the time variable sequence { , ,..., }, this variable is used as the horizontal axis time input to correspond to the RR interval value one by one to form a binary signal point ( , The system constructs a complete time domain signal from this, that is, the RR interval value is regarded as the signal amplitude that changes with time, thereby obtaining the function RR(t), which represents the change process of the heart rate interval over time.
[0061] Subsequently, the system inputs the time domain signal and its corresponding time variable into the preset wavelet transform analysis function to perform a continuous wavelet transform (CWT). In this process, a mother wavelet ψ, such as the Morlet or Daubechies type, is used to map the RR(t) signal into a two-dimensional time-frequency domain representation data W(a, b) by scaling and translating it on different scale parameters a (representing frequency) and displacement parameters b (representing time). This data reflects the energy intensity of specific frequency components in different time windows.
[0062] Based on the obtained time-frequency domain data, the system calculates parameters such as energy integration, energy density distribution, and frequency ratio within a preset frequency range (e.g., 0.04–0.15 Hz for low frequency (LF) and 0.15–0.4 Hz for high frequency (HF). Ultimately, it outputs time-frequency domain characteristic parameters representing the user's sympathetic and parasympathetic nervous system activity, including but not limited to key physiological indicators such as LF power, HF power, LF / HF ratio, frequency center, and SDNN. These characteristics provide a quantitative basis for the system to further assess the user's autonomic nervous system state and support subsequent physical and mental state recognition and content adjustment logic.
[0063] In this solution, by performing wavelet transform analysis on heart rate variability data and extracting high-precision time-frequency domain feature parameters, it is possible to more accurately identify the user's sympathetic and parasympathetic nerve activity states, and achieve dynamic quantitative assessment of the physical and mental state, thereby providing a scientific and reliable physiological basis for the generation and adjustment of personalized sleep-aid content.
[0064] On the basis of the above technical solution, an optional, preset wavelet transform formula is: ;
[0065] in, represents the data in the time-frequency domain; a is the preset scale parameter, which controls the expansion and contraction of the wavelet; b is the preset translation parameter, which controls the position of the wavelet on the time axis; is the time domain signal; is the analysis wavelet after the mother wavelet is stretched and translated, where represents the complex conjugate.
[0066] In this scheme, a controls the degree of "stretching" of the wavelet, determining the "frequency" or "temporal resolution" of the analyzed signal. When a is small, the wavelet is compressed, analyzing high-frequency / fast-changing components; when a is large, the wavelet is stretched, analyzing low-frequency / slow-changing components. This can be preset through experimentation.
[0067] b controls the position of the wavelet on the time axis, indicating the current time point at which the signal's frequency components are being observed. This is equivalent to the center position of a sliding window, continuously shifting the wavelet across the entire signal to obtain local frequency information. b typically slides the signal at a specific time step, such as every 0.5 seconds or every RR interval. Multiple b values can be sampled at equal intervals across the entire cumulative time axis, enabling a temporal distribution scan of the entire heart rate signal.
[0068] It is the form of the mother wavelet function after scale transformation (a) and time shift (b); plus the superscript asterisk Denotes the complex conjugate (applicable to complex-valued wavelets, such as Morlet wavelets), which ensures energy conservation and orthogonality of the transformation. It represents the degree of matching of the signal x(t) at different frequencies (determined by a) and times (determined by b). For each scale a and position b, the system calculates the inner product to obtain the time-frequency energy at that point, forming the time-frequency domain representation data W(a, b).
[0069] S102: Perform real-time analysis based on the user state feature vector and a preset state evaluation model to determine a quantitative index of the user's current physical and mental state.
[0070] The preset state assessment model can be an artificial intelligence prediction model that performs real-time analysis on the user state feature vector input by the user to output a classification result or numerical level of the user's current physical and mental state.
[0071] Quantitative indicators of physical and mental state can refer to measurable values output by the state assessment model that reflect the user's current psychological and physiological state, including relaxation level: a value reflecting whether the user is in a relaxed and quiet state; anxiety level: a value reflecting the user's negative emotional tendencies such as tension and irritability; concentration: a value indicating the degree of concentration of the user's attention.
[0072] The user's state feature vector is input into a set of pre-set state assessment models for real-time inference. These pre-set state assessment models are pre-trained lightweight convolutional neural networks (CNNs). Their network architecture consists of an input layer, two convolutional layers, ReLU activation functions, a max pooling layer, a fully connected layer, and a softmax output layer. The input vector is first processed through convolution to extract local feature patterns. Feature integration is then achieved through pooling dimensionality reduction and a fully connected operation. Finally, a probability distribution output for the multi-classification problem is generated at the output layer. Each dimension of this probability distribution corresponds to a level label for each of the three state types: relaxation, anxiety, and concentration. The system determines the current state level based on the classification label corresponding to the highest probability, and extracts its numerical probability as a confidence indicator for that state dimension. To improve the stability of the output results, a sliding window mechanism is introduced to perform a weighted average of multiple consecutive output results to suppress single-frame fluctuations caused by sporadic behavior. Finally, combining the above model inference output and state smoothing results, the system determines the quantitative indicators of the user's physical and mental state at the current moment, that is, the normalized values of the three-dimensional state, which are used to dynamically control the mindfulness meditation element embedding strategy and rhythm adjustment logic in subsequent sleep-aid interactive content.
[0073] The training process of the preset state assessment model is: The pre-set state assessment model is used to predict quantitative indicators of the user's current physical and mental state, such as relaxation, anxiety, and concentration, in real time based on the input user state feature vector. To ensure the model's ability to accurately recognize and generalize multimodal input features, the model undergoes offline training. The training process includes five main stages: training data construction, feature preprocessing, model structure design, training strategy setting, and performance evaluation.
[0074] First, during the model training phase, a large amount of experimental data was collected from the target user group. The data included breathing audio signals, original pulse wave signals, facial image sequences, acceleration signals, and auxiliary EEG signals during the interaction with sleep-aiding content. Each set of data was accompanied by subjective label results determined through psychometric questionnaires or expert annotation. The labels covered the user's current level of relaxation, anxiety, and concentration, which were used to supervise model training.
[0075] During data preprocessing, all raw signals undergo denoising, filtering, and normalization, and are then converted into a unified user state feature vector using a feature extraction algorithm. Features include, but are not limited to, respiratory rate and depth parameters extracted from breathing audio, RR interval sequences and heart rate variability features (SDNN, LF / HF) extracted from PPG signals, expression classification results extracted from image sequences, and physical activity intensity indicators extracted from acceleration data. All features are concatenated to form a standard input vector and aligned with the labels.
[0076] In terms of model structure, a lightweight convolutional neural network (CNN) is preferred. The input layer receives the aforementioned multimodal fusion vector. The core structure includes two convolutional layers (3×3 convolution kernels with ReLU activation), one max pooling layer, two fully connected layers, and a softmax output layer, which outputs the classification probability distributions of relaxation, anxiety, and concentration, respectively. This structure enables fast inference in resource-constrained terminal devices and is suitable for embedded deployment.
[0077] Training uses the cross-entropy loss function as the objective function and the Adam optimizer for parameter updates. To improve model generalization, data augmentation strategies (such as random noise perturbation and time series jitter) are introduced. Early stopping and validation set monitoring are used during training to prevent overfitting. After training, the model's accuracy, recall, and confusion matrix are verified on an independent test set to assess its stability in multi-state prediction tasks.
[0078] The final trained state assessment model is embedded in the terminal system as a basic judgment tool for real-time prediction of quantitative indicators of the user's current physical and mental state, providing a decision-making basis for the dynamic embedding of mindfulness elements and content rhythm control. At the same time, the extracted user state vector is input into the preset state assessment model. For example, for image features (such as facial expression spectrograms) or sequence features (such as treating physiological signal segments as one-dimensional images) based on convolutional neural networks (CNNs): the output fk(l) of the convolutional layer is usually calculated as:
[0079] For one-dimensional convolution, it is:
[0080] in: or is the output of the k-th feature map of the l-th layer at position (i, j) or (i).
[0081] is the activation function (such as ReLU, Sigmoid); is the cth input feature map of the previous layer (the l-1th layer); (x, y) or (x) is the value of the convolution kernel (weight) connecting the cth feature map of the l-1th layer to the kth feature map of the lth layer at position (x, y) or (x); , Or K is the width and height or length of the convolution kernel; is the bias term of the k-th feature map of the l-th layer; the final output layer of the model (such as the Softmax layer) will give a probability distribution indicating the possibility of the user belonging to each predefined state (such as "relaxed", "mildly anxious", and "focused").
[0082] S103, dynamically generating content based on the quantitative indicators of the physical and mental state and a preset mindfulness element template library to obtain sleep-aiding interactive content containing naturally embedded mindfulness meditation elements.
[0083] The pre-set mindfulness element template library is a structured collection of mindfulness-inducing materials. Each template element includes a mindfulness meditation element type, an adaptive state label (such as anxiety, high focus, low relaxation), and corresponding content embedding method, pacing requirements, and presentation format (voice / image / text). These rules are used to dynamically filter adaptive elements based on the user's current state. Each mindfulness element template defines: element name (such as deep breathing, five senses perception), scenario adaptation label (suitable for anxiety relief, attention maintenance, etc.), embedding method (such as voice instructions, natural background transitions), embedding rhythm (fast-paced guidance, slow-paced soothing), and media type (text, voice, image).
[0084] Mindfulness meditation elements can refer to the psychologically stimulating guided content segments commonly used in meditation training. They guide users through practices such as awareness, breathing, observation, and emotional acceptance. They are the core psychological conditioning structure embedded in sleep-aiding content. These include guided deep breathing, five senses (sight, hearing, touch, smell, and taste), non-judgmental observation, mental projection, and spatial relaxation imagery.
[0085] Sleep-aiding interactive content can refer to multimedia content generated by embedding mindfulness elements and presented to users. It contains elements such as narrative text, voice guidance, image scenes, and has the characteristics of continuous interaction with users (such as following voice rhythm, synchronizing breathing guidance, etc.), which is used to help users relax, relieve anxiety, and improve sleep quality.
[0086] The system uses real-time inference to quantify the user's physical and mental state—namely, the user's current state level across three dimensions: relaxation, anxiety, and focus—as input to drive content generation. This metric reflects the user's psychological and physiological state trends, and the system uses it as a key judgment to select the mindfulness meditation elements that best match their current state. The system then dynamically extracts appropriate mindfulness element combinations and embedding strategies from a pre-set library of mindfulness element templates.
[0087] The preset mindfulness element template library predefines a variety of common mindfulness guidance types, including but not limited to content segments such as guided deep breathing, five senses observation, non-judgmental observation and psychological projection, and configures a state adaptation label for each type of element (for example, deep breathing is highly adapted for anxiety, and five senses observation is less adapted for focus), as well as its corresponding media presentation method (voice, text, image), rhythm control requirements and embedding position rules (such as before the introduction, the peak of the story, the relaxation period at the end of the paragraph, etc.).
[0088] The system first ranks the templates' adaptability tags based on their current quantitative indicators of physical and mental state, selecting the optimal combination of mindfulness meditation elements. This combination is then fed into a pre-set story structure template engine, which constructs a complete narrative flow using a node-branch structure. The selected mindfulness elements are inserted at specific logical nodes, such as scene transitions, character emotional shifts, and climaxes.
[0089] During the element embedding process, the system uses the natural language generation engine, speech synthesis services (such as TTS), and image generation interfaces to convert the corresponding meditation elements into specific script content, audio clips, and accompanying images based on the rhythm rules and content type control methods defined in the template library. All generated content is formatted as a structured multimodal script file, including plot segments, audio paths, rhythm markers, and interactive prompts.
[0090] Finally, the system executes the content synthesis process based on the script file and outputs sleep-aiding interactive content. The embedded mindfulness meditation elements are presented in the form of voice guidance, text prompts, and visual image changes, and are naturally integrated with the original narrative plot. Users can perceive and participate without explicit operation.
[0091] Based on the above technical solution, optionally, dynamic content generation is performed based on the quantitative indicators of physical and mental state and a preset mindfulness element template library to obtain sleep-aiding interactive content containing naturally embedded mindfulness meditation elements, including: Matching the mindfulness guidance type based on the quantitative indicators of the physical and mental state and the element label data in the preset mindfulness element template library to obtain the currently adapted mindfulness meditation element combination; Based on the mindfulness meditation element combination and the node-branch structure data in the narrative structure template engine, element embedding position and rhythm control are performed to obtain a structured script template containing mindfulness element embedding rules; Audio data, image data, and text data are generated according to a structured script template, and multimodal content synthesis is performed based on the audio data, image data, and text data to obtain sleep-aiding interactive content that includes naturally embedded mindfulness meditation elements.
[0092] In this solution, the preset mindfulness element template library can be a data set pre-built in the system, which contains a variety of mindfulness meditation guidance elements (such as deep breathing, five senses perception, non-judgmental observation, psychological projection, etc.) and their parameterized descriptions, applicable state labels, semantic types, expression forms, etc.
[0093] The element label data can be the label set for each element in the mindfulness element template, which is used to indicate the user state it is adapted for (such as "high anxiety", "low relaxation", "medium concentration") and the corresponding functional goals (such as "soothing emotions", "guiding awareness"), and is used to achieve automatic matching with quantitative indicators of physical and mental states.
[0094] The combination of mindfulness meditation elements can be a set of mindfulness guidance types and content combinations dynamically selected by the system based on the matching results of the current user's quantitative indicators of physical and mental state and element label data, such as "deep breathing + non-judgmental observation", "five senses perception + psychological projection", etc., indicating the meditation intervention form suitable for the current user's state.
[0095] The narrative structure template engine can be a logic generation engine used to control the location, rhythm and form of embedding mindfulness elements in stories or interactive content, and a script structure control framework built based on node-branch logic.
[0096] The node-branch structure data can be the definition data of each semantic node and its possible branch direction in the narrative script structure, representing the logical path diagram of the interactive narrative process, which is used to specify where the mindfulness elements are embedded and how the content develops.
[0097] The mindfulness element embedding rules can be the strategic rules set by the system based on the user's current state, the characteristics of the mindfulness elements and the narrative structure, which determine at which node each element should be embedded, in what form it should be presented, and for how long it should last.
[0098] A structured script template can be a script file that combines a node-branch structure and mindfulness element embedding rules. It describes the logical flow, storyboard structure, and guiding element arrangement of the complete sleep-aiding interactive content, and is the basic framework for generating audio and video texts.
[0099] Audio data can refer to voice guidance tracks automatically generated using AI speech synthesis technology (such as TTS or Text-To-Speech) based on the voice content, rhythm, and emotional tags specified in a structured script template. This data is used to play mindfulness meditation guidance, such as deep breathing prompts, awareness guidance, and relaxation narration. It typically includes human voice audio with a steady rhythm, soft intonation, and customized style.
[0100] Image data can refer to static or dynamic visual material generated by image generation models (such as image library retrieval and image diffusion models) based on a combination of mindfulness elements and narrative rhythm. These images are used to create meditation scenes (such as natural scenery, light and shadow changes, and immersive images), and are combined with audio rhythm to create a relaxing atmosphere. The generated images can be realistic, low-polygon abstract, or meditative animations.
[0101] Text data refers to text content generated based on the semantic content of mindfulness meditation elements using language model generation algorithms (such as NLG, or Natural Language Generation) to assist with screen presentations, subtitle guidance, and task instructions. This text data can be displayed synchronously as subtitles for on-screen guidance or used as input for TTS speech, facilitating the alignment of speech and text output.
[0102] Based on the user's current quantitative indicators of their physical and mental state, the system analyzes their relaxation, anxiety, and concentration levels, uses these as query criteria, and matches them against the element label data stored in the mindfulness element template library. This element label data includes the adaptive state label (e.g., "low relaxation, high anxiety," "moderate concentration") associated with each mindfulness-guiding element (e.g., guided deep breathing, five-sense awareness, non-judgmental observation, and psychological projection), as well as the functional target label (e.g., "guided relaxation," "enhanced awareness," and "emotional relief"). Through label similarity matching and a state vector scoring mechanism, the system selects the set of mindfulness content that best matches the current state, forming a currently applicable combination of mindfulness meditation elements that serves as the core intervention plan for this round of content generation.
[0103] Next, the system searches for the corresponding semantic script fragments for each mindfulness guidance type within the mindfulness meditation element combination, and integrates this with the node-branch structure data from the pre-built narrative structure template engine to perform embedding control. Each node represents a semantic unit within the content structure, such as "story guidance," "transition emotional prompt," or "meditation guidance," while the branch structure defines the jump logic and rhythmic flow path between nodes. Based on the current user state and element characteristics, the system dynamically determines the node where each mindfulness element should be embedded (such as inserting the "Non-Judgemental Observation" element into the "Emotional Awareness" node), sets the embedding method (voice, image, subtitles) and rhythm control parameters (such as presentation time, pause intervals, etc.), and ultimately generates a structured script template containing the mindfulness element embedding rules, which serves as the logical blueprint for content production.
[0104] The system then uses the semantic content, emotional tags, and embedding strategies within the structured script template to call upon the speech synthesis service, image generation model, and natural language generation model, generating corresponding audio data (e.g., meditation guidance audio), image data (e.g., meditation background images, transition animations), and text data (e.g., screen display prompts, subtitles). All generated data is integrated and synchronized under a unified content arrangement strategy, performing cross-modal alignment and playback sequence control, completing a complete multimodal content synthesis process. The final output is personalized, immersive, and interactive sleep-enhancing content that naturally incorporates mindfulness meditation elements.
[0105] In this solution, by intelligently matching the user's quantitative indicators of physical and mental state with the mindfulness element template library, and combining it with the narrative structure template to achieve the rhythm embedding of guiding elements and multimodal content generation, it is possible to dynamically generate personalized and naturally embedded sleep-aiding interactive content, effectively improving the immersion and adaptability of mindfulness intervention, and enhancing the user's relaxation experience and willingness to continue using it.
[0106] S104: Acquire real-time user interaction data on the sleep-aid interactive content, perform reinforcement learning strategy optimization based on the real-time user interaction data on the sleep-aid interactive content, determine an optimal content adjustment strategy, and adjust playback parameters and / or mindfulness element combinations of the sleep-aid interactive content in real time based on the optimal content adjustment strategy.
[0107] Real-time interaction data refers to the interactive behavior information generated by users while watching or engaging in sleep-enhancing interactive content, which can be collected and processed in real time and used as a basis for content adaptive adjustment and reinforcement learning strategy training. This data may include whether the user continues to watch or actively exits the content, whether the user's breathing movements are synchronized with the rhythm, whether the user expresses positive emotions during the interaction (such as changes in facial expression and increased concentration), whether the user triggers adjustment actions (such as speeding up or slowing down the rhythm), breathing matching rate, interaction duration, and response frequency.
[0108] An optimal content adjustment strategy can refer to a policy network continuously updated during training using reinforcement learning methods (such as Deep Q-Networks (DQN)). Based on the user's current state and historical interaction feedback, it outputs the optimal action plan for controlling content playback behavior. This can include selecting the type of mindfulness element to continue or change, determining whether to speed up or slow down the playback rhythm, controlling content length and guidance density, mitigating the risk of user dropout and extending immersion time.
[0109] Playback parameters refer to a set of adjustable parameters that control the presentation of sleep-aiding interactive content during playback, used to optimize the content's cadence, presentation, and user fit. These parameters may include content playback speed (fast / slow voice tempo), content segment duration (compressed / lengthened), audio tempo curve (matching breathing rhythm), image switching frequency, background change speed, and content guidance density (number of prompts per minute).
[0110] A mindfulness element combination refers to a set of mindfulness meditation elements selected for embedding within the content structure during the current phase. The combination and sequence are dynamically generated based on the user's state and strategy model to guide relaxation and focus. This may include guided deep breathing, five-sense awareness, non-judgmental observation, and mental projection and imagery.
[0111] During the playback of interactive sleep-aiding content, the system continuously collects user behavior and physiological responses, forming real-time interaction data. This real-time interaction data includes: an interruption flag indicating whether the user actively terminated content playback, the matching rate of the user's breathing rhythm with system prompts, changes in emotional state inferred through facial expression recognition, and the degree of interaction coordination determined through motion detection. The breathing synchronization rate is achieved by phase-matching the real-time collected breathing rate with the system-set guidance rhythm; expression recognition uses a lightweight convolutional neural network (CNN) to classify camera image sequences and output the current emotion label and confidence level; interaction interruption behaviors are directly extracted by the behavior recorder in the content playback system.
[0112] This real-time interaction data is fed into the policy optimization process in real time, where it is combined with the current system state to form state-action pairs within a reinforcement learning environment. The system generates immediate reward signals based on the actions corresponding to the current content playback (such as increasing the tempo or replacing elements) and user feedback. The reward function is calculated based on multiple metrics: positive rewards are awarded if breathing synchronization improves, facial expressions become calm or focused, and the user maintains playback; negative rewards are awarded if interaction is interrupted, the user's state deteriorates, or their compliance rate decreases.
[0113] The system uses a Deep Q-Network (DQN) architecture for reinforcement learning training. It takes the current state (quantified indicators of physical and mental state + interaction feedback) as input and outputs a Q-value corresponding to the content-adjusted action. An experience replay pool stores the historical state-action-reward-next-state quadruple for periodic training of network parameters. Whenever new interaction data is generated, the system calculates an immediate reward and uses the replay pool samples to batch update the network, optimizing the strategy.
[0114] After the policy network is updated, the system inputs the current state into the network and infers the current optimal action as the optimal content adjustment strategy. The output of this policy control includes two aspects: Playback parameters: such as voice rhythm adjustment (controlled by TTS speech rate parameters), content paragraph length change (dynamic insertion or deletion of buffer segments), background beat change (adjusting background music rhythm through beat annotation), etc. Mindfulness element combination: Reselect mindfulness meditation elements that adapt to the current situation from the preset mindfulness element template library, such as replacing "guided deep breathing" with "non-judgmental observation", or adjusting the order of element insertion and presentation frequency.
[0115] Finally, the system modifies the subsequent content playback stream in real time based on the strategy output, so that the new sleep-aiding interactive content is more in line with the current user state in terms of rhythm, plot and meditation elements, thereby improving the guidance effect and immersion.
[0116] In an embodiment of the present application, the user's physiological data and user behavior data are obtained, and multimodal signal processing and feature extraction are performed on the user's physiological data and user behavior data to obtain a user state feature vector; real-time analysis is performed based on the user state feature vector and a preset state assessment model to determine the user's current physical and mental state quantitative index; dynamic content generation is performed based on the physical and mental state quantitative index and a preset mindfulness element template library to obtain sleep-aiding interactive content containing naturally embedded mindfulness meditation elements; the user's real-time interaction data on the sleep-aiding interactive content is obtained, and reinforcement learning strategy optimization is performed based on the user's real-time interaction data on the sleep-aiding interactive content to determine the optimal content adjustment strategy, and the playback parameters and / or mindfulness element combination of the sleep-aiding interactive content are adjusted in real time according to the optimal content adjustment strategy. Through the above-mentioned sleep-aiding interactive content generation method embedded with mindfulness meditation guidance, accurate perception of user status and adaptive content adjustment are achieved, which significantly improves the immersion, personalization and sustained effectiveness of mindfulness training, lowers the participation threshold, and is suitable for the sleep-aiding and psychological adjustment needs of a wider range of people.
[0117] On the basis of the above technical solution, optionally, a reinforcement learning strategy is optimized based on the real-time interaction data of the user on the sleep-aid interactive content to determine the optimal content adjustment strategy, including: Perform multimodal feature extraction on the real-time interaction data to obtain a comprehensive feature vector, and perform multi-dimensional evaluation calculation based on the comprehensive feature vector and preset reward rules to obtain an instant reward value for the sleep-aiding interactive content; Reinforcement learning strategy optimization is performed according to the instant reward value and a preset reward tuple, a strategy network for content decision-making is updated, and an optimal content adjustment strategy is determined according to the strategy network.
[0118] In this solution, the comprehensive feature vector refers to a unified representation vector formed by integrating multimodal state feedback information (such as facial expression changes, breathing changes, heart rate variability fluctuations, interaction delays, etc.) extracted from the user's real-time interaction data. This vector is used to characterize the actual feedback changes of the user's state while receiving the current sleep-aiding interactive content.
[0119] The pre-set reward rules can be a system of evaluation criteria used to assess whether the current content output achieves the goals of relaxation, stability, and immersion. This rule includes a multi-dimensional scoring mechanism based on target indicators (such as improved breathing stability, reduced heart rate fluctuations, and positive changes in facial expressions) to calculate the immediate reward value.
[0120] The instant reward value can refer to the numerical feedback signal calculated by the system based on the matching results of the current round of interaction feedback status (i.e., the comprehensive feature vector) and the reward rules. It is used to indicate the effectiveness of the current sleep-aid interactive content strategy in improving the user's status. It is usually a real value (such as the range of 0 to 1) or a discrete score.
[0121] The preset reward tuple can refer to the state-action-reward sequence stored in the reinforcement learning process, usually expressed as (S t , A t , R t ), which represents the current state, the content adjustment action taken, and the immediate reward value obtained. The system uses this tuple set for experience replay to optimize the strategy.
[0122] The policy network can be the core model in reinforcement learning algorithms, used to output the optimal action (i.e., the next content adjustment plan) based on the current state (such as user feedback). The network continuously optimizes content generation and adjustment strategies through accumulated experience training.
[0123] Multimodal feedback modeling is performed on the user's real-time interaction data. The real-time interaction data includes multiple types of passive physiological responses and active interaction behavior data of the user during the content playback process, such as: changes in the user's facial expression (image sequence captured by the camera), respiratory rate fluctuations (collected by audio sensors), dynamic characteristics of heart rate variability (through pulse wave analysis), and interaction behavior delays (such as gesture triggering or voice response delays). The system executes preset signal processing algorithms on the above multi-source data, such as using a CNN model to perform expression recognition on image sequences, frequency domain energy analysis on respiratory signals, wavelet transform processing on HRV sequences, etc., to extract feature dimensions such as emotional trends, rhythm smoothness, sympathetic / parasympathetic activity, etc., and finally constructs a unified comprehensive feature vector through a feature fusion mechanism (such as feature splicing and standardization) to express the user's physiological and behavioral response status under the stimulation of the current sleep-aiding content.
[0124] The system inputs this comprehensive feature vector into a pre-set reward rule for multi-dimensional evaluation and calculation. The reward rule is designed based on relaxation-oriented goals and includes a weighted scoring function for multiple physiological and behavioral indicators, such as whether the respiratory rate decreases, whether the heart rate variability (such as LF / HF) is stable, whether the facial expression transitions to a relaxed state, and whether the interaction is smooth. The system weights and sums these individual evaluation results to obtain an immediate reward value (e.g., a real value between 0 and 1) that reflects the current intervention effect. The higher this value, the more adapted the current sleep-aiding content is to the user's state and the better the intervention effect.
[0125] Then, the system considers the comprehensive feature vector corresponding to the current state (considered as state S t ), the content output method that has been adopted (Action A t ) and the immediate reward value just calculated (R t ) to form a triplet, which is added to the system's continuously updated set of preset reward tuples (i.e., the experience replay pool). The system employs deep reinforcement learning algorithms, such as DQN (Deep Q Network), using these reward tuples as training samples to iteratively update the policy network. Through backpropagation and loss function optimization, the policy network continuously improves its state-action mapping capabilities, enabling it to more accurately predict the optimal content adjustment strategy for similar user states.
[0126] Finally, in the new state judgment cycle, the system inputs the new comprehensive feature vector generated based on the current real-time interactive feedback into the updated strategy network and outputs the current optimal content adjustment strategy. This strategy will instruct the content playback engine to make targeted adjustments in meditation guidance, rhythm control, mindfulness element combination, etc., to achieve a closed-loop user state perception-reward feedback-strategy optimization process.
[0127] In this solution, by extracting multimodal features from real-time interaction data and combining reward rules with reinforcement learning strategy networks to dynamically optimize content output, adaptive adjustment of sleep-aiding content can be achieved, continuously improving intervention effects and individual matching, and enhancing users' immersive experience and relaxation guidance effects.
[0128] Figure 2 This is a flow chart of a method for generating interactive sleep-aiding content embedded with mindfulness meditation guidance provided by an embodiment of the present disclosure. The method may include the following steps: S201 , obtaining physiological data and user behavior data of a user, performing multimodal signal processing and feature extraction on the physiological data and user behavior data of the user, and obtaining a user state feature vector.
[0129] S202 : Inputting the user state feature vector into a preset state assessment model to obtain the user's relaxation probability, anxiety probability, and concentration probability.
[0130] The preset state assessment model can be a classification inference model built on a deep neural network (such as a multi-layer perceptron (MLP) or a lightweight convolutional neural network (CNN). Its input is the fused user state feature vector, and its output is the probability distribution of the user's current mental state across different dimensions. This model uses nonlinear modeling of multimodal features such as breathing, facial expression, and heart rate to output the confidence level of each mental state.
[0131] The relaxation probability may be a possibility that the user is currently in a state of physical and mental relaxation.
[0132] The anxiety probability may be the likelihood that the user will have an obvious nervous or stressful reaction.
[0133] The concentration probability can be the likelihood that the user is focused and has few external distractions.
[0134] The constructed user state feature vector is input into a pre-trained state assessment model to perform real-time state classification prediction. The user state feature vector is composed of multiple parameters after signal processing and feature extraction, including respiratory rate parameters, respiratory depth parameters, user emotional state classification results, and autonomic nervous system activity indicators. These characteristics reflect the user's multimodal physiological and behavioral state in the current time period.
[0135] This state assessment model is a multi-classification probabilistic output network built on deep learning architectures, such as lightweight multi-layer perceptrons (MLPs) or convolutional neural networks (CNNs). Pre-trained on multiple annotated datasets, the model is capable of nonlinear modeling and state label inference for high-dimensional input features. The model training data comes from real-user experiments and includes samples of relaxation, anxiety, and focus states labeled with corresponding physiological signals and professional questionnaire assessment results (such as STAI and MAAS scores). The cross-entropy loss function and iterative optimization using the Adam optimizer ensure good generalization.
[0136] During operation, the system inputs the user's state feature vector into the model. After hidden layer feature abstraction and weighting, the model finally uses the Softmax function at the output layer to normalize the state categories. The output is a three-dimensional probability vector, representing the predicted probability of the current state belonging to "relaxed," "anxious," and "focused," respectively. These three outputs are the probability of relaxation, anxiety, and focus, each fluctuating between 0 and 1, and the sum of the three is 1.
[0137] This three-dimensional probability output serves as a quantitative indicator of physical and mental state, and is used to drive the subsequent mindfulness guidance element matching and sleep-aiding content generation process, realizing a complete closed loop from user state perception to intervention content adjustment.
[0138] The training process of the preset state assessment model includes: During the model training phase, multi-source physiological and behavioral signals from real users are collected, including raw respiratory audio data, facial image sequences, motion signals, and pulse wave signals. Using a respiratory rate extraction algorithm, an expression recognition model, acceleration analysis, and HRV wavelet analysis, these raw data are extracted to construct structured feature values, such as respiratory rate and depth parameters, user emotional state classification results, and autonomic nervous system activity indicators. These are then uniformly constructed into standardized user state feature vectors. These features represent the user's comprehensive physical and mental state within a specific time window.
[0139] To generate the labeled data needed for supervised training, the system combines standard psychological assessment scales completed by subjects during the experiment (such as the STAI State Anxiety Inventory and the MAAS Mindfulness Awareness Scale) or expert-annotated emotion classification results to label each set of feature vectors as one of three psychological states: "relaxed," "anxious," or "focused." Thus, each piece of training data consists of a set of feature vectors and a state label.
[0140] To build the model, the system uses a lightweight neural network architecture, such as a multi-layer perceptron (MLP) or a streamlined convolutional neural network (CNN). It employs several hidden layers and activation functions (such as ReLU). The final layer outputs three neurons, one for each of the three states: "relaxed," "anxious," and "focused." The output layer uses a Softmax function for normalization, ensuring the output is a probability distribution of the three states, forming the model training target.
[0141] During training, the cross-entropy loss function is used as the objective function, and the model weights are iteratively updated through backpropagation and the Adam optimizer. To avoid overfitting, the dropout mechanism and training / validation set partitioning strategy are introduced. Model performance is evaluated and parameter optimization is performed in each iteration to ensure stable output on the validation set.
[0142] After training is complete, the system selects the parameter weight combination with the highest accuracy in the validation set as the final model, solidifying it as the preset parameters of the state assessment model for real-time user state recognition during the operational phase. The model has been proven to be highly sensitive to changes in the user's physiological and emotional signals, quickly outputting the probability of relaxation, anxiety, and focus, thereby providing accurate state information for dynamic matching of mindfulness elements and generation of sleep-enhancing content. This training method ensures the model's real-time performance, lightweightness, and personalized adaptability.
[0143] S203: Determine the user's relaxation level, anxiety level, and concentration level based on the relaxation probability, anxiety probability, and concentration probability and a preset state mapping rule, and determine a quantitative indicator of the user's current physical and mental state based on the relaxation level, anxiety level, and concentration level.
[0144] The preset state mapping rule can refer to a conversion rule that maps the probability values (relaxation, anxiety, and concentration) output by the model to corresponding level ranges. Typically, a threshold segmentation method is used. For example, the probability range [0, 1] is divided into multiple levels, such as: 0.0–0.3 → low (level 1), 0.3–0.7 → medium (level 2), and 0.7–1.0 → high (level 3).
[0145] The relaxation level may be a state level value obtained by discretizing the user's current relaxation probability according to a preset state mapping rule, and is used to represent the user's current relaxation level.
[0146] The anxiety level can be a numerical label obtained by the system based on the anxiety probability output by the model and discretized by the state mapping rule, which is used to quantify the user's current anxiety state.
[0147] The concentration level can be a discrete representation of the user's current concentration level. It is a level label formed by mapping the concentration probability value according to rules, and is used to evaluate the user's concentration level.
[0148] The physical and mental state metric represents the user's current psychological and physiological state, with adjustable label values across the three dimensions of relaxation, anxiety, and focus. This metric, composed of these three levels, forms a triplet (e.g., [relaxation = 3, anxiety = 1, focus = 2]), guiding the decision-making logic of subsequent modules like mindfulness element matching and content pacing. It not only reflects the user's overall physical and mental state but also supports dynamic updates and feedback adjustments, serving as a pivotal variable in the closed loop of "content generation, user feedback, and model adjustment."
[0149] After obtaining the user's relaxation, anxiety, and concentration probabilities, each probability value is converted into a corresponding discrete level value according to a set of pre-defined state mapping rules, forming the user's current relaxation, anxiety, and concentration levels. These state mapping rules are designed based on a threshold segmentation method, dividing the probability interval [0, 1] for each psychological state into multiple level intervals. For example, when the relaxation probability is between 0.0–0.3, it is mapped to relaxation level 1 (low relaxation); when it is between 0.3–0.7, it is mapped to level 2 (medium relaxation); and when it is greater than 0.7, it is mapped to level 3 (high relaxation). The same logic is used to determine the level of anxiety and concentration probabilities, mapping them to anxiety and concentration levels, respectively. The numerical ranges for these values can be adjusted based on experimental data to match the confidence distribution characteristics of the model output.
[0150] After completing the three-dimensional level conversion, the system combines these three level values into a three-dimensional vector, which serves as a quantitative indicator of the user's physical and mental state, for example, [Relaxation = 2, Anxiety = 1, Focus = 3]. This metric is a concise, structured representation of the user's current comprehensive mental state. It is interpretable and usable in real time, and can directly drive the subsequent generation of mindfulness meditation content and the regulation of playback rhythm.
[0151] S204 , dynamically generating content based on the physical and mental state quantitative indicators and a preset mindfulness element template library to obtain sleep-aiding interactive content containing naturally embedded mindfulness meditation elements.
[0152] S205: Acquire real-time user interaction data on the sleep-aid interactive content, perform reinforcement learning strategy optimization based on the real-time user interaction data on the sleep-aid interactive content, determine an optimal content adjustment strategy, and adjust playback parameters and / or mindfulness element combinations of the sleep-aid interactive content in real time based on the optimal content adjustment strategy.
[0153] In this embodiment, by inputting the user state feature vector into a trained state assessment model and converting it into relaxation level, anxiety level and concentration level in combination with the state mapping rules, structured physical and mental state quantitative indicators can be generated in real time, achieving fine perception and classification expression of the user's psychological state, providing a scientific basis for personalized and dynamic mindfulness sleep aid content adjustment, and effectively improving the system response accuracy and intervention effect.
[0154] Figure 3 This is a schematic block diagram of a sleep-aiding interactive content generation system embedded with mindfulness meditation guidance provided by an embodiment of the present disclosure. The system includes: The data acquisition and processing module 301 is used to obtain the user's physiological data and user behavior data, perform multimodal signal processing and feature extraction on the user's physiological data and user behavior data, and obtain a user state feature vector; A user state evaluation module 302 is configured to perform real-time analysis based on the user state feature vector and a preset state evaluation model to determine a quantitative index of the user's current physical and mental state; The sleep-aiding interactive content generation module 303 is configured to dynamically generate content based on the physical and mental state quantitative indicators and a preset mindfulness element template library to obtain sleep-aiding interactive content containing naturally embedded mindfulness meditation elements; The interactive feedback collection module 304 is used to obtain real-time user interaction data on the sleep-aid interactive content, perform reinforcement learning strategy optimization based on the real-time user interaction data on the sleep-aid interactive content, determine the optimal content adjustment strategy, and adjust the playback parameters and / or mindfulness element combination of the sleep-aid interactive content in real time according to the optimal content adjustment strategy.
[0155] Figure 4 A schematic block diagram of an electronic device 400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0156] The electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a ROM 402 or a computer program loaded from a storage unit 408 into a RAM 403. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An I / O interface 405 is also connected to the bus 404.
[0157] Multiple components in the electronic device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0158] Computing unit 401 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 401 performs the various methods and processes described above, such as the method for generating interactive sleep-aiding content embedded with mindfulness meditation guidance. For example, in some embodiments, the method for generating interactive sleep-aiding content embedded with mindfulness meditation guidance can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto electronic device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by computing unit 401, one or more steps of the method for generating interactive sleep-aiding content embedded with mindfulness meditation guidance described above can be performed. Alternatively, in other embodiments, the computing unit 401 may be configured in any other appropriate manner (for example, by means of firmware) to execute the method for generating sleep-aiding interactive content embedded with mindfulness meditation guidance.
[0159] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0160] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0161] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0163] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0164] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0165] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0166] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for generating interactive sleep-aiding content embedded with mindfulness meditation guidance, characterized in that: The method comprises: Acquiring physiological data and user behavior data of the user, performing multimodal signal processing and feature extraction on the user physiological data and user behavior data to obtain a user state feature vector; Performing real-time analysis based on the user state feature vector and a preset state assessment model to determine the quantitative indicators of the user's current physical and mental state; Dynamically generate content based on the quantitative indicators of physical and mental states and a preset mindfulness element template library to obtain sleep-aiding interactive content that includes naturally embedded mindfulness meditation elements; Real-time user interaction data on the sleep-aid interactive content is obtained, reinforcement learning strategy optimization is performed based on the real-time user interaction data on the sleep-aid interactive content, an optimal content adjustment strategy is determined, and playback parameters and / or mindfulness element combinations of the sleep-aid interactive content are adjusted in real time based on the optimal content adjustment strategy.
2. The method according to claim 1, characterized in that in, Obtain user physiological data and user behavior data, including: Acquiring breathing audio data, performing bandpass filtering on the breathing audio data to obtain a processed audio signal, and determining the user's breathing frequency data based on the audio signal; Acquire a facial image sequence, perform key point detection and expression classification on the image sequence to obtain a classification result, and determine the user's facial expression data based on the classification result; Acquiring a motion signal, performing activity intensity analysis on the motion signal to obtain an analysis result, and determining the user's physical activity data based on the analysis result; Acquiring original pulse wave signal data, performing bandpass filtering on the original pulse wave signal data, extracting the pulse wave peak position, and calculating a continuous pulse interval sequence based on the pulse wave peak position; Performing time-domain statistical analysis and frequency-domain wavelet transform processing on the continuous pulse interval sequence to obtain heart rate variability data representing the activity state of the user's autonomic nervous system; Determining user physiological data based on the respiratory rate data and the heart rate variability data; User behavior data is determined based on the facial expression data and the body activity data.
3. The method according to claim 2, characterized in that in, Performing multimodal signal processing and feature extraction on the user physiological data and user behavior data to obtain a user state feature vector, including: Performing short-time Fourier transform processing on the respiratory frequency data to obtain respiratory time-frequency characteristics, and determining a respiratory frequency parameter and a respiratory depth parameter according to the respiratory time-frequency characteristics; Inputting the facial expression data into a pre-trained expression recognition model for convolutional neural network processing to obtain a user emotional state classification result; performing wavelet transform analysis on the heart rate variability data, calculating time-frequency domain characteristic parameters, and determining an autonomic nervous system activity index based on the time-frequency domain characteristic parameters; The respiratory frequency parameter, the respiratory depth parameter, the user emotional state classification result and the autonomic nervous system activity index are subjected to feature fusion to obtain a user state feature vector.
4. The method according to claim 2, characterized in that in, Performing wavelet transform analysis on the heart rate variability data to calculate time-frequency domain characteristic parameters includes: determining a continuous interval sequence based on the heart rate variability data, constructing a cumulative time axis based on the continuous interval sequence, and determining a time variable based on the cumulative time axis; A time domain signal is constructed according to a continuous interval sequence and a time variable, and the time domain signal and the time variable are input into a preset wavelet transform formula to obtain time-frequency domain representation data, and the time-frequency domain characteristic parameters are determined according to the time-frequency domain representation data.
5. The method according to claim 4, characterized in that in, The preset wavelet transform formula is: ; in, represents the data in the time-frequency domain; a is the preset scale parameter, which controls the expansion and contraction of the wavelet; b is the preset translation parameter, which controls the position of the wavelet on the time axis; is the time domain signal; is the analysis wavelet after the mother wavelet is stretched and translated, where represents the complex conjugate.
6. The method according to claim 1, characterized in that in, Perform real-time analysis based on the user state feature vector and a preset state assessment model to determine the user's current physical and mental state quantitative indicators, including: Input the user state feature vector into the preset state assessment model to obtain the user's relaxation probability, anxiety probability and concentration probability; The user's relaxation level, anxiety level, and concentration level are determined based on the relaxation probability, anxiety probability, and concentration probability and according to a preset state mapping rule, and a quantitative index of the user's current physical and mental state is determined based on the relaxation level, anxiety level, and concentration level.
7. The method according to claim 1, characterized in that in, Dynamic content generation is performed based on the quantitative indicators of physical and mental states and a preset mindfulness element template library to obtain sleep-aiding interactive content that contains naturally embedded mindfulness meditation elements, including: Matching the mindfulness guidance type based on the quantitative indicators of the physical and mental state and the element label data in the preset mindfulness element template library to obtain the currently adapted mindfulness meditation element combination; Based on the mindfulness meditation element combination and the node-branch structure data in the narrative structure template engine, element embedding position and rhythm control are performed to obtain a structured script template containing mindfulness element embedding rules; Audio data, image data, and text data are generated according to a structured script template, and multimodal content synthesis is performed based on the audio data, image data, and text data to obtain sleep-aiding interactive content that includes naturally embedded mindfulness meditation elements.
8. The method according to claim 1, characterized in that in, Reinforcement learning strategy optimization is performed based on real-time user interaction data on the sleep-aid interactive content to determine the optimal content adjustment strategy, including: Perform multimodal feature extraction on the real-time interaction data to obtain a comprehensive feature vector, and perform multi-dimensional evaluation calculation based on the comprehensive feature vector and preset reward rules to obtain an instant reward value for the sleep-aiding interactive content; Reinforcement learning strategy optimization is performed according to the instant reward value and a preset reward tuple, a strategy network for content decision-making is updated, and an optimal content adjustment strategy is determined according to the strategy network.
9. A sleep-aiding interactive content generation system embedded with mindfulness meditation guidance, used to execute the method according to any one of claims 1 to 8, characterized in that: The system comprises: A data acquisition and processing module is used to obtain the user's physiological data and user behavior data, perform multimodal signal processing and feature extraction on the user's physiological data and user behavior data, and obtain a user state feature vector; A user state evaluation module is used to perform real-time analysis based on the user state feature vector and a preset state evaluation model to determine the quantitative indicators of the user's current physical and mental state; A sleep-aiding interactive content generation module is used to dynamically generate content based on the quantitative indicators of physical and mental states and a preset mindfulness element template library to obtain sleep-aiding interactive content containing naturally embedded mindfulness meditation elements; The interactive feedback collection module is used to obtain real-time interaction data of users on the sleep-aid interactive content, optimize the reinforcement learning strategy based on the real-time interaction data of users on the sleep-aid interactive content, determine the optimal content adjustment strategy, and adjust the playback parameters and / or mindfulness element combination of the sleep-aid interactive content in real time according to the optimal content adjustment strategy.
10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Interactive psychological counseling system adopting psychological data labeling modeling
CN112086169A
Multi-modal physiological signal-based personalized normal-feeling and pressure reduction system and multi-modal physiological signal-based personalized normal-feeling and pressure reduction method
CN119971244A
Emotional interaction regulation and control strategy generation method and device for rehabilitation training
CN120032791A
System
JP2025049368A
System for providing digital therapy service using meditation simulation
KR102508601B1