Sleep temperature regulation methods, systems and devices
Patent Information
- Application Number
- CN202610942704.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]本申请提供一种睡眠温度调整方法、系统及装置,用以解决由于睡眠数据获取需长时间实验,且人体不同部位或者不同阶段的温度调节并非独立,现有睡眠温度控制存在依赖人工或人工设计复杂的协同函数,导致生成的睡眠温度控制策略难以适应个体差异和动态变化的睡眠阶段需求的技术问题
[0065]本申请提供的睡眠温度调整方法、系统及装置,通过配置包含搜索中心向量、全局步长及用于表征各分区温控指令分区温控指令相关性的协方差矩阵的搜索空间,基于搜索空间采样生成温度方案种群,并结合决策周期中的多维生理特征信号确定各温度方案个体对应的奖励值,利用奖励值确定优选温度方案个体,再根据历史分区温控指令分区温控指令基于优选温度方案个体更新搜索空间并生成目标分区温控指令分区温控指令,能够在睡眠温度调节过程中同时考虑多分区温度之间的关联性以及生理反馈对方案优劣的表征,进而提升有限样本和反馈滞后条件下温度调节的协同优化能力、收敛效率以及对个体差异和睡眠阶段变化的动态适应性。
Smart Images

Figure CN122664633A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sleep technology, and in particular to a method, system and device for adjusting sleep temperature. Background Technology
[0002] In modern life, sleep quality is closely related to human health, and it is significantly affected by ambient temperature, individual physiological differences, and dynamic changes in sleep stages. For example, the body needs a lower ambient temperature in the early stages of falling asleep to promote melatonin secretion, while maintaining body temperature stability is necessary during the REM sleep stage. In addition, there is a synergistic relationship between the temperature needs of different body parts (such as the head, torso, and feet); for example, the combination of a cool head and warm feet can effectively induce deep sleep.
[0003] In existing technologies, sleep temperature control typically relies on fixed rule-based methods, adjusting the ambient temperature through preset temperature thresholds or empirical formulas. This reliance on manually defined rules fails to dynamically adapt to individual user differences or changes in sleep stages, and it struggles to handle the collaborative optimization of multi-dimensional temperature parameters. Furthermore, sleep temperature control can be achieved through neural network strategies, learning the optimal temperature strategy through trial and error. However, acquiring sleep data requires extensive experimentation, and temperature regulation in different parts of the body or at different stages is not independent. Existing neural network strategies rely on manually designed complex collaborative functions, making it impossible to generate real-time sleep temperature control strategies.
[0004] Therefore, since acquiring sleep data requires long-term experiments, and the temperature regulation of different parts of the human body or different stages is not independent, existing sleep temperature control relies on artificial or artificially designed complex cooperative functions, making it difficult for the generated sleep temperature control strategies to adapt to individual differences and dynamically changing sleep stage needs. Summary of the Invention
[0005] This application provides a method, system, and device for adjusting sleep temperature, in order to solve the technical problem that existing sleep temperature control relies on manual or artificially designed complex cooperative functions, which is difficult to adapt to the needs of individual differences and dynamically changing sleep stages because sleep data acquisition requires long-term experiments and the temperature regulation of different parts of the human body or different stages is not independent.
[0006] In a first aspect, this application provides a method for adjusting sleep temperature, including:
[0007] Obtain multidimensional physiological feature signals for the current decision-making cycle. Based on the multidimensional physiological feature signals, obtain the reward value corresponding to each temperature scheme individual. Each temperature scheme individual is a decision vector containing multiple temperature control feature dimensions.
[0008] Based on the individual temperature schemes and their corresponding reward values in the current decision-making cycle, the individual temperature schemes for the next decision-making cycle are determined using an adaptive covariance matrix evolution strategy.
[0009] Based on the temperature scheme, the target zone temperature control command is obtained to drive the temperature control equipment to execute the target zone temperature control command.
[0010] This sleep temperature adjustment method addresses the shortcomings of traditional sleep temperature regulation technologies, such as insufficient multi-dimensional coordination, low sample efficiency, delayed attribution errors, and poor dynamic adaptability. It significantly improves temperature control by automatically capturing the correlation between zone temperatures using a covariance matrix. Second-order optimization characteristics enable rapid convergence with limited samples, reducing user waiting time. Future physiological improvements are factored into the decision-making time, eliminating policy bias caused by erroneous attribution. A sliding window mechanism and context parameter correction allow the system to adapt to changes in the user's physiological state in real time. By optimizing the mapping function between context parameters and temperature commands, it automatically learns the user's personalized temperature sensitivity, improving policy generalization and achieving personalized, real-time sleep temperature regulation, thus significantly improving the user's sleep quality.
[0011] In one possible embodiment, an adaptive covariance matrix evolution strategy is used to determine the individual temperature schemes for the next decision cycle, including:
[0012] Configure the search space, which includes the search center vector, global step size, and covariance matrix. The covariance matrix is configured to characterize the correlation between various temperature control feature dimensions.
[0013] Using the reward value corresponding to each temperature scheme, determine the optimal temperature scheme and update the search space based on the optimal temperature scheme.
[0014] Based on the search space, a population consisting of individuals with multiple temperature schemes is generated by sampling.
[0015] By constructing an iteratively updatable search space using the search center vector, global step size, and covariance matrix, and forming a closed-loop feedback loop of reward values with multi-dimensional physiological feature signals, and then combining the optimal temperature scheme with historical zone temperature control commands for joint updates, joint search and continuous optimization of multi-zone temperature can be achieved under conditions of limited samples and feedback lag. Furthermore, modeling the correlation between temperature control commands in each zone using the covariance matrix allows the system to learn the relationships between areas such as the head, torso, and feet, avoiding the problem that traditional independent zone control cannot reflect synergistic effects. The introduction of reward values for real sleep physiological feedback gives the search process individual dynamic adaptability, thereby improving the adaptability of target zone temperature control commands to individual user differences, environmental fluctuations, and sleep stage switching, ultimately achieving stable, precise, and coordinated temperature regulation throughout the entire sleep process.
[0016] In one possible embodiment, the temperature scheme individual is defined as a parameter vector of a mapping function between context parameters and zone temperature control commands;
[0017] The method also includes: obtaining the current context parameters in each decision cycle, and calculating the target partition temperature control command under the current context parameters based on the current context parameters and the mapping function.
[0018] By defining individual temperature schemes as parameter vectors of a mapping function, the search object is transformed from a single fixed temperature value into learnable policy parameters, thereby improving the ability to express complex contexts. Since the target partition temperature control command is calculated based on the current context parameters within each decision cycle, the system can adapt to environmental fluctuations and individual differences in real time during sleep, reducing biases caused by fixed strategies and improving the stability and convergence efficiency of partitioned collaborative temperature control.
[0019] In one possible embodiment, the method further includes:
[0020] Obtain the zone temperature control command and corresponding reward value for the current sleep cycle. By introducing a sliding window mechanism, construct a history buffer and store the zone temperature control command and corresponding reward value for the current sleep cycle into the history buffer.
[0021] Within each decision cycle, execute a temperature scheme individual, obtain the reward value corresponding to the temperature scheme individual, store the latest executed temperature scheme individual and its corresponding reward value into the history buffer, and remove the oldest temperature scheme individual from the history buffer.
[0022] The covariance matrix is updated based on the individual temperature schemes and their corresponding reward values within the historical buffer.
[0023] The buffer maintenance mechanism ensures that search space updates continuously rely on recent sleep feedback, preventing outdated samples from interfering with current individual differences and environmental changes. The covariance matrix thus more accurately reflects the correlation between temperature control commands in each zone, allowing subsequent sampling to focus on high-yield areas and improving the response speed to sleep stage transitions and circadian rhythm shifts. Consequently, the system maintains good convergence and stability under limited sample conditions, enhancing the optimization accuracy and individual adaptability of zone-based coordinated temperature regulation.
[0024] In one possible embodiment, the covariance matrix is updated based on the temperature scheme individuals within the historical buffer and their corresponding reward values, including:
[0025] Based on the chronological order, reward values within the historical buffer are assigned exponentially decaying weights, and the covariance matrix is updated according to these decaying weights.
[0026] Alternatively, the environmental memory corresponding to the partition temperature control instructions and reward values in the historical buffer can be obtained, the importance weights can be calculated and corrected, and the covariance matrix can be updated using the corrected importance weights.
[0027] In this way, the covariance matrix can continuously reflect the correlation changes between temperature control commands in different zones, and incorporate recent effective feedback or representative environmental memories into the update process, so that the sampling distribution of the search space gradually converges towards more effective temperature control combinations. Since the update process takes into account both time decay and environmental importance correction, it can reduce interference from old samples, alleviate attribution bias caused by feedback lag, and improve the stability and individual adaptability of zone temperature control collaborative optimization.
[0028] In one possible embodiment, calculating and correcting the importance weights further includes:
[0029] Based on the environmental memory corresponding to the partition temperature control instructions and reward values in the historical buffer, calculate the current distribution probability density and the historical distribution probability density;
[0030] The importance weights are calculated based on the current probability density and the historical probability density.
[0031] Based on the reward value corresponding to each temperature scheme, the importance weight is adjusted according to a preset threshold.
[0032] By incorporating environmental similarity and reward feedback into the weighting calculation process, the system reduces bias caused by relying solely on frequency statistics and suppresses the interference of low-quality samples on covariance updates. Consequently, the search direction for zoned temperature control commands becomes more stable across different sleep environments, improving the generation accuracy of target zoned temperature control commands and thus enhancing the individual adaptability and continuous optimization capability of sleep temperature regulation.
[0033] In one possible embodiment, the preferred temperature scheme is determined using the reward value corresponding to each individual temperature scheme, including:
[0034] The sleep scores of each sleep stage in the current sleep cycle are combined with the pre-acquired historical average physiological expectation values to obtain the improvement increment as the reward value.
[0035] Using a preset time response function, the improvement increment is attributed to time series. By converting the instantaneous improvement increment at future moments to the temperature control decision moment, the optimal temperature scheme can be determined by ranking the improvement increments.
[0036] This reward construction method combines changes in sleep quality during different sleep stages with historical expected baselines and uses a time response function to achieve delayed effect attribution, ensuring that the evaluation of the temperature scheme aligns with the actual timing of the temperature regulation effect. This reduces attribution bias caused by delayed physiological feedback, improves the accuracy of individual selection for optimal temperature schemes, and consequently enhances the stability of subsequent search space updates and the individual adaptability of temperature regulation.
[0037] In one possible embodiment, the individual temperature scheme is a spatial decision vector, which is obtained by combining the zoned temperature control commands of each temperature control device; the method further includes:
[0038] The search space is updated and the target partition temperature control command is generated by using the reward value corresponding to each spatial decision vector.
[0039] By abstracting the zone temperature control commands of each temperature control device into individual spatial decision vectors and updating the search space based on the corresponding reward values, the controller can continuously approach the optimal zone temperature control combination under limited sample conditions, enhance the modeling ability of multi-region correlation, and improve the adaptability to changes in sleep stages and individual differences, thereby improving the stability and adjustment accuracy of zone temperature control.
[0040] In one possible embodiment, the search space is updated based on the preferred temperature scheme individual, including:
[0041] Calculate the statistical mean of the individual optimal temperature schemes across each temperature control feature dimension, and update the search center vector to the statistical mean;
[0042] Based on the individual preferred temperature scheme, calculate the standardized deviation vector between the individual preferred temperature scheme and the current search center point;
[0043] The evolutionary path of the covariance matrix is maintained and updated based on the global step size and the standardized deviation vector.
[0044] The first update term is obtained by combining the vector cross product of the evolutionary paths with the evolutionary path direction of the individuals in the optimal temperature scheme;
[0045] By utilizing the spatial distribution of individuals with the preferred temperature scheme in the current sleep cycle, the empirical covariance is calculated, resulting in the second update term. The covariance matrix, the first update term, and the second update term are then weighted and fused to adjust the axial length and rotation angle of the search space.
[0046] Through the above update method, the search space can simultaneously absorb historical directional accumulation, recent high-quality sample trends, and spatial distribution information of the current sleep cycle, enabling multi-zone temperature control search to maintain good directional consistency and adaptability under limited sample conditions, thereby improving the convergence speed, stability, and individual adaptability of coordinated temperature regulation.
[0047] Secondly, this application provides a method for adjusting sleep temperature, including:
[0048] During multiple consecutive decision cycles, the corresponding target zone temperature control command is output to the temperature control equipment;
[0049] The temperature control commands for each target zone generated within multiple consecutive decision cycles are used as a multidimensional temperature scheme vector, and the corresponding sample covariance matrix conforms to the following characteristics:
[0050] (i) The sample covariance matrix is a strictly positive definite matrix, and the ratio of its minimum eigenvalue to its maximum eigenvalue is greater than a preset proportion threshold. The preset proportion threshold is configured to be greater than the upper limit of the random variation ratio caused by hardware measurement and execution noise of the temperature control system, and the preset proportion threshold is configured to be 5%.
[0051] (ii) There exists at least one non-zero off-diagonal element in the sample covariance matrix;
[0052] (iii) The eigenvector corresponding to the largest eigenvalue in the sample covariance matrix has components that are not zero on the temperature setpoint coordinate axes of at least two temperature control zones.
[0053] The sample covariance matrix is constructed by sending target zone temperature control commands to each temperature control device. The dynamic update mechanism of the covariance matrix automatically learns the correlation between zone temperatures, breaking through the limitations of traditional independent adjustment, significantly improving the effect of sleep quality regulation, and ensuring that the zone temperature control strategy meets the user's personalized physiological needs.
[0054] In one possible embodiment, the temperature scheme is a time decision vector, which is obtained by combining the zoned temperature control commands for each sleep stage in the sleep cycle; the method further includes:
[0055] The search space is updated and the target partition temperature control command is generated by using the reward values corresponding to each time decision vector.
[0056] By incorporating the sleep stage dimension into the decision space, the temperature control changes over time are matched with the sleep process, and the search direction is continuously corrected through reward feedback. This implementation improves the adaptability of staged temperature regulation during sleep, reduces deviations caused by fixed curve control, enhances the responsiveness to individual differences and stage switching, and improves the convergence efficiency and regulation stability of temperature control commands for the target zone.
[0057] Thirdly, this application provides a sleep temperature regulation system, including: a zoned temperature control device, a biofeedback sensor, and a controller.
[0058] The zoned temperature control device includes at least one temperature control unit for performing thermal regulation on the user's sleeping area;
[0059] Biofeedback sensing devices are used to acquire multidimensional physiological characteristic signals of users in real time;
[0060] The controller is electrically connected to the zone temperature control device and the biofeedback sensing device. The controller is used to receive multi-dimensional physiological characteristic signals, execute the sleep temperature adjustment method as described in any of the first aspects, generate a target zone temperature control command and send it to the zone temperature control device, thereby controlling the zone temperature control device to execute the target zone temperature control command to adjust the thermal parameters of different parts of the user's body.
[0061] Fourthly, this application provides a sleep temperature adjustment device, comprising:
[0062] The temperature scheme individual acquisition module is used to acquire multi-dimensional physiological feature signals of the current decision cycle. Based on the multi-dimensional physiological feature signals, the reward value corresponding to each temperature scheme individual is acquired. The temperature scheme individual is a decision vector containing multiple temperature control feature dimensions.
[0063] The temperature scheme individual determination module is used to determine the temperature scheme individuals for the next decision period based on the temperature scheme individuals and their corresponding reward values in the current decision period, using an adaptive covariance matrix evolution strategy.
[0064] The target zone temperature control command acquisition module is used to obtain the target zone temperature control command based on the individual temperature scheme, so as to drive the temperature control device to execute the target zone temperature control command.
[0065] The sleep temperature adjustment method, system, and apparatus provided in this application configure a search space including a search center vector, a global step size, and a covariance matrix used to characterize the correlation between temperature control commands for each zone. A population of temperature schemes is generated based on sampling from the search space. The reward value corresponding to each individual temperature scheme is determined by combining multi-dimensional physiological characteristic signals during the decision-making cycle. The optimal temperature scheme is determined using the reward value. Then, the search space is updated based on the optimal temperature scheme based on historical temperature control commands for each zone, and target temperature control commands for each zone are generated. This allows for simultaneous consideration of the correlation between temperatures in multiple zones and the characterization of scheme quality by physiological feedback during sleep temperature regulation. This improves the collaborative optimization capability, convergence efficiency, and dynamic adaptability to individual differences and sleep stage changes under limited sample and feedback lag conditions. Attached Figure Description
[0066] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0067] Figure 1 A schematic flowchart illustrating a sleep temperature adjustment method provided in an embodiment of this application;
[0068] Figure 2 This is a schematic diagram of a sleep temperature regulation system provided in an embodiment of this application;
[0069] Figure 3 This is a schematic diagram of the structure of a sleep temperature adjustment device provided in an embodiment of this application;
[0070] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0071] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0072] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the embodiments below.
[0073] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0074] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0075] It should be noted that the sleep temperature adjustment method, system and device provided in this application can be used in the field of sleep monitoring technology, or in any field other than sleep monitoring technology. The application field of the sleep temperature adjustment method, system and device in this application is not limited.
[0076] Intelligent sleep regulation technology is typically applied in sleep devices with zoned temperature control capabilities, such as smart mattresses, smart pillows, bedding temperature control modules, and their associated control terminals. In these scenarios, the system usually manages temperature throughout the user's entire nighttime sleep process, adjusting the thermal environment for different areas of the head, torso, feet, or bed surface by setting temperatures for each stage of sleep (falling asleep, light sleep, deep sleep, and REM sleep). To achieve a more stable temperature control effect, the system generally works in conjunction with physiological signal acquisition devices and environmental sensing devices, such as collecting information on electroencephalograms (EEG), heart rate variability, respiratory rhythm, body movement, room temperature, and humidity, and transmitting the relevant data to the controller. The controller then generates temperature control commands based on preset strategies or feedback results, driving the corresponding zoned temperature control devices to execute, thus forming a closed-loop application scenario where sleep state perception, temperature control decision-making, and device execution work together.
[0077] In existing technologies, common solutions include using preset temperature curves, time-segmented rule-based control, PID control, and closed-loop temperature regulation based on single physiological feedback to adjust temperature during sleep. Preset temperature curves typically divide the night into several time periods based on general population experience, providing fixed temperature values or temperature fluctuation trends for each period. Rule-based control triggers temperature increases or decreases based on threshold conditions. PID control continuously adjusts the temperature around a target temperature or a single error value. Some feedback-based solutions adjust the temperature based on the user's subjective sensation of hot or cold, or on limited indicators such as heart rate and body movement. While these solutions can achieve automatic temperature regulation to some extent, their basic assumptions are that each temperature control zone is independent, making it difficult to accurately characterize the correlation between temperatures in different parts of the body, such as the head, torso, and feet, and also failing to reflect the synergistic effects of different combinations on sleep latency, deep sleep maintenance, and micro-arousal inhibition. Meanwhile, sleep optimization itself is characterized by expensive samples and slow feedback. A single night's sleep often yields only a limited number of effective samples, and the physiological improvements following temperature adjustment often exhibit significant lag. This leads to existing solutions being prone to slow convergence, feedback attribution bias, and distorted search direction during the optimization process. Furthermore, when users experience changes in daytime activity levels, drink alcohol before bed, experience fluctuations in environmental conditions, or switch sleep stages, fixed-parameter or low-dimensional feedback schemes often fail to adapt promptly. Consequently, temperature regulation strategies struggle to accommodate individual differences and dynamic changes, limiting both overall adjustment accuracy and stability.
[0078] In view of this, how to improve the collaborative optimization capability and individual dynamic adaptability of sleep temperature regulation under limited sample size and feedback lag conditions has become an urgent technical problem to be solved. To solve the above technical problem, this application provides a sleep temperature adjustment method, system, and device. In the intelligent sleep regulation system, the controller can work in conjunction with the zone temperature control device and the physiological signal acquisition device. When entering the corresponding decision cycle, a search space including the search center vector, global step size, and covariance matrix is first configured, and multiple temperature scheme individuals are generated by sampling using this search space; then, multidimensional physiological feature signals in the decision cycle are collected to obtain the reward value corresponding to each temperature scheme individual; then, the optimal temperature scheme individual is determined based on the reward value, and the search space is updated by combining historical zone temperature control instructions to generate target zone temperature control instructions to drive the temperature control device to perform the corresponding zone temperature control operation. By using the covariance matrix to characterize the correlation of temperature control commands in each zone, this technical approach enables joint search and continuous optimization of multi-zone temperatures during sleep temperature regulation, thereby providing a foundation for improving the multi-zone coordinated temperature regulation effect and enhancing the adaptability to individual differences, aiming to solve the above-mentioned technical problems of existing technologies.
[0079] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0080] Figure 1 This is a flowchart illustrating a sleep temperature adjustment method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes:
[0081] S101: Obtain multidimensional physiological feature signals for the current decision-making cycle, and based on the multidimensional physiological feature signals, obtain the reward value corresponding to each temperature scheme for each individual.
[0082] In this embodiment of the application, the individual temperature scheme is a decision vector containing multiple temperature control feature dimensions.
[0083] In this step, multidimensional physiological characteristic signals are used to reflect the user's sleep state and physiological response, serving as a direct basis for evaluating the effectiveness of temperature control. These signals may include one or more of the following: electroencephalogram (EEG), heart rate variability, respiratory rhythm, body movement signals, skin temperature, peripheral perfusion index, frequency of turning over, number of micro-arousals, and derivative indicators related to sleep thermal comfort. The reward value is an evaluation indicator used to quantify the individual temperature control effectiveness of candidate temperature schemes. A higher value indicates that the corresponding candidate temperature scheme is more beneficial to the current sleep goals, such as shortening sleep latency, increasing the probability of maintaining deep sleep, reducing the frequency of nighttime micro-arousals, or reducing the amplitude of autonomic nervous system fluctuations. The decision cycle corresponds to one execution and feedback collection process for a candidate temperature scheme, and its length can be set from several minutes to tens of minutes depending on the sleep stage and signal stability.
[0084] In practice, the controller acquires raw signal data by communicating with physiological signal acquisition devices. These devices can be installed in headband-mounted EEG acquisition devices, wristband-mounted heart rate sensors, mattress-embedded pressure and respiration sensors, non-contact millimeter-wave respiratory monitoring devices, infrared body surface temperature sensors, and ambient temperature and humidity sensors. During the execution of each individual temperature scheme, the controller performs time alignment, artifact removal, filtering, and segmentation processing on the acquired raw data. For example, bandpass filtering can be performed on EEG signals to extract features such as slow wave power and sleep spindle density; time-domain and frequency-domain indicators can be extracted from heart rate signals to form heart rate variability features; respiratory signals can extract respiratory rate, rhythm stability, and the number of apnea events; and body movement signals can extract the number of turns per unit time, the amplitude of local pressure changes, and the duration of movement. Ambient temperature, humidity, and real-time bed surface temperature can also be used as auxiliary contextual information in reward value calculation to reduce the interference of external fluctuations on the attribution results.
[0085] The reward value can be calculated using a weighted combination method or a mapping model. In one possible embodiment, a scoring model of the following form can be constructed: R = w1·f1 + w2·f2 + w3·f3 - w4·f4, where R represents the reward value, f1 represents the score of deep sleep-related features, f2 represents the score of autonomic nervous system stability, f3 represents the score of respiratory rhythm stability, f4 represents the micro-arousal or body movement penalty item, and w1 to w4 represent the corresponding weights. Each weight can be set according to the user's individual goals, the current sleep stage, and the historical model training results. For example, the weight of indicators related to shortening sleep latency can be increased during the sleep onset stage, and the weight of indicators related to micro-arousal inhibition and thermoregulation stability can be increased in the second half of the night. If a mapping model is used, multidimensional physiological feature signals can be input into a trained scoring network, regression model, or Bayesian estimation model, and the model outputs a reward value to represent the individual's comprehensive contribution to improving sleep quality under the current temperature scheme.
[0086] Considering the significant lag in sleep feedback, this step can also incorporate time decay attribution and baseline correction mechanisms. The time decay attribution mechanism assigns time-weighted physiological improvements observed at the end of the decision cycle to the currently executing temperature strategy individual, reducing interference from preceding control actions on the current reward value. The baseline correction mechanism compares the current physiological characteristics with the user's resting baseline, historical mean values from the same period, or adjacent unadjusted window data to obtain a relative improvement rather than an absolute value. For example, if a user's heart rate variability is consistently low, directly using absolute values may underestimate the temperature regulation effect; baseline correction can more accurately reflect the gain brought by the current temperature control strategy. In cases of insufficient signal quality, the controller can generate a confidence coefficient based on the artifact ratio, sensor connection status, and data integrity, reducing or marking reward values as low-confidence samples to prevent abnormal data from misleading subsequent search updates.
[0087] In terms of specific execution results, the controller generates a corresponding record for each temperature scheme individual in the population. This record includes at least the individual identifier, execution start and end times, multidimensional physiological feature vectors, environmental context parameters, reward value, and data reliability label. After evaluating all individuals, the system obtains a one-to-one correspondence between temperature scheme individuals and reward values. Based on the above analysis, by collecting multidimensional physiological feature signals during the decision-making cycle and calculating reward values accordingly, the system can directly incorporate sleep physiological feedback into the temperature control optimization closed loop. This allows the quality of temperature schemes to no longer depend on a single temperature error or fixed empirical rules, but rather on a comparison based on a comprehensive quantitative result of the actual sleep state. This improves the effectiveness of the search direction and the accuracy of attribution under conditions of limited samples and feedback lag. It should be understood that the above example is for demonstration purposes only and is not a limitation.
[0088] S102: Based on the individual temperature schemes and corresponding reward values in the current decision-making cycle, determine the individual temperature schemes for the next decision-making cycle using an adaptive covariance matrix evolution strategy.
[0089] S103: Based on the individual temperature scheme, obtain the target zone temperature control command to drive the temperature control device to execute the target zone temperature control command.
[0090] In one implementation scenario, temperature regulation in different parts of the human body or at different stages is not independent. Traditional control or simple reinforcement learning usually assumes that variables are independent, or requires the design of extremely complex cooperative functions. However, human thermoregulation is highly interconnected, and isolating individual parts cannot achieve optimal results. For example, during the sleep onset stage, DPG (hand and foot temperature minus trunk temperature) is a physiological indicator predicting sleep onset speed (SOL). When blood vessels in the feet dilate, core heat rapidly diffuses to the limbs through blood circulation, inducing a drop in core body temperature (CBT), which is a key biological signal for initiating sleep.
[0091] Meanwhile, acquiring sleep data is extremely costly, typically generating only one valid sample per night. Deep reinforcement learning usually requires a large number of interactions to converge, and under the constraint of "one sample per night," it takes a very long time to complete personalized evolution. The cross-night iteration cycle is too long. Gathering a population λ usually requires λ days or λ sleep cycles, making it impossible to track the user's physiological rhythm drift in real time.
[0092] Therefore, this application provides a sleep temperature control method based on an adaptive covariance matrix evolution strategy, applying the CMA-ES algorithm to the sleep temperature adjustment scenario and capturing the correlation between dimensions through the covariance matrix C. For example, it can spontaneously learn physiological relationships in different dimensions. This co-evolutionary capability enables spontaneous modeling of the user's personalized physiological response manifold without the need for pre-defined physical formulas. The advantage of CMA-ES, as a population-based second-order optimization algorithm, is that it guides the search direction by learning the "shape" (matrix approximation) of the search space. It has extremely high convergence efficiency in small-sample, high-dimensional searches, enabling it to find the optimal temperature curve more quickly within a limited number of sleep samples. Specifically, the covariance matrix C of CMA-ES is used to capture the correlation between the three-dimensional vectors of "head, torso, and feet." If the "cool head + warm feet" combination strategy produces a high reward, the rank-μ update immediately rotates the eigenvectors of the covariance matrix (through the B matrix), causing the axis of the search distribution to point towards this specific cooperative direction. The golden ratio of "cool head + warm feet" is automatically discovered using matrix updates (Rank-μ updates). By defining a nonlinear kernel function, low weights are assigned during the initial stage of action execution, while peak weights are assigned during the thermal steady-state establishment period (20-30 minutes), thus accurately capturing the causal relationship between temperature-controlled actions and physiological feedback. A sliding window mechanism is used to establish a first-in, first-out historical buffer, with incremental updates performed immediately after acquiring a new sample each night, squeezing out older data and achieving daily parameter evolution.
[0093] Furthermore, the physiological baselines of the human body are completely different at different sleep stages. For example, the physiological baselines at 1 AM (the first N2) and 4 AM (the fourth N2) differ in that core body temperature (CBT) is usually at its lowest point at 4 AM, and heart rate and HRV baselines at this time show a significant natural drift compared to the early stages of sleep. This natural shift (periodic rhythm) over time is a huge background noise for the algorithm. If the four cycles of a night are ranked together directly, the algorithm will consider the parameters of the second half of the night to be better (or worse), simply because the body's physiological baseline has changed, not because of the effect of temperature control parameters. There is a 15-30 minute lag in heat conduction and metabolic response; the current physiological improvement may be due to an action taken half an hour earlier. This causal mismatch can cause the algorithm to fail.
[0094] The CMA-ES population (λ) evaluation is "vertically sliced" on the timeline instead of "horizontally accumulated" using a homogeneous background ranking strategy. The execution process includes: the system labels each 90-minute sample (e.g., Night_1_N2_Index_1, Night_1_N2_Index_2, etc.). A homogeneous buffer is constructed: the algorithm maintains four independent search spaces or training sets, corresponding to the 1st, 2nd, 3rd, and 4th N2 stages each night. Cross-night collection: on the first night, schemes Φ_1,1 and scores J_1,1 from the 1st N2 stage are collected. On the second night, schemes Φ_2,1 and scores J_2,1 from the 1st N2 stage of the second night are collected. This continues until λ samples belonging to the same "1st N2 stage" are collected. Under the same "3 N2" background, the interference source of physiological evolution is removed from the comparison formula.
[0095] Based on the above analysis, this application provides a sleep temperature adjustment method, including configuring a search space, which includes a search center vector, a global step size, and a covariance matrix. The covariance matrix is used to characterize the correlation of temperature control commands for each zone. Based on the search space, a population of temperature schemes is generated by sampling, and the population includes multiple temperature scheme individuals. Multidimensional physiological feature signals are collected during the decision-making cycle, and reward values corresponding to each temperature scheme individual are obtained based on the multidimensional physiological feature signals. Using the reward values corresponding to each temperature scheme individual, a preferred temperature scheme individual is determined. Based on historical zone temperature control commands and the preferred temperature scheme individuals, the search space is updated and a target zone temperature control command is generated to drive the temperature control device to execute the target zone temperature control command. In this embodiment, by constructing an iteratively updatable search space using the search center vector, global step size, and covariance matrix, and forming a closed-loop feedback of reward values using multidimensional physiological feature signals, and then combining the preferred temperature scheme individuals with historical zone temperature control commands for joint updates, joint search and continuous optimization of multi-zone temperatures can be achieved under conditions of limited samples and feedback lag. Furthermore, the modeling of the correlation between temperature control commands in each zone by the covariance matrix enables the system to learn the relationships between areas such as the head, torso, and feet, avoiding the problem that traditional independent zone control is difficult to reflect synergistic effects. The introduction of reward values for real sleep physiological feedback enables the search process to have individual dynamic adaptability, thereby improving the adaptability of target zone temperature control commands to individual user differences, environmental fluctuations, and sleep stage switching, ultimately achieving stable, precise, and coordinated temperature regulation throughout the entire sleep process.
[0096] Specifically, datasets are set up for each of the four sleep stages, and the zoned temperature control command Φ_target is used for each of the four stages (N1, N2, N3, REM) in the current cycle. This embodiment takes N3 as an example, where each individual x is a 3-dimensional vector. Optionally, it can also be combined with the air conditioning temperature to form a four-dimensional vector.
[0097] Expert Initialization (Seed Assignment): Based on the expert manual for sleep medicine research, an independent initial target heat exchange rate Φ_0 is set for each sleep stage. Before the system is first started, the initial heat flux density Φ_initial(S) is set in the expert database. Four empty training sets are initialized, and subsequent calculations are performed separately for the four sleep stages.
[0098] CMA-ES maintains a multidimensional Gaussian distribution to describe the region where the optimal solution may exist, automatically stretching and shifting within the search space until it encompasses the true optimal reward point. CMA-ES transforms the optimization problem into an evolution problem involving a multivariate normal distribution. In the g-th generation (i.e., the g-th decision period), the search distribution is defined as:
[0099]
[0100] In the formula, The current search center (mean) is used to characterize what is currently considered the most likely personalized optimal temperature point; This is the global step size, used to control the granularity of the search. The covariance matrix is used to determine the shape of the search distribution and capture the correlation between temperature control actions of different body parts; This represents the number of individuals in the population.
[0101] (1) Sample generation (Mutation): The system will generate samples all at once under the same distribution parameters. Individual candidate temperature schemes, then... The plan is executed in each of the N3 stages, waiting for this... After all the options have been evaluated, the covariance is updated, which constitutes a decision cycle.
[0102] Covariance matrix decomposition: Here, D is a diagonal matrix used to store the square roots of the eigenvalues of the covariance matrix C. It determines the scaling ratio of the search ellipse along each principal axis. For example, if the variance of foot temperature is large, the ellipse will be stretched along that dimension. B is an orthogonal matrix whose column vectors are the eigenvectors of the covariance matrix C, determining the rotation angle of the search ellipse. This captures synergistic effects such as "cool head + warm feet," allowing the search to move beyond a single coordinate axis and proceed along the oblique dimensions of physiological correlations. The covariance matrix C is initially set as an identity matrix I (at which point B=I, D=I), used to characterize the initial uncorrelated temperature control in each zone. Random vectors are drawn from a standard normal distribution. The actual heat flux density obtained by transformation is: (Where B and D originate from the eigenvalue decomposition of C).
[0103] (2) Second step: Evaluation and ranking (Selection & Recombination) - Homogeneous background ranking: Calculate the sleep reward value J_i corresponding to the partition temperature control command Φ_i. The CMA-ES population ( Evaluation is performed by "vertical slicing" along the timeline. The system sequentially labels each N3 sample for a given night (e.g., Night_1_N3_Index_1, Night_1_N3_Index_2, etc.). A homogeneous buffer is constructed: four independent search spaces or training sets are maintained, corresponding to the 1st, 2nd, 3rd, and 4th N3 stages each night. On the first night, schemes Φ_1,1 and scores J_1,1 for the 1st N2 stage are collected. On the second night, schemes Φ_2,1 and scores J_2,1 for the 1st N3 stage of the second night are collected. This continues until λ samples belonging to the same "1st N3 stage" are gathered. Intra-group competitive optimization: ranking is only performed among these λ samples with the same "time baseline." Under the same "3rd N3" context, all samples experience similar sleep stress and body temperature drop rates. By comparing samples at the same stage and ranking, the interference of physiological evolution is eliminated from the comparison formula, achieving competition among samples at the same ranking and stage: under the same "third N3" background, all samples experience similar sleep pressure and body temperature drop rates. The optimal sample selected through ranking then truly benefits from the improved reward score J derived from the temperature control parameters. The parallel processes evolve independently, enabling the adoption of different temperature control strategies for different sleep stages within the same night.
[0104] Selection and weighting: Sort all individuals in descending order of reward value J. Select the top μ performers (usually μ-λ / 2), and update the center: new center. It is the weighted average of these preferred temperature schemes:
[0105]
[0106] in, As a weight, the higher the ranking, the greater the weight.
[0107] Optionally, data from different time periods can be grouped together using stage expectation correction, while normalizing the reward score itself. For example, the system pre-calculates or statistically analyzes the "historical average physiological expectation" for 1 AM and 4 AM, and then subtracts the baseline expectation when calculating the current score. The calculation logic is: r(t) = r(t) - r_stage(t). The algorithm no longer evaluates the absolute sleep score, but rather the "incremental improvement relative to the natural state during that time period." This essentially eliminates the unfairness caused by the time context. The ranking mechanism only cares about the individual's ranking. Even if external interference causes all options to decline overall that night, as long as the relative order of good and bad among these options remains unchanged, the search center m generated by the algorithm will still evolve in the correct direction, demonstrating strong resistance to noise.
[0108] (3) Third step: Step size control (CSA: Cumulative step size adaptive):
[0109] We maintain an evolution path p_σ, which records the direction of movement of the search center across multiple generations. If the movement direction is largely the same for several generations, the length of p_σ will be much larger than the expected length of a random walk. This means that the search is "sprinting" in a definite direction, and σ should be increased in this case. If the search direction jumps back and forth, the positive and negative vectors in p_σ will cancel each other out, causing the path to shorten, and σ should be decreased in this case.
[0110] The evolutionary path is an exponentially weighted moving average vector. In the (g+1)th generation, the update formula is as follows:
[0111]
[0112] In the formula, The learning rate accumulated for the path (usually set to approx4 / n, where n is the dimension); The effective population size (a constant related to the weights); It is the inverse square root of the covariance matrix, used to normalize the current step size to an isotropic space, so that the step sizes of each dimension are comparable.
[0113] The observed path length Expected path length generated by random search Comparison:
[0114]
[0115] Based on the comparison results, σ is adjusted according to the exponential law.
[0116]
[0117] In the formula, is the damping parameter, which is usually greater than 1 and is used to determine the rate of change of σ. If |p_σ|>E, the exponential term is positive and σ increases. If |p_σ|<E, the exponential term is negative and σ decreases.
[0118] (4) Covariance update: In the sleep temperature control scenario with spatial partition collaboration (head, torso, feet), two problems are faced: The human body's thermal regulation is not isolated. For example, in the sleep onset stage, in order to induce a decrease in core body temperature (CBT) to enter deep sleep, simply cooling the torso often causes muscle trembling; the correct approach is "slight cooling of the head + heating of the feet to induce vasodilation". This interdependence between partitions (physiological collaborative ratio) is unknown, nonlinear, and varies from person to person. Traditional PID control or independent variable optimization algorithms cannot learn this combination. The single-night sleep score J is severely interfered by random noise such as pre-sleep alcohol consumption, daytime fatigue, and turning over. If the control strategy is blindly changed only based on several good samples (local optima) of tonight, it is very easy to fall into overfitting; if the response is too slow, it cannot keep up with the user's physiological rhythm drift.
[0119] CMA updates the covariance matrix C through two complementary mechanisms: Rank-1 Update focuses on temporal correlation, it uses the long-term evolution path p_c to capture the global trend of the function space, and stretches the distribution along the direction of successful movement. Rank-μ Update focuses on population spatial information. It uses the distribution of the top μ temperature schemes (individuals) with the best performance on the current night relative to the current center to quickly correct the search shape.
[0120] A. Define the standardized step vector: Extract the preferred temperature scheme individuals ranked top μ in the sliding window, and calculate them relative to the current center standardized deviation vector, which describes the distribution shape of the preferred individuals relative to the center in three-dimensional space.
[0121]
[0122] B. Update the covariance evolution path , which is a vector of length n, used to accumulate the successful movement directions of multiple consecutive decision cycles:
[0123]
[0124] In the formula, is the learning rate for path update; is a piecewise function, used to temporarily freeze when the step size σ increases too fast The update prevents the covariance matrix from diverging due to a sudden increase in step size. The covariance evolution path is similar to momentum. C. Core matrix update formula: Rank-1 update takes the outer product of the evolution path vectors to generate a matrix with rank 1. The search distribution (multivariate Gaussian bubble) is "stretched" along the direction indicated by p_c. The weight of the matrix in the success direction is increased by the outer product of the evolution path. If the optimal temperature control point moves in the same direction for several consecutive nights, the matrix will be stretched in that direction, forming a long ellipse, thus allowing for a deeper search in that direction.
[0125]
[0126] The rank-μ update utilizes the spatial distribution of individuals in the preferred temperature scheme for that night to calculate the empirical covariance.
[0127]
[0128] Merge the old matrix with the two updated terms in a weighted ratio:
[0129]
[0130] In the formula, , These represent the learning rates for the two update methods. Using the Rank-μ update mechanism, the axial length and rotation angle of the search distribution are dynamically adjusted based on the spatial distribution of the best-performing candidate temperature schemes within the current decision-making cycle.
[0131] Based on the above analysis, this step utilizes reward values to determine the optimal temperature scheme for each individual and combines historical zone temperature control commands to jointly update the search center vector, global step size, and covariance matrix. This allows the system to continuously correct the search direction under conditions of multi-region correlation, feedback lag, and limited samples. Specifically, updating the search center vector focuses on the current optimal temperature combination, updating the global step size achieves a dynamic balance between exploration and convergence, and updating the covariance matrix enables continuous learning of zone linkage relationships. The combined effect of these three factors ensures that the target zone temperature control command is no longer a fixed rule output, but rather continuously approaches the personalized optimal state along with the sleep process and individual responses. Therefore, the system can improve its temperature regulation adaptability in different stages of sleep onset, stable sleep, and the later stages of sleep, improve the accuracy of multi-region coordinated temperature regulation, and enhance control stability under environmental changes, daytime activity changes, and sleep stage switching. It should be understood that the above example is for demonstration purposes only and is not limiting.
[0132] Since sleep patterns at a single night are influenced by random factors such as caffeine and stress, the system needs to distinguish between physiological improvements caused by temperature control and those caused by external fluctuations. The system struggles to differentiate between physiological improvements stemming from temperature control and those from external random factors (such as daytime stress or caffeine intake). Optionally, candidate temperature schemes are defined as parameter vectors of the mapping function between context parameters and zone temperature control commands. The method also includes: within each decision cycle, obtaining the current context parameters and calculating the target zone temperature control command under the current context parameters based on the current context parameters and the mapping function.
[0133] In this scheme, context parameters are used to characterize other factors that may affect sleep scores, including room temperature, humidity, pre-sleep HRV, daytime activity level, and daytime caffeine intake. The system reads these parameters in each decision cycle and uses them as input to the mapping function. The parameter vector corresponding to each temperature scheme carries adjustable mapping coefficients, bias terms, or partition weights. This parameter vector, together with the context parameters, determines the output partition temperature control command, enabling the same parameter vector to generate different target temperature control results under different contexts. The partition temperature control command can correspond to the target temperature, temperature rise / fall range, or duty cycle of different partitions such as the head, torso, and feet. After substituting the current context parameters into the mapping function, the controller obtains the target partition temperature control command under the current context parameters and sends it to each partition temperature control module for execution.
[0134] In practical implementation, the mapping function can be constructed from a linear model, a piecewise function, or a lightweight neural network, and the parameter vector can be updated jointly by historical zone temperature control commands and reward values. After receiving the current context parameters, the controller first performs normalization and feature concatenation, then calculates the output value corresponding to the parameter vector, and generates the target control quantity for each zone based on the output value. Since this parameter vector directly depicts the mapping relationship from the context to the zone temperature control commands, it can maintain continuous adjustment of the temperature control strategy in different sleep stages, enabling the temperature control output to be updated synchronously with changes in the environment and user status. In practical applications, this mapping function can also adopt other model forms, which are not limited in this embodiment.
[0135] Specifically, the following steps are used to eliminate the contribution of "non-thermostatic factors" from complex physiological feedback and obtain a pure thermostatic benefit signal:
[0136] The system collects various context parameters and packages them into a context parameter feature vector X_{context}. A low-dimensional metabolic load vector can be obtained by compressing (1) biological feature context, (2) behavioral feature context and (3) environmental feature context. Among them, (1) biological feature context can include: a. Neural excitation factor (BiologicalContext), which is used to reflect the excitation state of the user's autonomic nervous system: reflecting the user's autonomic nervous tension and sensitivity to external interference. Input parameters: pre-sleep HRV (heart rate variability), caffeine intake, alcohol intake, environmental noise. b. Thermal load factor, which is used to characterize the extra energy stored in the body before the user falls asleep. Input parameters: daytime activity (METs), pre-sleep high-protein diet (food thermic effect), environmental humidity, initial room temperature. The higher this factor, the stronger the cold stimulus usually needs in the initial stage of the system. Daytime activity (METs): the muscle metabolic thermogeneity (thermal effect) after high-intensity exercise will last for several hours. Intake: drinking alcohol at dinner and caffeine intake will directly change the sleep latency and deep sleep structure. High-protein foods consumed within two hours of bedtime produce a significant thermic effect of food, increasing metabolic heat. Humidity: Increased humidity reduces heat dissipation efficiency. c. Rhythm and metabolic baseline factors, used to reflect the natural shift of the body's core body temperature (CBT) baseline. Input parameters: Time of day, menstrual cycle stage (female users), basal metabolic rate (determined by age / gender). Melatonin and core body temperature baselines at 10:00 PM and 1:00 AM are completely different. Basal body temperature rises during menstruation, requiring predictive triggering of compensation logic.
[0137] By defining individual temperature schemes as parameter vectors of a mapping function, the search object is transformed from a single fixed temperature value into learnable policy parameters, thereby improving the ability to express complex contexts. Since the target partition temperature control command is calculated based on the current context parameters within each decision cycle, the system can adapt to environmental fluctuations and individual differences in real time during sleep, reducing biases caused by fixed strategies and improving the stability and convergence efficiency of partitioned collaborative temperature control.
[0138] In one possible implementation, an adaptive covariance matrix evolution strategy is used to determine the individual temperature schemes for the next decision cycle, including: configuring a search space. In this embodiment, the search space includes a search center vector, a global step size, and a covariance matrix, where the covariance matrix is used to characterize the correlation between temperature control commands in each zone. The covariance matrix is configured to characterize the correlation between various temperature control feature dimensions; using the reward value corresponding to each individual temperature scheme, a preferred individual temperature scheme is determined; based on the preferred individual temperature scheme, the search space is updated; and based on the search space, a population including multiple individuals temperature schemes is generated by sampling.
[0139] In this step, the search space is used to describe the mathematical space of the possible range of temperature control parameters. The search center vector represents the center position of the current search, i.e., the mean of the current optimal temperature strategy. Each dimension corresponds to the target value of the temperature control command for different zones, such as the temperature setpoint, temperature difference compensation value, or adjustment range for the head region, torso region, foot region, or multiple independent temperature control zones on the bed surface. The global step size controls the size of the search range. Its value reflects the overall scale when expanding the scheme around the search center vector. When the global step size is large, the temperature scheme covers a wider range of parameters, which is conducive to exploring new temperature control combinations; when the global step size is small, the search will focus on the current better region, which is conducive to stable convergence. In this application, the dimensions of the covariance matrix correspond to the N temperature control feature dimensions contained in the decision vector. The off-diagonal elements of this matrix represent the intensity of the cooperative change of different temperature control feature dimensions (e.g., head temperature control dimension, torso temperature control dimension, and foot temperature control dimension) in the search distribution. Specifically, the covariance matrix is an N-by-N matrix, and its off-diagonal elements C_{ij} reflect the distributional synergy between the i-th temperature control feature dimension and the j-th temperature control feature dimension in the search space, thereby depicting the correlation of different body parts (or different time stages) in temperature control strategies.
[0140] In specific implementation, the execution entity in this application embodiment can be a controller in an intelligent sleep regulation system. The controller can be located on the main control board of the smart mattress, a bedside control terminal, a gateway device, a cloud server, or a control architecture composed of edge processors and the cloud. Before entering a new decision cycle, the controller first reads historical zone temperature control instructions, physiological feedback data corresponding to historical execution periods, and current environmental parameters from the memory. Historical zone temperature control instructions may include multi-period head temperature setpoints, torso temperature setpoints, foot temperature setpoints, and their adjustment rates throughout the previous night's sleep, and may also include several rounds of control instructions already executed in the current night. Based on these historical zone temperature control instructions, the controller determines the initial value of the search center vector, so that the search starting point falls in an area that is relatively close to the user's previous sleep habits. In one possible embodiment, the zone temperature control instructions with higher reward values in the most recent effective decision cycles can be weighted and averaged to obtain the search center vector; in another exemplary implementation, the most recently executed and stable target zone temperature control instruction can also be selected as the initialization benchmark for the search center vector.
[0141] The global step size can be determined based on the fluctuation range of historical zone temperature control commands, the user's sensitivity to temperature changes, and the current seasonal environment. For example, if historical data shows that the user's suitable nighttime temperature distribution is relatively concentrated and the indoor temperature and humidity changes are small, the global step size can be set to a smaller value to reduce invalid searches. If the user's daily activity level is high, the ambient temperature fluctuates significantly, or the current user status deviates greatly from the historical average, the global step size can be increased to improve the search space's coverage of individual dynamic changes. The covariance matrix can be initialized as a scaled identity matrix to represent the initial approximate independence of temperature control commands for each zone; alternatively, the sample covariance matrix can be calculated based on the collaborative changes in historical zone temperature control commands to obtain an initial correlation structure that better reflects the individual's actual thermal preferences. For example, if historical data shows that when foot temperature rises, trunk temperature often needs to decrease slightly to avoid excessive overall heat load, the corresponding dimensions of the foot and trunk in the covariance matrix can form a negative or weakly positive correlation structure.
[0142] To ensure the feasibility of the search space, physical and safety constraints can be added to the search center vector and covariance matrix in this step. The temperature setpoint can be limited to the device's allowable range, such as 15 to 45 degrees Celsius. The temperature difference between zones can be limited to human comfort and safety thresholds. The rate of temperature change can be limited to no more than a preset value per minute to avoid user discomfort due to rapid temperature changes. If the calculated covariance matrix does not meet the positive definite condition, the controller can adjust it to a positive definite matrix through eigenvalue correction, diagonal loading, or matrix projection to ensure stable subsequent multivariate normal sampling. Based on the above analysis, by configuring a search space containing the search center vector, global step size, and covariance matrix at the beginning of each decision cycle, the system can not only provide a unified parameter basis for subsequent temperature scheme generation but also explicitly express the correlation between multi-zone temperature control commands using the covariance matrix. This avoids the search distortion problem caused by fragmented processing of each temperature control zone, ensuring that multi-zone joint temperature control is established on a learnable correlation model from the outset. It should be understood that the above example is for demonstration purposes only and is not a limitation.
[0143] In this embodiment, the covariance matrix is configured to characterize the correlation between various temperature control feature dimensions.
[0144] In this step, the temperature scheme population is a set of multiple individual temperature schemes used for parallel evaluation of different temperature control strategies. Each individual temperature scheme is a combination of individual temperature strategies within the population; essentially, it is a set of parameter vectors that can be directly converted into zone temperature control commands, with each dimension corresponding to the search center vector. When a candidate temperature scheme individual is parameterized for a multi-zone temperature control device, it is instantiated as a spatial decision vector individual, i.e., a vector composed of zone temperature control commands from each temperature control unit, used to characterize the parameter correlation of different spatial regions (such as head, torso, and feet) under coordinated temperature regulation. When a candidate temperature scheme individual is parameterized for staged temperature regulation within a sleep cycle, it is instantiated as a temporal decision vector individual, i.e., a vector composed of the temperature control command sequence corresponding to each sleep stage, used to characterize the dynamic temperature regulation logic of different time dimensions (such as N1, N2, N3, and REM) during the sleep process.
[0145] Sampling based on the search space refers to generating multiple individual temperature schemes based on a multivariate distribution jointly defined by the search center vector, the global step size, and the covariance matrix. Since the search center vector provides the mean position, the global step size controls the overall dispersion, and the covariance matrix determines the axial scaling and rotation direction of the search distribution, the population sampled from this distribution can not only cover different combinations near the current optimal solution but also reflect the correlation and changing trends between temperature control commands in different zones.
[0146] In practical implementation, after completing the search space configuration, the controller can use a multivariate normal distribution for sampling, which can be expressed as xk = m + σ·A·zk, where xk represents the k-th candidate temperature scheme, m represents the search center vector, σ represents the global step size, A represents the covariance matrix decomposition matrix, and zk represents a random vector following a standard normal distribution. In the above expression, the search center vector determines the reference position of the candidate temperature scheme, the global step size determines the overall search radius, and the covariance matrix decomposition matrix makes the sampling points exhibit a correlation distribution in different dimensions. This allows the generated candidate temperature schemes to include not only independent temperature rise and fall changes in a single zone, but also combined changes of coordinated temperature rise, coordinated temperature fall, or one rise and one fall in multiple zones. For a temperature control device with three zones, the candidate temperature scheme can be represented as a three-dimensional vector, corresponding to the head set temperature, torso set temperature, and foot set temperature, respectively. For a higher-dimensional bed matrix temperature control device, the candidate temperature scheme can be expanded into a multi-dimensional vector, corresponding to the control parameters of each grid region.
[0147] In one possible embodiment, the population size can be set based on computing resources, decision cycle length, and device response speed. For example, 4 to 20 candidate temperature scheme individuals can be generated per decision cycle, preferably the top 50%. If the decision cycle is short, the population size can be reduced and the reuse of historical search results can be increased to reduce temperature control fluctuations perceived by the user; if the decision cycle is long and allows for more exploration, the population size can be increased to improve search coverage. To prevent sampling results from exceeding physical control boundaries, the controller can perform boundary clipping, mirror reflection, or penalty shrinkage processing after generating each temperature scheme individual. For example, when the foot temperature setting in a certain temperature scheme individual exceeds the device's upper limit, it can be clipped to the upper limit value or reflected back to the effective range in a boundary symmetry manner. To avoid excessive temperature jumps within consecutive decision cycles, a rate of change constraint can be introduced after sampling. The new temperature scheme individual is compared with the instructions executed in the previous cycle. If the temperature difference change in the same partition in adjacent cycles exceeds a threshold, it is smoothed according to a preset ratio.
[0148] After individual temperature schemes are generated, the controller can assign a unique identifier to each individual and establish a correspondence between the individual and the execution time period, facilitating attribution analysis when collecting multidimensional physiological characteristic signals. In one exemplary implementation, the system does not require simultaneous execution of multiple individuals at the same time. Instead, different individuals can be executed sequentially in multiple consecutive micro-decision windows, or population evaluation can be completed by combining offline model estimation with online real-world execution. For the typical application scenario of a single user and a single bed, a decision cycle can be divided into several sub-windows, with each sub-window executing one individual temperature scheme and collecting corresponding physiological feedback after the window ends. For systems with simulation models, some individuals can be pre-screened using a user-individualized thermal comfort model and a sleep response prediction model before selecting individuals with higher scores to enter the online execution stage, thereby reducing ineffective perturbations.
[0149] This step can also incorporate noise control mechanisms to improve sampling stability. For example, some high-reward temperature scheme individuals can be retained as preferred individuals between adjacent decision cycles and directly incorporated into the population of new temperature schemes to enhance search continuity; alternatively, low-discrepancy random sequences can be used instead of completely random sampling during sampling to reduce coverage bias under small sample conditions. Based on the above analysis, it can be seen that generating a population of temperature schemes based on search space sampling can simultaneously examine multiple possible regional temperature control combinations under limited sample conditions. This allows the system to move beyond single-path adjustments and improve its ability to discover multi-regional coordinated temperature control relationships through population search with relevant structures, providing sufficient candidates for subsequent reward value evaluation and search space updates. It should be understood that the above examples are for demonstration purposes only and are not limiting.
[0150] In this step, the optimal temperature scheme is the one with the best reward value among multiple temperature schemes, serving as the core basis for guiding the next round of search. Updating the search space based on historical zone temperature control commands means not only referencing the best-performing temperature schemes in the current decision-making cycle but also combining previously executed zone temperature control trajectories to ensure temporal continuity and consistency of individual habits in the update direction. The target zone temperature control command is the final output command generated based on the updated search space, used to drive the temperature control equipment to perform specific temperature adjustment operations on each zone.
[0151] In practice, the controller first sorts the reward values of all individual temperature schemes and then filters them based on data reliability, execution integrity, and safety constraints. The preferred temperature scheme can be the individual with the highest reward value, or a set of several top-ranked individuals. For example, in a population of 8, the top 2 to 4 candidates can be selected as the preferred temperature schemes. If multiple individuals have similar reward values, a secondary sort can be performed based on factors such as temperature change smoothness, user subjective comfort feedback, or distance from historically high-scoring schemes. Subsequently, the controller updates the search center vector based on the preferred temperature schemes, using a weighted average approach where the weight decreases with reward value ranking, ensuring that higher-scoring individuals contribute more to the new search center vector. If the preferred individual is xi and the weight is wi, the updated search center vector can be represented as m′=Σwi·xi. This update method gradually moves the search center towards the zoned temperature control combinations with higher sleep improvement potential.
[0152] The global step size can be adaptively adjusted based on the evolutionary path or the dispersion of the preferred individuals. When the preferred temperature schemes improve in similar directions over multiple consecutive decision cycles, it indicates that the current search direction is relatively stable, and the global step size can be appropriately reduced to enhance local fine-grained optimization. When the preferred individuals are more dispersed or the reward value does not show a significant increase, it indicates that an effective region may not have been found yet, and the global step size can be increased to expand the search. The covariance matrix is updated to learn the true correlation structure between temperature control commands in each partition. In one possible embodiment, a combination of rank-1 and rank-μ update mechanisms can be used, where rank-1 update emphasizes the directional information of the continuous evolutionary path, and rank-μ update emphasizes the statistical distribution information of multiple preferred individuals. Through this update, the direction and scale of the characteristic axes of the covariance matrix will be continuously adjusted according to the preferred temperature schemes, thereby gradually approaching the user's individualized thermoregulation pattern. For example, if multiple high-reward temperature schemes consistently show a combination trend of slightly cooling the head, slightly warming the feet, and maintaining a stable torso, the covariance matrix will gradually strengthen the correlation direction between these dimensions, making subsequent sampling more likely to generate similar synergistic combinations.
[0153] In this step, historical zone temperature control commands not only provide an initialization basis but also suppress aggressive updates that deviate excessively from the user's existing comfort zone. The controller can form a sliding window trajectory from the historical zone temperature control commands and calculate the offset between the new search center vector and this trajectory. When the offset exceeds a threshold, the update result is pulled back using an inertia coefficient or regularization term to ensure that the newly generated target zone temperature control command reflects the improvement direction brought about by recent physiological feedback without drastic jumps due to a single noise reward. Subsequently, the controller generates target zone temperature control commands based on the updated search space. The generation method can be to directly output the updated search center vector as the control target for the next stage, or to further sample from the updated distribution and combine safety, smoothness, and energy consumption constraints to select the optimal execution scheme. The target zone temperature control command can specifically include the target temperature value of each zone, the heating and cooling rate, the execution duration, and the start and stop sequence of each zone. For example, during a certain late-night period, the target zone temperature control command can be set to lower the head area by 0.5 degrees Celsius, keep the torso area unchanged, and raise the foot area by 0.8 degrees Celsius, and be completed gradually within 5 minutes to balance physiological improvement and comfort.
[0154] After receiving the target zone temperature control command, the temperature control device executes the corresponding action through bus communication, wireless communication, or an embedded interface. The temperature control device can be a zone temperature control unit composed of a built-in liquid cooling circulation module, a semiconductor cooling and heating module, an electric heating film, an air duct supply module, or a phase change temperature control unit. Each zone actuator adjusts the heating power, cooling power, fluid flow rate, valve opening, or air volume according to the control command to gradually bring the actual bed surface or contact interface temperature closer to the target value. The controller can also simultaneously read execution feedback, such as real-time temperature returned by the zone temperature sensor, actuator operating status, and energy consumption data, to confirm that the target zone temperature control command has been effectively executed. The controller then writes the execution result into the historical zone temperature control command database for continued use in the next round of search space updates.
[0155] In one possible implementation, the method further includes: obtaining the partition temperature control command and corresponding reward value of the current sleep cycle; constructing a historical buffer by introducing a sliding window mechanism; storing the partition temperature control command and corresponding reward value of the current sleep cycle into the historical buffer; executing a temperature scheme individual within each decision cycle; obtaining the reward value corresponding to the temperature scheme individual; storing the latest executed temperature scheme individual and its corresponding reward value into the historical buffer; and removing the oldest temperature scheme individual from the historical buffer; and updating the covariance matrix based on the temperature scheme individuals and their corresponding reward values in the historical buffer.
[0156] The current sleep cycle's zone temperature control command refers to the temperature control parameters output for zones such as the head, torso, or feet within the same sleep cycle. The reward value characterizes the comprehensive improvement effect of the zone temperature control command on sleep quality, body movement inhibition, and thermal comfort. A history buffer stores time-series samples of finite length. A sliding window mechanism automatically discards the oldest samples when new samples are written to the buffer, maintaining a constant sample size and retaining recent valid information. Candidate temperature schemes can be represented as parameter vectors of zone temperature control commands, and the covariance matrix characterizes the correlation and joint search direction among the zone temperature control parameters.
[0157] In specific implementation, at the end of each sleep cycle or each decision cycle, the controller writes the currently obtained partition temperature control command and corresponding reward value into a historical buffer. The historical buffer can be implemented using a ring memory or a queue structure, and its capacity can be set to cover the length of multiple recent decision cycles to retain sufficient dynamic samples. Whenever a new temperature scheme individual is executed and generates a reward value, the controller appends the individual and reward value to the end of the buffer, while deleting the earliest sample at the beginning of the buffer, so that subsequent covariance updates are based only on the most recent temperature control feedback. The update of the covariance matrix can be re-estimated based on the mean offset and reward weight of the samples in the buffer, and high-contribution samples are given higher weights in combination with the reward value to enhance the ability to characterize effective temperature control combinations. In practical applications, the covariance matrix can also be implemented in the form of numerical matrices of different sizes, which is not limited in this embodiment. In one possible embodiment, the method further includes: outputting the corresponding target partition temperature control command to the temperature control device within multiple consecutive decision cycles; using each target partition temperature control command generated within multiple consecutive decision cycles as a multi-dimensional temperature scheme vector, the corresponding sample covariance matrix formed by them conforms to the following characteristics:
[0158] (i) The sample covariance matrix is a strictly positive definite matrix, and the ratio of its minimum eigenvalue to its maximum eigenvalue is greater than a preset proportion threshold. The preset proportion threshold is configured to be greater than the upper limit of the random variation ratio caused by hardware measurement and execution noise of the temperature control system, and the preset proportion threshold is configured to be 5%.
[0159] (ii) There exists at least one non-zero off-diagonal element in the sample covariance matrix;
[0160] (iii) The eigenvector corresponding to the largest eigenvalue in the sample covariance matrix has components that are not zero on the temperature setpoint coordinate axes of at least two temperature control zones.
[0161] The sample covariance matrix is constructed by sending target zone temperature control commands to each temperature control device. The dynamic update mechanism of the covariance matrix automatically learns the correlation between zone temperatures, breaking through the limitations of traditional independent adjustment, significantly improving the effect of sleep quality regulation, and ensuring that the zone temperature control strategy meets the user's personalized physiological needs.
[0162] In one implementation scenario, the setpoints (columns, spatial dimension) of each independent temperature control zone collected within a continuous observation period (rows, time dimension) are aligned to construct... The original observation matrix of dimension (in For the number of periods, (This represents the number of partitions). Each row corresponds to a temperature scheme feature vector in the multidimensional control space.
[0163] For the observation matrix Perform column-wise arithmetic averaging to generate a mean vector representing the expected temperature control of each zone in the system. Subsequently, using Data centralization is performed to eliminate systematic biases in the spatial topology distribution caused by the initial reference temperature and constant terms of each partition. The centered matrix... Its transpose Perform matrix multiplication inner product operations and introduce Bessel correction coefficients. Unbiased statistical estimation is performed, and the final solution is obtained. dimensional sample covariance matrix :
[0164]
[0165] This matrix serves as a second-order statistical operator for the control space. Its diagonal elements directly measure the exploratory variability of instructions in each single partition, while its off-diagonal elements correspond to the non-independent collaborative linkage characteristics between hardware partitions.
[0166] Traditional sleep temperature control systems often employ PID control or piecewise expert curves based on independent sensor feedback for each zone. In such systems, the logic of each temperature control zone is independent of each other, or only depends on triggering within a fixed time period. If the system in question uses this traditional control, since each zone is orthogonal and does not interfere with each other during fine-tuning, the off-diagonal elements of the sample covariance matrix calculated from the output zone temperature control command dataset are zero. In this application, "zero" refers to a theoretical zero; in actual engineering, it is supported only by extremely weak hardware random measurement noise, thus approaching zero.
[0167]
[0168] In some traditional control engineering projects, one or more fixed multi-zone linkage rules may be used. While this hard-coded control can generate linkage, resulting in non-zero off-diagonal elements of the covariance matrix (satisfying feature two) and a skewed eigenvector of the largest eigenvalue (satisfying feature three), this linkage is a completely deterministic fixed rule, and the system lacks the ability to perform probabilistic exploration in multidimensional space. This causes the resulting scatter plot to geometrically collapse to a one-dimensional line or two-dimensional plane, and the corresponding sample covariance matrix theoretically suffers from eigenvalue degradation (the existence of zero eigenvalues). The tiny secondary eigenvalues actually observed are merely random measurement noise from the system hardware and lack statistical significance for spatial exploration.
[0169] By introducing the constraint that "the sample covariance matrix is a strictly positive definite matrix, and the ratio of its minimum eigenvalue to its maximum eigenvalue is greater than a preset proportional threshold (e.g., 5%)", this threshold constraint requires the system to have an active, non-degenerate probabilistic exploration volume in the multidimensional control space, thereby completely excluding traditional deterministic control that only has the appearance of linkage but lacks multidimensional evolution capabilities.
[0170] Some intelligent temperature control systems may employ conventional genetic algorithms for parameter optimization. Ordinary genetic or evolutionary algorithms typically use independent one-dimensional Gaussian mutation during mutation operations, meaning that the temperature of each temperature control zone is increased by an independent random perturbation. Since conventional mutation lacks the ability to adaptively and collaboratively adjust direction, the covariance matrix corresponding to its mutation operator is a diagonal matrix (all values outside the diagonal are 0). Although it can generate a full-rank exploration volume, its scatter plot geometric envelope is a sphere or an axisymmetric hyperellipsoid. The eigenvector corresponding to its largest eigenvalue must be parallel to a single hardware zone temperature coordinate axis (e.g., [1,0,0]), meaning the eigenvector's components on other coordinate axes are zero (the principal axis is not tilted). Therefore, it does not meet condition three: the eigenvector corresponding to the largest eigenvalue in the sample covariance matrix must have non-zero components on at least two temperature control zone temperature setpoint coordinate axes.
[0171] If Bayesian optimization or deep reinforcement learning (such as PPO, the existing λ-sample algorithm plus deterministic policy network) is used to adjust sleep temperature: the decision output of Bayesian optimization (BO) is controlled by the maximum point of the acquisition function. In the later stage of optimization, the acquisition function guides the sampling to be highly concentrated near a few optimal points, which is essentially equivalent to repeated sampling in a very small neighborhood. At this time, most feature values of the sample set are close to the level of hardware measurement noise, and the ratio of the minimum to the maximum feature value will be far below the 5% threshold, thus not satisfying the strict positive definiteness and proportionality conditions in feature (I). If all samples are combined for calculation, the disordered jumps in the early stage and the convergence point clusters in the later stage coexist, and the main eigenvector of the covariance matrix formed does not stably reflect the direction of multi-partition collaborative exploration, lacking the stable multi-partition tilt structure required by feature (III).
[0172] When Deep Reinforcement Learning (RL) outputs actions, its exploration noise typically employs isotropic action space noise (such as diagonal Gaussian noise). Since the mean of the policy network output is deterministic, its multidimensional output manifests as independent circular jitters around a deterministic trajectory, failing to form a full-rank tilted ellipsoid that autonomously rotates along its principal axis as the reward value changes. The sample covariance matrix of the zoned temperature control commands output by the RL system within a continuous decision-making cycle typically has eigenvectors corresponding to its largest eigenvalues that are approximately parallel to the temperature coordinate axis of a single zone, rather than having significant components on the coordinate axes of at least two zones, thus failing to satisfy feature (iii).
[0173] In one implementation scenario, a sliding window mechanism is introduced for rolling updates. CMA-ES is generational, requiring sampling of the entire population at once, uniform evaluation and reward calculation, followed by a batch update of the mean and covariance matrix as a decision cycle. However, in the physical scenario of sleep temperature control, due to the long cycle of complete sleep stages, only a very small number of effective feedback samples can be generated daily. Directly using traditional CMA-ES would result in an excessively long overnight iteration cycle and an inability to track the user's circadian rhythm drift in real time. Therefore, a sliding window mechanism is introduced. The system generates and executes one partition temperature control command Φ_1 each night. The next day, the reward value J_1 corresponding to the partition temperature control command is obtained, and a first-in-first-out (FIFO) historical buffer of length λ (λ is the population size) is maintained. Each day, the latest partition temperature control command and its corresponding reward value (Φ_i, J_i) are pushed into the queue, while the oldest data is removed. This continuously sliding window is used to perform approximate local covariance updates.
[0174] In one example, the newly generated tuple (Φ_{new}, J_{adjusted_new}) is added to the sliding window, while the oldest historical tuple in the window is removed. Then, among the existing λ samples in the current window, the system performs intra-group comparisons and rank-based sorting based solely on the incremental score J_{adjusted} to determine the weight w_i of each sample.
[0175] In the CMA-ES framework based on sliding window, a weight decay factor is introduced. By applying a weight that decreases exponentially over time to the samples within the window, the algorithm can automatically forget the physiological feedback under the old mean and thus focus on the latest search region.
[0176] Specifically, the data format of the first-in-first-out (FIFO) history buffer is modified. In addition to storing the partition temperature control command Φ and the reward value J, the time step (number of days or decision cycle) of the sample entering the window is recorded, and a window storage tuple (Φ_i, J_i, ...) is constructed. ),in, It's the time offset, the latest sample. =0, the oldest sample =λ-1.
[0177] In each decision cycle, the λ samples within the sliding window are sorted from largest to smallest according to their reward score J. The individual ranked 1st is assigned the highest original weight w_1. The individual ranked λth has the lowest weight.
[0178] Define a decay coefficient alpha (suggested range: 0.7 - 1.0). For each sorted individual i, calculate its decayed weight using the following formula:
[0179]
[0180] In the formula, For individual weights after decay, The original weights are used if a well-performing solution was generated 7 days ago. =6), even if it ranks high, its actual influence will be weakened by the time decay mechanism.
[0181] Weight normalization. To ensure mathematical consistency between CMA-ES mean and covariance updates, it is essential to ensure that the sum of the weights of the top μ preferred temperature schemes equals 1.
[0182]
[0183] use Replace the weight terms in the original algorithm and update the rank-μ terms of the mean m and covariance matrix C.
[0184] The buffer maintenance mechanism ensures that search space updates continuously rely on recent sleep feedback, preventing outdated samples from interfering with current individual differences and environmental changes. The covariance matrix thus more accurately reflects the correlation between temperature control commands in each zone, allowing subsequent sampling to focus on high-yield areas and improving the response speed to sleep stage transitions and circadian rhythm shifts. Consequently, the system maintains good convergence and stability under limited sample conditions, enhancing the optimization accuracy and individual adaptability of zone-based coordinated temperature regulation.
[0185] In one possible implementation, the covariance matrix is updated based on the individual temperature schemes and their corresponding reward values within the historical buffer. This includes: assigning exponentially decaying weights to the reward values within the historical buffer based on their chronological order, and updating the covariance matrix according to the decayed weights; or, obtaining the environmental memory corresponding to the partitioned temperature control instructions and reward values within the historical buffer, calculating and correcting the importance weights, and updating the covariance matrix using the corrected importance weights.
[0186] In this embodiment, the history buffer is used to store executed zone temperature control commands, corresponding individual temperature schemes, and reward values calculated from multidimensional physiological characteristic signals. It can also synchronously record the environmental memory corresponding to the command. The environmental memory includes the historical distribution parameters of the system at that moment, specifically the historical mean, historical step size, and historical covariance matrix.
[0187] When using a time-ordered update method, the system first arranges the reward values in the historical buffer from most recent to oldest according to their execution time, and then assigns an exponential decay weight to each reward value. This exponential decay weight can be determined by a preset decay coefficient and a time interval, ensuring that reward values from more recent times have a higher contribution and reward values from more distant times have a lower contribution. Subsequently, the system jointly calculates the sample covariance by combining the weighted reward values with the parameter vectors of the corresponding temperature scheme individuals, and then smoothly updates the covariance matrix accordingly to make the new matrix parameters more consistent with the recent sleep temperature regulation effect.
[0188] When employing the importance sampling correction method, the system first extracts the historical distribution parameters (including mean, step size, and covariance matrix) corresponding to the sample's generation time from the partition temperature control instructions in the historical buffer, using them as environmental memory. The system calculates the target probability density of each old sample under the current latest distribution and the proposal probability density under the corresponding historical distribution in the environmental memory, and calculates the importance weight (likelihood ratio) based on the ratio of these two values. The system then truncates and corrects the importance weights, setting a threshold to prevent a single outdated extreme sample from dominating the update of the entire covariance matrix. The truncated importance weights are multiplied by the initial weights based on reward ranking to obtain the corrected weights. Subsequently, the aforementioned corrected weights are normalized, and they are used to weight and accumulate the rank-mu update term of the covariance matrix, thereby updating the covariance matrix. This covariance matrix can be stored in the controller's memory in two-dimensional or multi-dimensional real symmetric matrix form.
[0189] In this way, the covariance matrix can continuously reflect the correlation changes between temperature control commands in different zones, and incorporate recent effective feedback or representative environmental memories into the update process, so that the sampling distribution of the search space gradually converges towards a more effective temperature control combination.
[0190] In one possible implementation, calculating and correcting the importance weights further includes: calculating the current distribution probability density and the historical distribution probability density based on the environmental memory corresponding to the partition temperature control instructions and reward values in the historical buffer; calculating the importance weights based on the current distribution probability density and the historical distribution probability density; and correcting the importance weights based on the reward values corresponding to each individual temperature scheme and a preset threshold.
[0191] In this embodiment, environmental memory is used to characterize the correlation between zone temperature control commands and reward values within the historical buffer, and it can be composed of the environmental feature vector corresponding to the current sleep stage. The historical buffer can adopt a circular cache structure to store control commands, reward values, and corresponding environmental memories from the most recent decision cycles, which facilitates maintaining a stable statistical basis when the sample size is limited and feedback is delayed.
[0192] When calculating the current and historical probability densities, the controller can establish a current distribution model based on samples in the historical buffer that are similar to the current environment memory, using Gaussian kernel density estimation or a Gaussian mixture model, and establish a historical distribution model based on all historical samples. The current probability density is used to characterize the likelihood of temperature control commands occurring in each zone under the current sleep environment, while the historical probability density is used to characterize the overall distribution characteristics in long-term accumulated samples. The probability density can be calculated separately for each zone's temperature dimension and then normalized to form a joint density value.
[0193] After obtaining the current probability density and the historical probability density, the system can calculate the importance weight based on the ratio or logarithmic difference between the two to characterize the contribution of the current sample relative to historical samples. This importance weight can be used to measure the representativeness of the individual with the current temperature scheme in the current sleep environment. The higher the current probability density and the lower the historical probability density, the greater the corresponding importance weight, thus giving control commands that are more in line with the current environment a higher weight in covariance updates.
[0194] When adjusting the importance weights based on the reward values corresponding to each individual temperature scheme, the controller compares the reward values with preset thresholds. When the reward value of an individual temperature scheme is higher than the threshold, the corresponding importance weight is increased; when the reward value is lower than the threshold, the corresponding importance weight is decreased. The preset thresholds can be determined by the quantiles of the historical reward distribution or set separately according to different sleep stages to reflect the different requirements for comfort and stability at different stages. The adjusted importance weights are then used for subsequent covariance matrix updates, enabling high-reward and highly matched zone temperature control commands to have stronger retention and propagation capabilities.
[0195] In one implementation scenario, importance sampling (IS) is used to correct for distribution bias. Within the sliding window CMA-ES framework, this allows the algorithm to utilize older data from the past few days, while mathematically offsetting biases caused by changes in the search center (mean m) and shape (covariance C).
[0196] To implement importance sampling, each tuple in the sliding window cannot only store scores, but must also store the "environmental memory" at the time of sampling. Therefore, each cache entry should contain: the partition temperature control instruction Φ_i, the multidimensional temperature vector at the time of sampling; the sleep reward J_i: the observed physiological feedback score; and the historical distribution parameters: the mean, step size, and covariance matrix of the system at the time the sample was generated.
[0197] For any old sample Φ_i within the window, calculate the ratio of its probability of occurrence under the current latest distribution to its probability under the old distribution at the time of sampling. The probability density of the target distribution is calculated... The probability density of the proposal distribution is obtained by calculating... get.
[0198] Importance weight (Likelihood Ratio):
[0199]
[0200] PDF calculation for a multivariate normal distribution is performed using the following formula:
[0201]
[0202] wherein is a composite covariance matrix.
[0203] Since directly using L_i may lead to excessively large weight when old samples are extremely rare under the current distribution. Calculate the corrected weight by multiplying the original reward ranking-based weight w_i of CMA-ES by L_i: , a threshold (e.g., 0.1<L_i<10) is set to prevent a single outdated extreme sample from dominating the update of the entire covariance matrix.
[0204] normalize the weights of the first μ preferred temperature scheme individuals within the window to make their sum equal to 1, then update m and C according to the standard CMA-ES formula. The Rank-μ update of CMA-ES requires that all individuals must come from the current probability distribution. IS, through mathematical transformation (ratio correction), forcibly "disguises" outdated samples as sampled from the current distribution, so that the formula is mathematically valid again.
[0205] Through the above method, the system can introduce both environmental similarity and reward feedback into the weight calculation process, reduce the deviation caused by simply relying on frequency statistics, and suppress the interference of low-quality samples to covariance update. Accordingly, the search direction of the zoned temperature control instruction in different sleep environments is more stable, the generation accuracy of the target zoned temperature control instruction is improved, and thereby the individual adaptability and continuous optimization capability of sleep temperature regulation are enhanced.
[0206] On the basis of the foregoing embodiment, further, determining preferred temperature scheme individuals by using the reward values corresponding to each temperature scheme individual includes: obtaining an improvement increment as the reward value by combining the sleep scores of each sleep stage in the current sleep cycle with a pre-obtained historical average physiological expected value; adopting a preset time response function to perform time series attribution on the improvement increment, and converting the instantaneous improvement increment at a future time to the temperature control decision moment, so as to determine the preferred temperature scheme individuals by sorting the improvement increments.
[0207] In this embodiment, the sleep score is used to characterize the comprehensive sleep quality of each sleep stage, which can be obtained by weighting indicators such as sleep onset latency, deep sleep proportion, number of micro-awakenings, body movement intensity and heart rate variability. The historical average physiological expected value is used to characterize the baseline expected level of the same user under similar environments and similar sleep conditions, which can be obtained by offline statistics from historical sleep data, and a difference calculation is performed with the phased sleep score in the current sleep cycle, thereby forming the improvement increment. The improvement increment can be used as the reward value of the temperature scheme individual to reflect the promotion degree of the current temperature scheme to the user's sleep state.
[0208] The time response function describes the lag propagation of physiological improvements through temperature regulation. It can take the form of an exponentially decaying kernel, a piecewise linear kernel, or a convolutional kernel, mapping the instantaneous improvement increment generated in the future to the current temperature control decision moment. Through this attribution method, the system can trace back the improvement effect that only appears in the later stages of sleep to the corresponding temperature control action. This ensures that the reward evaluation of individual temperature plans is not limited to a single instantaneous feedback but covers its cumulative effect in subsequent sleep stages. Therefore, the controller can rank the individual temperature plans based on the calculated reward value and select the top-ranked individual as the preferred temperature plan.
[0209] This reward construction method combines changes in sleep quality during different sleep stages with historical expected baselines and uses a time response function to achieve delayed effect attribution, ensuring that the evaluation of the temperature scheme aligns with the actual timing of the temperature regulation effect. This reduces attribution bias caused by delayed physiological feedback, improves the accuracy of individual selection for optimal temperature schemes, and consequently enhances the stability of subsequent search space updates and the individual adaptability of temperature regulation.
[0210] Based on the aforementioned embodiments, the temperature scheme individual is further defined as a spatial decision vector, which is obtained by combining the partition temperature control instructions of each temperature control device. The method also includes: using the reward value corresponding to each spatial decision vector to update the search space and generate the target partition temperature control instruction.
[0211] In this embodiment, the spatial decision vector individual is used to characterize the combination of zoned temperature control commands from different temperature control devices within the same decision cycle. Each component corresponds to the target temperature setpoint for different areas of the head, torso, feet, or bed surface. The reward value can be composed of changes in sleep score, body movement inhibition effect, improvement in heart rate variability, and environmental comfort evaluation. After receiving the actual execution status of each temperature control device, the controller maps the physiological feedback and environmental feedback after execution to the corresponding reward value to reflect the degree to which the combination of zoned temperature control commands promotes sleep. The search space can be defined by the search center vector, global step size, and covariance matrix, where the covariance matrix describes the correlation between different zoned temperature control commands.
[0212] After acquiring the reward values corresponding to each spatial decision vector, the controller selects the optimal temperature scheme individuals based on the ranking of reward values. It then updates the search space by incorporating the current search center point, moving the new search center towards the high-reward region. Simultaneously, it updates the covariance matrix based on the individual distribution to adjust the axial length and rotation angle of the search space. This updated search space more accurately reflects the user's individualized temperature control preferences during the current sleep cycle, thereby generating target zone temperature control commands and outputting them to the corresponding temperature control devices for execution. These target zone temperature control commands can simultaneously apply to multiple zones within the same control cycle, thus achieving coordinated regulation of the sleep environment.
[0213] By abstracting the zone temperature control commands of each temperature control device into individual spatial decision vectors and updating the search space based on the corresponding reward values, the controller can continuously approach the optimal zone temperature control combination under limited sample conditions, enhance the modeling ability of multi-region correlation, and improve the adaptability to changes in sleep stages and individual differences, thereby improving the stability and adjustment accuracy of zone temperature control.
[0214] In one implementation scenario, CMA-ES no longer optimizes a specific temperature value, but rather a context-temperature function. In context-aware CMA-ES, an individual is a matrix (i.e., weight parameters). The following example uses stage N3; other sleep stages can be described similarly. Input (k context parameters). Output: .
[0215] Define a linear mapping In the formula, W is a 3×k weight matrix and b is a 3×1 bias vector. Flatten all elements of W and b to form a high-dimensional vector of length 3k+3, which serves as the individual for CMA-ES evolution. (1) Initialization steps: Initialize the mean vector Its length is 3k+3. Using expert initialization, [the following is done]: The portion corresponding to 'b' is set as the expert-recommended reference temperature. At this point... This represents the initial "general control logic" of the system.
[0216] (2) Sampling steps: Based on the current distribution Sample generation individual Each Each is a vector of length 3k+3. The current context parameters are collected. Calculate the actual instructions corresponding to each candidate temperature scheme: .
[0217] (3) Evaluation and ranking: Group temperature command Execute the command and obtain the corresponding sleep reward score J. After accumulating a specific number of cycles, score the rewards according to the value of J. Rank them.
[0218] (4) Mean update: Utilize the top-ranked preferred temperature scheme individuals (weight vector) Update the mean m. The updated m represents a better mapping function. This means the system has learned how to calculate more scientifically as X (the context parameter) changes. .
[0219] (5) Step size adaptation: based on the evolutionary path Adjustment If feedback from several consecutive cycles proves that "increasing a certain weight" is effective, This will increase and accelerate the evolution of the policy matrix W, solving the problem of adaptive policy adjustment speed under different environmental sensitivities.
[0220] (6) Matrix adaptation: Update the covariance matrix: The variance distribution on the diagonal directly reflects the “weight sensitivity” of different contexts to different parts. The off-diagonal elements of C begin to describe “coordination between strategies”.
[0221] Based on the aforementioned embodiments, the search space is further updated based on the preferred temperature scheme individuals, including: calculating the statistical mean of the preferred temperature scheme individuals in each temperature control feature dimension, and updating the search center vector to the statistical mean; calculating the standardized deviation vector between the preferred temperature scheme individuals and the current search center point based on the preferred temperature scheme individuals; maintaining and updating the evolutionary path of the covariance matrix based on the global step size and the standardized deviation vector; obtaining the first update term based on the vector outer product of the evolutionary path and the evolutionary path direction of the preferred temperature scheme individuals; calculating the empirical covariance using the spatial distribution of the preferred temperature scheme individuals in the current sleep cycle, obtaining the second update term; and weightedly fusing the covariance matrix, the first update term, and the second update term to adjust the axial length and rotation angle of the search space.
[0222] The current search center point characterizes the central position of the temperature control search distribution within the current decision-making cycle. The standardized deviation vector eliminates the influence of differences in the dimensions of temperature control commands across different zones on the update results, ensuring a unified representation of the deviation direction and magnitude of the preferred temperature schemes relative to the center point. The evolutionary path of the covariance matrix accumulates the changes in search direction over multiple decision-making cycles, while the global step size controls the scale of path updates, ensuring the search space maintains a stable expansion or contraction trend across different sleep stages. The outer product of the evolutionary paths and the evolutionary path direction of the preferred temperature schemes together form the first update term, reinforcing the directional information dominated by recent superior samples. The empirical covariance, obtained from the spatial distribution statistics of preferred temperature schemes within the current sleep cycle, reflects the correlation of multiple superior schemes across different zone temperature dimensions within this cycle. The second update term supplements the constraints of the sample distribution on the search shape.
[0223] In practical implementation, the controller extracts temperature values for each dimension from the partition temperature control command vector corresponding to the individual preferred temperature scheme, differs them from the current search center point, and then normalizes them according to the historical standard deviation of each dimension to obtain a standardized deviation vector. The evolutionary path can be stored as a vector with the same length as the partition dimension, and is updated exponentially smoothed in each decision cycle based on the global step size and the direction of the current superior sample. The empirical covariance can be calculated based on the set of superior samples ranked first in the current sleep cycle, combined with the sample mean, and summed by outer product of the deviation vector, and normalized according to the number of samples. Subsequently, the controller merges the covariance matrix, the first update term, and the second update term according to preset weights to generate a new covariance matrix, thereby changing the axial length and rotation angle of the search ellipsoid, so that subsequent sampling is closer to the effective temperature combination in the current sleep state.
[0224] Through the above update method, the search space can simultaneously absorb historical directional accumulation, recent high-quality sample trends, and spatial distribution information of the current sleep cycle, enabling multi-zone temperature control search to maintain good directional consistency and adaptability under limited sample conditions, thereby improving the convergence speed, stability, and individual adaptability of coordinated temperature regulation.
[0225] Based on the aforementioned embodiments, the temperature scheme individual is further defined as a time decision vector, which is obtained by combining the zone temperature control instructions of each sleep stage in the sleep cycle; the method also includes: using the reward value corresponding to each time decision vector to update the search space and generate the target zone temperature control instruction.
[0226] In this embodiment, the time decision vector individual is used to characterize the combination of zoned temperature control commands corresponding to different sleep stages within a sleep cycle. The zoned temperature control commands can consist of target temperatures and their durations for different temperature zones such as the head, torso, and feet. The reward value reflects the comprehensive contribution of the time decision vector individual to sleep stability, deep sleep maintenance, and micro-arousal inhibition. The controller can quantify the sleep improvement effect after executing temperature control commands at each stage based on collected EEG, heart rate variability, respiratory rhythm, body movement, and environmental temperature and humidity information, and generate corresponding reward values accordingly. The time decision vector individual can be expressed numerically as a sequence of temperature control parameters arranged according to sleep stages to facilitate unified calculation with historical samples, covariance structure, and search center vector. In practical applications, other encoding methods can also be selected, and this embodiment does not limit this approach.
[0227] After obtaining the reward values corresponding to each time decision vector individual, the controller considers individuals with higher reward values as candidate temperature schemes more favorable to the current sleep state, and updates the search space based on their temperature control parameter distribution to reduce inefficient areas and increase the sampling probability of high-yield areas. The updated search space can be further used to generate target zone temperature control instructions, enabling it to output zone temperature combinations that better meet user needs in subsequent sleep stages according to current individual differences and stage characteristics, thereby driving the temperature control device to perform corresponding adjustments.
[0228] In one implementation scenario, in this embodiment, the temperature control unit (e.g., a temperature-controlled pillow / mattress) is treated as a single temperature control unit, but the algorithm treats the four sleep stages (N1, N2, N3, REM) of a sleep cycle as a continuous control sequence. The search entity is a 4-dimensional vector:
[0229]
[0230] Expert initialization: Based on the sleep medicine manual, an initial mean is set to ensure that the algorithm starts to evolve from a high-probability correct region that conforms to the body's metabolic rhythms (such as cooling down when falling asleep and keeping warm in the early morning), reducing the slow convergence problem caused by high experimental costs.
[0231] In this embodiment, executing one individual requires one sleep cycle. When the population has λ individuals, executing λ individuals in λ sleep cycles constitutes one decision cycle. The covariance matrix C of CMA-ES is used to discover synergistic associations between sleep stages. Matrix transformation: The covariance matrix is updated via Rank-μ, stretching the search space from a "perfect circle" to a "tilted ellipse," representing the mutual influence between temperatures at different stages. This "shape adaptation" can automatically match the user's unique physiological response manifold (e.g., some users are extremely sensitive to N1 temperature but insensitive to REM).
[0232] By incorporating the sleep stage dimension into the decision space, the temperature control changes over time are matched with the sleep process, and the search direction is continuously corrected through reward feedback. This implementation improves the adaptability of staged temperature regulation during sleep, reduces deviations caused by fixed curve control, enhances the responsiveness to individual differences and stage switching, and improves the convergence efficiency and regulation stability of temperature control commands for the target zone.
[0233] Figure 2 This is a schematic diagram of a sleep temperature regulation system provided in an embodiment of this application, as shown below. Figure 2 As shown, the sleep temperature regulation system 20 provided in this embodiment includes: a zoned temperature control device 201, a biofeedback sensor, and a controller.
[0234] The zoned temperature control device 201 includes at least one temperature control unit for performing thermal regulation on the user's sleeping area;
[0235] The biofeedback sensing device 202 is used to acquire multidimensional physiological characteristic signals of the user in real time.
[0236] The controller 203 is electrically connected to the zone temperature control device 201 and the biofeedback sensing device 202. The controller 203 is used to receive multi-dimensional physiological characteristic signals, execute the sleep temperature adjustment method provided above, generate a target zone temperature control command and send it to the zone temperature control device 201, thereby controlling the zone temperature control device 201 to execute the target zone temperature control command to adjust the thermal parameters of different parts of the user's body.
[0237] In this system, zoned temperature control devices are set up with temperature control units corresponding to different areas of the head, torso, feet, or bed surface. Biofeedback sensors continuously output multi-dimensional physiological characteristic signals such as EEG, heart rate variability, respiratory rhythm, and body movement. Based on these signals, the controller executes sleep temperature adjustment methods and generates target zoned temperature control commands. This allows the thermal parameters of each zone to be adjusted in a coordinated manner, rather than in isolation, but in conjunction with the user's real-time physiological state for optimization. Thermal parameter adjustments can be fixed or used as a target average. For example, assuming a temperature of 26 degrees Celsius is implemented in stage N3, this includes both a fixed implementation of 26 degrees Celsius and using 26 degrees Celsius as an average value, setting a fluctuation range of 2 degrees Celsius above and below the target center value, or incorporating other perturbations at the starting point; no limitations are imposed in this application.
[0238] Because the controller can dynamically control the heat generation of different body parts based on multidimensional physiological feedback, it can improve the adaptability to changes in temperature requirements during sleep onset, restful sleep, and the later stages of sleep. Therefore, it can improve the effect of multi-zone coordinated temperature regulation, reduce temperature regulation deviation caused by fixed rules or single feedback, and improve the accuracy, stability and individualized adaptation of sleep temperature regulation.
[0239] Figure 3 This is a schematic diagram of a sleep temperature adjustment device provided in an embodiment of this application, as shown below. Figure 3 As shown, the sleep temperature adjustment device 30 provided in this embodiment includes:
[0240] Temperature scheme individual acquisition module 301 is used to acquire multi-dimensional physiological feature signals of the current decision cycle, and based on the multi-dimensional physiological feature signals, acquire the reward value corresponding to each temperature scheme individual. The temperature scheme individual is a decision vector containing multiple temperature control feature dimensions.
[0241] Temperature scheme individual determination module 302 is used to determine the temperature scheme individuals for the next decision period based on the temperature scheme individuals and their corresponding reward values in the current decision period, using an adaptive covariance matrix evolution strategy.
[0242] The target zone temperature control command acquisition module 303 is used to obtain the target zone temperature control command according to the individual temperature scheme, so as to drive the temperature control device to execute the target zone temperature control command.
[0243] By introducing a search center vector, global step size, and covariance matrix through the search space configuration module, the evolutionary trends and partition correlations in historical partition temperature control commands can be incorporated into a unified model, thus avoiding the isolated treatment of each temperature control region. The population sampling module generates multiple individual temperature schemes based on this, enabling the system to compare various partition combinations in parallel within a finite decision-making cycle. The reward value acquisition module combines multi-dimensional physiological characteristic signals to form a comprehensive evaluation of the schemes, making the feedback basis no longer limited to a single indicator, and thus more accurately reflecting the effects of sleep onset, stable sleep, and micro-arousal inhibition. The covariance update module updates the search space based on the optimal temperature scheme individuals and outputs the target partition temperature control command, thus continuously correcting the search direction and partition correlation, enabling the system to maintain good convergence efficiency, coordinated temperature regulation capability, and individual dynamic adaptability even under conditions of feedback lag and limited samples.
[0244] The sleep temperature adjustment device provided in this embodiment can perform the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0245] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 40 may include a memory 401 and a processor 402. Optionally, the electronic device may also include a transceiver 403, wherein the memory 401 and the processor 402 communicate with each other; for example, the memory 401, the processor 402 and the transceiver 403 may communicate via a communication bus 404, the memory 401 is used to store a computer program, and the processor 402 executes the computer program to implement the method of the above embodiments.
[0246] Optionally, the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps in the method embodiments disclosed in this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0247] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the methods in any of the above method embodiments.
[0248] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods in any of the above method embodiments.
[0249] All or part of the steps in the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.
[0250] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0251] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0252] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1The steps of the function specified in one or more boxes.
[0253] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of this application and its equivalents, this application also intends to include these modifications and variations.
[0254] In this application, the term "comprising" and its variations can refer to non-limiting inclusion; the term "or" and its variations can refer to "and / or". The terms "first", "second", etc., in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0255] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0256] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0257] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0258] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0259] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0260] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0261] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0262] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0263] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for adjusting sleep temperature, characterized in that, The method includes: Obtain multidimensional physiological feature signals for the current decision-making cycle, and based on the multidimensional physiological feature signals, obtain the reward value corresponding to each temperature scheme individual, wherein the temperature scheme individual is a decision vector containing multiple temperature control feature dimensions; Based on the individual temperature schemes and their corresponding reward values in the current decision-making cycle, the individual temperature schemes for the next decision-making cycle are determined using an adaptive covariance matrix evolution strategy. Based on the individual temperature scheme, a target zone temperature control command is obtained to drive the temperature control device to execute the target zone temperature control command.
2. The method according to claim 1, characterized in that, The process of determining individual temperature schemes for the next decision cycle using an adaptive covariance matrix evolution strategy includes: Configure a search space, which includes a search center vector, a global step size, and a covariance matrix. The covariance matrix is configured to characterize the correlation between the temperature control feature dimensions. Using the reward value corresponding to each of the temperature schemes, a preferred temperature scheme is determined, and the search space is updated based on the preferred temperature scheme. Based on the search space, a population comprising multiple individuals of the temperature scheme is generated by sampling.
3. The method according to claim 1, characterized in that, The temperature scheme is defined individually as a parameter vector of a mapping function between context parameters and zone temperature control commands; The method further includes: within each decision cycle, obtaining the current context parameters, and calculating the target partition temperature control command under the current context parameters based on the current context parameters and the mapping function.
4. The method according to claim 2, characterized in that, The method further includes: Obtain the zone temperature control command and corresponding reward value for the current sleep cycle. By introducing a sliding window mechanism, construct a history buffer and store the zone temperature control command and corresponding reward value for the current sleep cycle into the history buffer. Within each decision cycle, execute one of the temperature schemes, obtain the reward value corresponding to the temperature scheme, store the latest executed temperature scheme and its corresponding reward value in the historical buffer, and remove the oldest temperature scheme from the historical buffer. The covariance matrix is updated based on the individual temperature schemes and corresponding reward values within the historical buffer.
5. The method according to claim 4, characterized in that, The step of updating the covariance matrix based on the temperature scheme individuals and their corresponding reward values within the historical buffer includes: Based on the chronological order, the reward values in the historical buffer are assigned exponentially decaying weights, and the covariance matrix is updated according to the decayed weights. Alternatively, the environmental memory corresponding to the partition temperature control instructions and reward values in the historical buffer can be obtained, the importance weights can be calculated and corrected, and the covariance matrix can be updated using the corrected importance weights.
6. The method according to claim 5, characterized in that, The calculation of importance weights and the correction of those importance weights also include: Based on the environmental memory corresponding to the partition temperature control instructions and reward values in the historical buffer, calculate the current distribution probability density and the historical distribution probability density; The importance weight is calculated based on the current probability density and the historical probability density. Based on the reward value corresponding to each individual temperature scheme, the importance weight is adjusted according to a preset threshold.
7. The method according to claim 5, characterized in that, The step of determining the preferred temperature scheme individual using the reward value corresponding to each of the individual temperature schemes includes: The improvement increment is obtained by combining the sleep score of each sleep stage in the current sleep cycle with the pre-acquired historical average physiological expectation value as the reward value. Using a preset time response function, the improvement increment is attributed to a time series. By converting the instantaneous improvement increment at future moments to the temperature control decision moment, the optimal temperature scheme can be determined by ranking the improvement increments.
8. The method according to claim 2, characterized in that, The individual temperature scheme is a spatial decision vector, which is obtained by combining the zoned temperature control commands of each of the temperature control devices; the method further includes: The search space is updated using the reward values corresponding to each of the spatial decision vectors, and a target partition temperature control command is generated.
9. The method according to claim 2, characterized in that, The process of updating the search space based on the preferred temperature scheme includes: Calculate the statistical mean of the individual preferred temperature schemes in each of the temperature control feature dimensions, and update the search center vector to the statistical mean; Based on the individual preferred temperature scheme, calculate the standardized deviation vector between the individual preferred temperature scheme and the current search center point; The evolutionary path of the covariance matrix is maintained and updated based on the global step size and the standardized deviation vector. Based on the vector outer product of the evolutionary paths, and combined with the evolutionary path direction of the individuals in the preferred temperature scheme, the first update term is obtained; Using the spatial distribution of individuals with the preferred temperature scheme in the current sleep cycle, the empirical covariance is calculated to obtain the second update term. The covariance matrix, the first update term, and the second update term are then weighted and fused to adjust the axial length and rotation angle of the search space.
10. A method for adjusting sleep temperature, characterized in that, include: During multiple consecutive decision cycles, the corresponding target zone temperature control command is output to the temperature control equipment; Each target zone temperature control command generated within the consecutive decision-making cycles is taken as a multi-dimensional temperature scheme vector, and its corresponding sample covariance matrix conforms to the following characteristics: (i) The sample covariance matrix is a strictly positive definite matrix, and the ratio of its minimum eigenvalue to its maximum eigenvalue is greater than a preset proportion threshold. The preset proportion threshold is configured to be greater than the upper limit of the random variation ratio caused by hardware measurement and execution noise of the temperature control system, and the preset proportion threshold is configured to be 5%. (ii) At least one off-diagonal element in the sample covariance matrix is non-zero; (iii) The eigenvector corresponding to the largest eigenvalue in the sample covariance matrix has components that are not zero on the temperature setpoint coordinate axes of at least two temperature control zones.