A method and device for sleep thermal comfort perception and indoor sleep environment regulation
By constructing a sleep thermal comfort perception model, monitoring and predicting the sleep thermal comfort status in real time, and adjusting the HVAC system based on the predicted results, the problem of the impact of contact devices on the thermal comfort of sleepers is solved, and efficient energy use and a comfortable sleep environment are achieved.
Patent Information
- Application Number
- CN202510196505.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art is difficult to effectively solve the thermal comfort effect of contact devices on sleepers, resulting in waste of energy and discomfort sleeping environments.
By obtaining the video data, physiological data and environmental data of the subjects during sleep, data annotation, feature extraction and data fusion are carried out, sleep thermal comfort perception model is constructed, sleep thermal comfort status is monitored and predicted in real time, and the operating parameters of the HVAC system are adjusted according to the prediction results.
The thermal comfort monitoring and regulation of sleepers is achieved, which significantly improves energy use efficiency, reduces energy waste, and reduces overall energy consumption of buildings.
Smart Images

Figure CN119665417B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for sleep thermal comfort perception and indoor sleep environment regulation, belonging to the technical fields of computer vision and heating, ventilation and air conditioning control. Background Art
[0002] Sleep is an important physiological activity of the human body. Nearly one-third of a person's life is spent in sleep. The principle of sleep thermal comfort refers to people's perception and acceptance of the temperature of the sleep environment. This principle involves the physiological characteristics of the human body and the influence of environmental conditions on the human body. The perception of sleep thermal comfort by the human body is affected by various factors, including environmental temperature, relative humidity, wind speed, radiant temperature, etc. According to the thermal comfort principle, the human body will adjust the skin surface temperature and sweating to adapt to different environmental temperatures to maintain the thermal balance in the body. In recent years, using visual perception technology to obtain the sleep thermal comfort state of personnel has become a hot research field. This technology can greatly improve the efficiency of the heating, ventilation and air conditioning system and reduce energy waste caused by too low or too high temperature settings. By using visual perception technology, such as infrared cameras and night vision devices, the body temperature and movements of sleepers can be monitored in real time, so as to accurately judge their thermal comfort state. By analyzing these data, the operating parameters of the heating, ventilation and air conditioning system, such as indoor temperature, humidity and air flow, can be adjusted. This intelligent heating, ventilation and air conditioning control strategy based on visual perception can not only ensure the thermal comfort of sleepers, but also significantly improve the energy use efficiency. By avoiding unnecessary overcooling or overheating, energy waste can be greatly reduced, thus reducing the overall energy consumption of the building. Against the background of the large proportion of urban building energy consumption, the application of this technology is of great significance for promoting energy conservation, emission reduction and sustainable development of cities. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method and device for sleep thermal comfort perception and indoor sleep environment regulation. By predicting the sleep thermal comfort of personnel, the operating parameters of the heating, ventilation and air conditioning system can be adjusted accordingly, solving the influence of contact devices on sleepers, not only ensuring the thermal comfort of sleepers, but also significantly improving the energy use efficiency, thereby reducing the overall energy consumption of the building and achieving energy conservation and emission reduction.
[0004] To achieve the above purpose, the present invention is implemented by adopting the following technical solutions:
[0005] In the first aspect, the present invention provides a method for sleep thermal comfort perception and indoor sleep environment regulation, including:
[0006] Obtaining video data, physiological data and environmental data of a subject during sleep measured in advance;
[0007] After data annotation, feature extraction, and data fusion of video data, physiological data, and environmental data, a dataset is created.
[0008] Input the created dataset into a pre-constructed sleep thermal comfort perception model to obtain sleep thermal comfort perception prediction results.
[0009] Regulate the indoor sleep environment according to the obtained sleep thermal comfort perception prediction results.
[0010] Furthermore, the acquisition of pre-measured video data, physiological data, and environmental data of the subject during sleep includes:
[0011] Obtain video data of the subject during sleep through a thermal imaging dual-spectrum network barrel camera and a vision camera.
[0012] Use a physiological detection device to record the physiological data of the subject during sleep. The physiological data includes the subject's body temperature, heart rate, and electroencephalogram data.
[0013] Use an indoor thermal comfort meter to obtain indoor environmental data. The environmental data includes air temperature, relative humidity, air velocity, and carbon dioxide concentration.
[0014] Furthermore, the creation of a dataset after data annotation, feature extraction, and data fusion of video data, physiological data, and environmental data includes:
[0015] Extract frames from the video data at fixed intervals into time-series images, and perform video processing and annotation on the time-series images.
[0016] Use computer vision technology to process the time-series images after video processing and annotation to obtain images after extracting human pose features.
[0017] Synchronously record the physiological data and perform normalization processing on the physiological data.
[0018] Fuse the images after extracting human pose features, the normalized physiological data, and the environmental data to create a dataset.
[0019] Furthermore, the input of the created dataset into a pre-constructed sleep thermal comfort perception model to obtain sleep thermal comfort perception prediction results includes:
[0020] Input the created dataset into the sleep thermal comfort perception model. The sleep thermal comfort perception model includes a sleep thermal comfort pose detection module and a multi-modal data fusion module. The sleep thermal comfort pose detection module includes a human bone point detection module and a sleep comfort pose estimation module.
[0021] In the sleep thermal comfort posture detection module, the following steps are executed:
[0022] Detect the human body bone points in the images of the dataset through the human body bone point detection module, and identify the images with human body posture features;
[0023] Send the images with human body posture features into the sleep comfort posture estimation module for the recognition and analysis of sleep thermal comfort postures, and obtain each sleep posture and its corresponding thermal sensation voting value ;
[0024] In the multi-modal data fusion module, detect the physiological data in the dataset, and obtain each sleep posture and its corresponding thermal sensation voting value ;
[0025] The , are subjected to weighted average processing, and the obtained value is used as the final prediction result.
[0026] Furthermore, the detection of the human body bone points in the images of the dataset through the human body bone point detection module and the identification of the images with human body posture features include:
[0027] The images in the input dataset are first subjected to preliminary downsampling through two convolutional layers configured with 3x3 convolutional kernels and a stride of 2, and then feature extraction and scale dynamic adjustment are performed through an initialization layer, a Transition structure, and a Stage structure. Finally, feature abstraction and fusion are performed through a Transformer module and a multi-scale fusion module to detect the bone points of the subject during sleep and obtain the images with human body posture features.
[0028] Furthermore, the sending of the images with human body posture features into the sleep comfort posture estimation module for the recognition and analysis of sleep thermal comfort postures and obtaining each sleep posture and its corresponding thermal sensation voting value , includes:
[0029] The images with human body posture features are subjected to feature extraction through multiple convolutional layers and pooling layers, then the features at different levels are integrated using a multi-level feature fusion strategy, and feature transformation is performed through a Swin Transformer architecture to detect each sleep posture and its corresponding thermal sensation voting value .
[0030] Furthermore, the detection of the physiological data in the dataset in the multi-modal data fusion module and obtaining each sleep posture and its corresponding thermal sensation voting value , includes:
[0031] Input the physiological data in the dataset into the multi-modal data fusion module, perform preliminary processing through the LSTM module, then conduct feature fusion through the multi-head attention mechanism, and finally input the fused features into the fully connected layer for final prediction or classification to obtain each sleep posture and its corresponding thermal sensation voting value. 。
[0032] Further, the 、 are subjected to weighted average processing, and the calculation formula is as follows:
[0033] ;
[0034] Among them, is the thermal sensation voting value of each sleep posture and its corresponding after weighted average, is the weight of the sleep thermal comfort posture detection module, is the weight of the multi-modal data fusion module, and + = 1.
[0035] Second, the present invention provides a sleep thermal comfort perception and indoor sleep environment regulation device, including:
[0036] A data acquisition module for acquiring video data, physiological data, and environmental data of a subject during sleep measured in advance;
[0037] A dataset production module for making a dataset after data annotation, feature extraction, and data fusion of video data, physiological data, and environmental data;
[0038] A prediction module for inputting the made dataset into a pre-constructed sleep thermal comfort perception model to obtain a sleep thermal comfort perception prediction result;
[0039] A regulation module for regulating the indoor sleep environment according to the obtained sleep thermal comfort perception prediction result.
[0040] Third, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the foregoing are implemented.
[0041] Fourth, the present invention provides a computer device, including:
[0042] A memory for storing computer programs / instructions;
[0043] A processor for executing the computer programs / instructions to implement the steps of the method described in any one of the foregoing.
[0044] Fifth aspect, the present invention provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the method described in any one of the foregoing.
[0045] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0046] The present invention provides a method and device for sleep thermal comfort perception and indoor sleep environment regulation. By using a sleep thermal comfort perception model to detect the human body's thermal comfort level and adjusting the operating parameters of indoor air conditioning equipment accordingly, in this way, the indoor environment can be automatically adjusted according to the real-time thermal comfort state of the human body to ensure the provision of the best comfort level. It not only realizes precise thermal comfort adjustment but also effectively improves the adaptability and convenience of the system. In addition, through intelligent regulation, the system can significantly save energy, avoid energy waste, and thus achieve more efficient energy utilization and environmental protection goals. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is the overall block diagram of the sleep thermal comfort perception and indoor sleep environment regulation method provided by the embodiment of the present invention;
[0048] Figure 2 It is the overall block diagram of the sleep thermal comfort pose estimation model STCE-Net (Sleep Thermal Comfort Pose Recognition Network) in the embodiment of the present invention;
[0049] Figure 3 It is the structure diagram of the human body bone point detection module in the sleep thermal comfort pose estimation model STCE-Net of the embodiment of the present invention;
[0050] Figure 4 It is the structure diagram of the sleep pose estimation module in the sleep thermal comfort pose estimation model STCE-Net of the embodiment of the present invention;
[0051] Figure 5 It is the network architecture diagram of the multi-modal data fusion model MDF-Net in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0053] Embodiment 1. This embodiment introduces a method for sleep thermal comfort perception and indoor sleep environment regulation, including:
[0054] Obtain the video data, physiological data, and environmental data of the subject during sleep measured in advance;
[0055] After data annotation, feature extraction, and data fusion of video data, physiological data, and environmental data, a dataset is created.
[0056] The created dataset is input into a pre-constructed sleep thermal comfort perception model to obtain the sleep thermal comfort perception prediction result.
[0057] According to the obtained sleep thermal comfort perception prediction result, the indoor sleep environment is regulated.
[0058] As Figure 1 shown, the sleep thermal comfort perception and indoor sleep environment regulation method provided in this embodiment can accurately identify the sleep posture of the subject by collecting parameters related to indoor thermal comfort and physiological parameters of the subject during sleep, and combine the normalized analysis of the subject's body temperature, heart rate, and brain waves using a multi-modal data fusion model to predict the indoor thermal comfort, and at the same time feedback to the indoor heating, ventilation, and air conditioning control system. The specific steps are as follows:
[0059] S1: Data collection. A sleep video of the subject is obtained through a thermal imaging dual-spectrum network barrel camera and a vision camera. The body temperature of the subject is measured using an i-Button (temperature information button), and the changes in the heart rate and brain waves of the subject during sleep are recorded through a physiological detection device. The indoor thermal comfort meter mainly obtains the indoor air temperature, relative humidity, air velocity, and carbon dioxide concentration.
[0060] S2: Dataset creation. First, the video stream containing human sleep behavior is preprocessed, and frames are extracted at fixed intervals into a time series of images, then video processing and annotation are performed, and physiological data such as body temperature, heart rate, and brain waves are synchronously recorded and normalized. Then, computer vision technology is used to process the pictures to extract human posture and thermal map information. Finally, the extracted image features are fused with physiological and environmental data, and the sleep comfort is annotated to integrate into a structured sleep thermal comfort dataset.
[0061] S3: Sleep thermal comfort perception. The sleep posture of the subject is estimated through a sleep thermal comfort posture recognition algorithm, combined with the normalized processing and analysis of physiological data such as body temperature, heart rate, and brain waves using a multi-modal data fusion module, and a multi-modal-based sleep thermal comfort perception system is used to perform real-time perception of the sleep thermal comfort of indoor personnel.
[0062] S4: Intelligent control and adjustment. Using a multi-modal-based sleep thermal comfort perception system, through the sleep video of indoor personnel, physiological data such as body temperature, heart rate, and brain waves, and indoor environmental thermal comfort data collected in S2, the sleep thermal comfort of current indoor personnel is predicted, and based on the prediction result of the system, the indoor heating, cooling, ventilation, and air conditioning control system is controlled.
[0063] Specifically, step S3 is divided into two modules: the sleep thermal comfort posture estimation module and the multimodal data fusion module.
[0064] S301: The sleep thermal comfort posture estimation module STCE-Net is used to obtain the sleep postures of the subjects during sleep. STCE-Net mainly includes a human body bone point detection module and a sleep comfort posture estimation module, and its content is as follows:
[0065] In step S2, static images were extracted from the sleep videos recorded in the sleep experiment environment at an appropriate frame rate to construct a sleep thermal comfort posture dataset. Then, the human body bone points in the images, such as the head, shoulders, and legs, were labeled. Each image was classified and marked according to the comfort feedback during sleep, indicating whether each posture was related to uncomfortable sleep thermal comfort. Finally, these labeled images were segmented into a training set, a validation set, and a test set after data cleaning and verification, preparing for further analysis and model training.
[0066] As Figure 2 shown, the sleep thermal comfort posture estimation module STCE-Net will first load the dataset containing the labeled data. These images are preprocessed and marked with human body bone points, and each posture is classified according to its correlation with thermal comfort. Subsequently, the system performs human body bone point detection through the human body bone point detection module, identifies important posture features, and these posture features are sent to the sleep comfort posture estimation module for the recognition and analysis of sleep thermal comfort postures. Finally, the STCE-Net output will detail the relationship between each recognized sleep posture and thermal comfort.
[0067] Specifically, as Figure 3 shown, the human body bone point detection module is used to detect the bone points of the subjects during sleep, and its steps are as follows:
[0068] First, the input image first passes through two convolutional layers configured with 3x3 convolutional kernels and a stride of 2. After each layer, batch normalization and ReLU (activation function) are connected to achieve the initial downsampling process of the image, reducing its spatial size to 1 / 4 of the original size. Then, the image data is transferred to the initialization layer, which is composed of multiple repeatedly stacked Bottleneck (bottleneck layer) structures. Its main function is to adjust the number of feature channels while keeping the spatial size unchanged. Then, the image data flows to the initialization layer, which is mainly composed of repeatedly stacked Bottleneck structures, used to adjust the number of channels of the feature map without changing its spatial size.
[0069] The human body bone point detection module adopts a Transition (transition) structure and a Stage (stage) structure to achieve dynamic adjustment of scales and comprehensive utilization of features. In each Transition structure, the feature map is further downsampled through two parallel 3x3 convolutional layers to form multiple scale branches. Transition1 forms scale branches with a downsampling factor of 4 times and 8 times based on the output of the initialization layer; Transition2 further adds a scale branch with a downsampling factor of 16 times on the basis of these two scales, which is achieved by further downsampling the feature map with a downsampling factor of 8 times; Transition3 first receives the feature maps of each scale from Stage3, including feature maps with magnification scales such as 1x, 2x, 4x, etc. Starting from a convolutional layer applying a group of 3x3 convolutional kernels with a stride set to 2, this operation aims to reduce the spatial dimension of the feature map, reducing the feature map with a downsampling factor of 8 times from 16x12x128 to 8x6x256. During this process, not only is the size of the feature map reduced, but also the feature processing ability of the network is enhanced by increasing the number of output channels. The feature maps of each scale are processed similarly to ensure that all feature maps are effectively adjusted to be more suitable for further high-level feature abstraction and fusion processing in Stage4.
[0070] In the Stage structure, each scale branch first processes features through four BasicBlocks in ResNet (residual network). Subsequently, in the scale fusion stage, the outputs of each scale are upsampled by different multiples, added to the original scale output, and processed through the ReLU activation function to form a fused output. This fusion strategy ensures the effective integration of features at different scales and enhances the model's ability to capture details and context information.
[0071] One of the innovations of the present invention is that through each Transition structure, not only multiple scale branches are formed, but also a Transformer (transformer) module is introduced at the end of each Stage structure to further abstract and process features. The introduction of the Transformer module enables the network to more effectively process and integrate features at each scale, enhances the interaction ability between features using the self-attention mechanism, and thus enhances the recognition of complex patterns in the image. Subsequently, through the multi-scale fusion module, a multi-scale feature fusion strategy is implemented. The outputs of each scale not only include basic upsampling and addition operations, but also effectively fuse features at different scales through the Transformer module. This way ensures that the network can capture and integrate rich information from different scales, thereby significantly improving the overall recognition performance.
[0072] Among them, the sleep comfort posture estimation module is to detect the posture of the subject during sleep and its corresponding thermal sensation voting value TSV, as Figure 4 shown, and the steps are as follows:
[0073] Specifically, the input of the network is a preprocessed sleep image with a human skeleton. First, it is processed by a convolutional layer with a 3x3 convolutional kernel containing 32 filters. This layer is followed by batch normalization BatchNorm and the ReLU activation function, and then a 2x2 max pooling operation is performed to reduce the spatial size of the feature map. Continuing, the image data flows to the second convolutional layer, which uses 64 filters, also with a 3x3 convolutional kernel, followed by BatchNorm and ReLU, and max pooling is performed again. This process is repeated. As the network deepens, the number of filters in the convolutional layer increases to 128, 256, and finally reaches 512. After each step of increasing the number of filters, corresponding batch normalization and ReLU activation are performed, and max pooling is performed in some layers to continue reducing the feature map size. In this way, the network gradually extracts more and more complex features from the original image, while reducing the spatial dimension of the data and increasing the depth of the features.
[0074] After feature extraction, the network adopts a multi-level feature fusion strategy to integrate features at different levels. In this strategy, the lower-level feature maps are adjusted to the same scale as the higher-level feature maps through upsampling operations, and then these feature maps are merged and fused through further convolutional layers.
[0075] Immediately afterwards, it comes to the architecture of the Swin Transformer (a transformer model introducing the Shifted Window mechanism). It first splits the input RGB image into non-overlapping patches of equal size through a patch splitting module. Each patch is regarded as a patch token, and a total of N effective input sequence lengths of the Transformer are split out. More specifically, using of size and the number of channels for the patch, so the flattened feature dimension of each patch is , and there are patch tokens in total. In other words, each image is processed into image patches, each patch is flattened into a 48-dimensional token vector, and overall it is a flattened N×( ×3)=( ) × 48-dimensional 2D patch sequence.
[0076] It will then go through four stages, each stage consisting of a fully connected layer and a Swin Transformer block. The fully connected layer projects the tensor with the current dimension of to an arbitrary C dimension, obtaining a Linear Embedding with a dimension of . Subsequently, these patch tokens are fed into several Swin Transformer blocks with improved self-attention. The first Swin Transformer block keeps the number of input and output tokens constant at and is jointly designated as Stage 1 with the linear embedding layer.
[0077] To generate a hierarchical representation, as the network deepens, the number of tokens is gradually reduced through the patch merging layer. The first patch merging layer concatenates each group of 2×2 adjacent patches, and the number of patch tokens becomes of the original, that is , while the dimension of the patch tokens is expanded 4 times, that is 4C. Then, a linear layer is used for the 4C-dimensional patch concatenated features to reduce the output dimension to 2C. Then, Swin Transformer blocks are used for feature transformation, and its resolution remains unchanged. The first Patch merging layer and this feature transformation Swin Transformer block are designated as Stage 2. Repeat the same process as Stage 2 twice, and they are respectively designated as Stage 3 and Stage 4. The output resolution / number of patch tokens is respectively and . Each stage changes the dimension of the tensor, thus forming a hierarchical representation. Thus, this architecture can easily replace the backbone networks of various existing visual tasks.
[0078] Finally, the output of the entire STCE-Net (Sleep Thermal Comfort Pose Recognition Network) mainly includes the specific pose types of the sleeper and their corresponding thermal sensation voting values , and these outputs are intended to reflect whether the pose of the sleeper helps to maintain a suitable body temperature and comfort.
[0079] S302: Among them, the multi-modal data fusion model MDF-Net is used to simultaneously process the body temperature, heart rate, and brain wave information of the subject during sleep, so as to realize the detection of the subject's sleep thermal comfort, such asFigure 5 As shown, the content is as follows:
[0080] MDF-Net is used to efficiently process and fuse various sensor data to achieve comprehensive analysis and prediction of complex time series data. The network structure of MDF-Net includes multiple LSTM (Long Short-Term Memory) modules, a multi-head attention mechanism, a fully connected layer, and a data processing unit, and realizes the fusion and analysis of body temperature, heart rate, and electroencephalogram data through a series of processing steps. The following is a detailed description of this network structure and an analysis in combination with formulas.
[0081] First, the input module receives three-dimensional data from different sensors, including body temperature, heart rate, and electroencephalogram. The input data enters three LSTM modules respectively for preliminary processing. The internal structure of each LSTM module includes a forget gate, an input gate, an output gate, and the update of the cell state.
[0082] For each time step t, the calculation steps of the LSTM module are as follows:
[0083] The forget gate determines the information to be forgotten:
[0084] ;
[0085] Where, is the hidden state of the previous time step, is the current input, and are the weights and biases, is the sigmoid activation function, represents the output of the forget gate at time step t.
[0086] The input gate determines the information to be updated:
[0087] ;
[0088] ;
[0089] Where, is the new candidate value, is the output of the input gate at time step t, is the weight matrix of the input gate, is the bias term of the input gate, is the weight matrix of the candidate memory unit, is the bias term of the candidate memory unit, represents the hyperbolic tangent function.
[0090] Update of the cell state:
[0091] ;
[0092] Among them, is the output of the memory unit at time step t, that is, the updated memory, is the state of the memory unit at the previous moment.
[0093] The output gate determines the output information:
[0094] ;
[0095] ;
[0096] Among them, is the output of the output gate at time step t, which determines how much information of the current hidden state will be output, is the weight matrix of the output gate, is the bias term of the output gate, is the hidden state at the current moment, which determines the output of the LSTM, is the state of the memory unit at the current moment.
[0097] Each LSTM module processes the input data and outputs the hidden state representing the characteristics of the time series data.
[0098] The feature data processed by the LSTM module enters the multi-head attention mechanism. The multi-head attention mechanism processes the data through the following steps. First, the input feature vector is split into multiple heads for the multi-head splitting operation, and each head represents a different feature subspace.
[0099] The principle in the multi-head self-attention mechanism is to divide the input data into multiple parts, and each part has an independent attention head. Each attention head calculates the weights related to the input data, and then sums these weights weighted to obtain the final output result. For each feature , here represents , , , which represent the feature information of body temperature, heart rate, and brain waves respectively, and apply the multi-head attention mechanism. Suppose there are H heads, and each head has its own query, key, and value matrices, which are represented as Among them, is the index of the head. The query, key, and value matrix calculation formulas for each feature are as follows:
[0100] ;
[0101] Among them, Represents each feature Query matrix of Represents each feature Key matrix of Represents each feature Value matrix;
[0102] The calculation formula of the attention mechanism is as follows:
[0103] ;
[0104] Among them, Is a scaling factor used to prevent the inner product value from being too large, resulting in the disappearance of the gradient of the Softmax activation function; Represents the output obtained through the calculation of the attention mechanism, Represents The transpose matrix of the matrix.
[0105] After the calculation through the attention mechanism, the output of the attention mechanism is added to the input feature to form a residual connection:
[0106] ;
[0107] Among them, Represents the residual connection;
[0108] Perform layer normalization on the result of the residual connection to ensure the stable distribution of data at each level. The feed-forward neural network layer includes two linear transformations and a ReLU activation function for further processing features:
[0109] ;
[0110] Among them Is the weight matrix from the input layer to the hidden layer, responsible for mapping the input data x to the hidden layer, Is the weight matrix from the hidden layer to the output layer, mapping the output of the hidden layer to the final result, Is the bias term of the hidden layer, which is a constant term added to the calculation result of the neuron to adjust the value of the hidden layer output, Is the bias term of the output layer, which is a constant term added to the output calculation result to adjust the final output, Represents a feed-forward neural network, Is the input vector, which is the output from the previous layer or the input data of the model.
[0111] Perform the operations of residual connection and layer normalization again, and perform residual connection and layer normalization with the output of the feed-forward neural network layer:
[0112] ;
[0113] Among them, represents the information after being processed by the feedforward neural network, but it still retains the information of the original input x.
[0114] The processed features are input into the fully connected layer for final prediction:
[0115] ;
[0116] Among them, represents the output after further processing, which passes the feature input to the fully connected layer for final prediction; is the output feature of the previous layer, is the weight of the fully connected layer, is the bias of the fully connected layer.
[0117] The formula of the multi-head self-attention mechanism can be expressed in the following form:
[0118] ;
[0119] Among them , is the concatenated linear transformation matrix, represents the attention calculation performed on the th head to obtain the output of this head, represents the concatenation operation, which concatenates multiple vectors or matrices along the column direction to form a larger vector or matrix.
[0120] Concatenate the outputs after the multi-head attention mechanism processes body temperature, heart rate, and brain waves:
[0121] ;
[0122] represents the output after the multi-head attention mechanism processes the body temperature, heart rate, and brain wave feature information, , , respectively represent the output of this head obtained by the multi-head attention calculation of body temperature, heart rate, and brain waves.
[0123] The concatenated output is further processed and classified through the fully connected layer:
[0124] ;
[0125] Among them, and are the weights and biases of the fully connected layer, Is the output processed by the fully connected layer and the softmax activation function, and is usually used for the final prediction in classification tasks.
[0126] After multi-modal data fusion and feature processing, the feature vector is input into the fully connected layer for linear transformation. Assume the feature vector is The output of the fully connected layer is:
[0127] ;
[0128] Among them, is the weight matrix, is the bias vector, is the output of the fully connected layer.
[0129] The output of the fully connected layer is input into the Softmax function to calculate the probability of each class. The Softmax function is defined as follows:
[0130] ;
[0131] Among them, is the score of the m-th output, is the sum of the exponents of all output scores, represents the m-th element in the input vector of the Softmax function.
[0132] The output of the Softmax function is a probability distribution for classification tasks. In this case, assume there are K classes, corresponding to the range of TSV values. Map these probabilities to specific TSV values. Assume the range of TSV values is (-3, -2,..., 3), and there are a total of 7 classes.
[0133] To normalize the Softmax output to the specified TSV value range, first determine the output class. Through the probability distribution of the Softmax output, select the TSV value corresponding to the class with the highest probability. Assume the Softmax output is P, then:
[0134] ;
[0135] Among them, represents the index of the maximum value in the probability distribution P, that is, the class corresponding to the highest probability.
[0136] Normalize the probability values of the Softmax output to the range (-3, -2,..., 3). The specific steps are as follows:
[0137] Determine the TSV value range: Assume the range is , where K is the number of categories, and here v1 = -3, v2 = -2, …, v7 = 3.
[0138] Calculate the normalized TSV value:
[0139] ;
[0140] Where, is the probability of the nth category output by Softmax, is the corresponding TSV value.
[0141] S401: The TSV value output by the sleep thermal comfort posture detection module is , and the TSV value output by the multi-modal data fusion module is . The calculation formula for the weighted average TSV value of the two modules is:
[0142] ;
[0143] Where, is the weighted average TSV value of the two modules, is the weight of the STCE-Net model, is the weight of the MDF-Net model;
[0144] Through a large number of experimental tests and verifications, the weight ratios of the STCE-Net model and the MDF-Net model in the comprehensive calculation of the TSV value have been determined. Among them, the weight of the STCE-Net model is determined to be 0.4, while the weight of the MDF-Net model is determined to be 0.6. These weight values are obtained through experimental data analysis and optimization adjustment, which can effectively reflect the importance and contribution degree of each model in sleep thermal comfort perception, so as to ensure that the finally output TSV value is more accurate and reliable. The comprehensive TSV value will be used for the control decision-making of the subsequent sleep thermal comfort perception system and intelligent regulation system, providing a more accurate and comfortable sleep environment regulation scheme.
[0145] S402: Based on human thermal comfort, this paper comprehensively calculates the TSV value through the STCE-Net model and the MDF-Net model to detect human thermal comfort, and adjusts the operating parameters of indoor air conditioning equipment accordingly. In this way, the system can automatically adjust the indoor environment according to the real-time thermal comfort state of the human body to ensure the best comfort. This not only realizes precise thermal comfort adjustment, but also effectively improves the adaptability and convenience of the system. In addition, through intelligent regulation, the system can significantly save energy, avoid energy waste, and thus achieve more efficient energy utilization and environmental protection goals.
[0146] Embodiment 2. This embodiment provides a sleep thermal comfort perception and indoor sleep environment regulation device, including:
[0147] A data acquisition module, configured to acquire video data, physiological data, and environmental data of a subject during sleep measured in advance;
[0148] A data set production module, configured to perform data annotation, feature extraction, and data fusion on the video data, physiological data, and environmental data, and then produce a data set;
[0149] A prediction module, configured to input the produced data set into a pre-constructed sleep thermal comfort perception model to obtain a sleep thermal comfort perception prediction result;
[0150] A regulation module, configured to regulate the indoor sleep environment according to the obtained sleep thermal comfort perception prediction result.
[0151] For the specific function implementation of the above modules, refer to the relevant content in the method of Embodiment 1, which will not be elaborated here.
[0152] Embodiment 3. This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of Embodiment 1 are implemented.
[0153] Embodiment 4. This embodiment provides a computer device, including:
[0154] A memory, configured to store computer programs / instructions;
[0155] A processor, configured to execute the computer programs / instructions to implement the steps of the method described in any one of Embodiment 1.
[0156] Embodiment 5. This embodiment provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method described in any one of Embodiment 1 are implemented.
[0157] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
[0158] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0159] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to produce a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure rather than to limit the scope of its protection. Although the present disclosure has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: after reading the present disclosure, those skilled in the art can still make various changes, modifications, or equivalent replacements to the specific implementation manners of the invention, but these changes, modifications, or equivalent replacements are all within the scope of the claims of the pending disclosure.
Claims
1. A method for sensing thermal comfort during sleep and regulating an indoor sleeping environment, characterized in that: include: Obtain the video data, physiological data and environmental data of the subject during sleep that were measured in advance; including: Video data of the subjects sleeping was obtained through a thermal imaging dual-spectrum webcam and a visual camera; Physiological detection equipment is used to record the subject's physiological data during sleep, including the subject's body temperature, heart rate and brain wave data; An indoor thermal comfort measuring instrument is used to obtain indoor environmental data, including air temperature, relative humidity, air flow rate and carbon dioxide concentration; After data annotation, feature extraction and data fusion of video data, physiological data and environmental data, a data set is created; including: Extract video data into time series images at fixed intervals, and process and annotate the time series images into videos; Use computer vision technology to process the video-processed and annotated time series images to obtain images after human posture features are extracted; Synchronously record physiological data and normalize the physiological data; The images after human posture features are extracted, the physiological data after normalization, and the environmental data are fused to form a data set; Input the prepared data set into the pre-built sleep thermal comfort perception model to obtain the sleep thermal comfort perception prediction results; including: Input the prepared data set into a sleep thermal comfort perception model, wherein the sleep thermal comfort perception model includes a sleep thermal comfort posture detection module and a multimodal data fusion module, wherein the sleep thermal comfort posture detection module includes a human skeleton point detection module and a sleep comfort posture estimation module; In the sleeping thermal comfort posture detection module, perform the following steps: The human skeleton point detection module detects the human skeleton points of the images in the dataset and identifies images with human posture features; including: The images in the input dataset are initially downsampled through two convolutional layers with 3x3 convolution kernels and a stride of 2. Then, the initialization layer, transition structure, and stage structure are used for feature extraction and dynamic scale adjustment. Finally, the Transformer module and multi-scale fusion module are used for feature abstraction and fusion to detect the skeleton points of the subjects when they are sleeping and obtain images with human posture features. The image with human posture features is sent to the sleep comfort posture estimation module to identify and analyze the sleep thermal comfort posture, and obtain each sleep posture and its corresponding thermal sensation voting value ;include: The image with human posture features is extracted through multi-layer convolutional layers and pooling layers, and then a multi-level feature fusion strategy is used to integrate features at different levels. The Swin Transformer architecture is used to perform feature transformation to detect each sleeping posture and its corresponding thermal sensation voting value. ; In the multimodal data fusion module, the physiological data in the data set is detected to obtain each sleeping posture and its corresponding thermal sensation voting value. ;include: The physiological data in the dataset is input into the multimodal data fusion module, preliminarily processed by the LSTM module, and then feature fused by the multi-head attention mechanism. Finally, the fused features are input into the fully connected layer for final prediction or classification to obtain each sleeping posture and its corresponding thermal sensation voting value. ; Will , Perform weighted average processing and obtain the value as the final prediction result; According to the obtained sleep thermal comfort perception prediction results, the indoor sleeping environment is regulated. The calculation formula is as follows: ; in, is the weighted average of each sleeping posture and its corresponding thermal sensation voting value, is the weight of the sleeping thermal comfort posture detection module, is the weight of the multimodal data fusion module, and + = 1, where Determined to be 0.4, Determined to be 0.
6.
2. A device for sensing thermal comfort during sleep and regulating indoor sleeping environment, using the method for sensing thermal comfort during sleep and regulating indoor sleeping environment according to claim 1, characterized in that: include: A data acquisition module, used to acquire video data, physiological data and environmental data of the subject during sleep that are measured in advance; The data set production module is used to produce data sets after data annotation, feature extraction and data fusion of video data, physiological data and environmental data; A prediction module is used to input the prepared data set into a pre-built sleep thermal comfort perception model to obtain a sleep thermal comfort perception prediction result; The control module is used to control the indoor sleeping environment according to the obtained sleep thermal comfort perception prediction results.
3. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to claim 1 are implemented.
Citation Information
Patent Citations
Motion recognition method and device based on video image, equipment and storage medium
CN114511931A
Living space thermal environment adjusting method and system based on behavior activities
CN117628669A