A vital sign monitoring method and device based on deep learning

Through a deep learning-based vital sign monitoring device, combined with infrared perception and image acquisition module, using Mask-RCNN and transformer networks, contactless and highly integrated vital sign monitoring is achieved, solving the problems of single functions and poor integration of traditional equipment, and improving monitoring accuracy and prediction capabilities.

CN116098596BActive Publication Date: 2025-07-29BEIJING UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310119325.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-15
Publication Date
2025-07-29
Estimated Expiration
2043-02-15

AI Technical Summary

Technical Problem

Traditional vital sign monitoring equipment has a single function and is mostly contact-type, which is difficult to apply to certain scenarios, and has poor integration, so it is impossible to efficiently monitor important signs such as body temperature, pulse, blood pressure, and breathing at the same time.

Method used

The vital sign monitoring device based on deep learning is adopted, combined with infrared perception module, image acquisition module and calculation processing module, and real-time monitoring of body temperature, pulse, and breathing through deep learning algorithms. The Mask-RCNN algorithm is used to detect the facial and chest and abdominal areas, and the transformer network is used to analyze and predict sign information.

Benefits of technology

It realizes contactless and highly integrated vital sign monitoring, improves the accuracy and prediction ability of monitoring, and can promptly detect dangerous fluctuations and provide early warnings, reducing the pain to patients and the equipment space occupied.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116098596B_ABST
    Figure CN116098596B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for vital sign monitoring based on deep learning, including: a computing and processing module, an infrared sensing module, an image acquisition module, a user interaction module, and a communication and transmission module. The body temperature information is obtained through the infrared sensing module, and the user's facial image is obtained through the image acquisition module. The user's facial image and thoracic activity image are obtained from the facial image, and the above information is transmitted to the computing and processing module. The present invention not only gets rid of the limitation of the single function of the monitoring device, but also ensures the accuracy of vital sign monitoring. The monitoring algorithm uses the deep learning algorithm to monitor the three important vital sign parameters of body temperature, pulse, and respiration in real time, and analyzes the vital sign information of the above parameters through deep learning technology, timely discovers dangerous fluctuations and change trends, and gives early warnings of dangers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vital sign monitoring, and specifically to the technical field of vital sign monitoring based on deep learning and computer vision. Background Art

[0002] Traditional vital sign monitoring means are relatively single and mostly contact-based, requiring direct contact with the skin, which may be difficult to apply in some scenarios. For example, for patients with extensive burns, if a traditional electrocardiogram detector is used to detect the heart rate, the electrodes need to be attached to the patient's chest, which is very unfavorable for the skin rehabilitation of the patient and is likely to cause unnecessary pain to the patient. At the same time, the single function of the product often means that multiple instruments need to be used to detect multiple vital signs, further squeezing the limited space in the ward.

[0003] In the field of intensive care of patients, there are four crucial vital signs that need to be monitored in real time: body temperature, pulse, blood pressure, and respiration. Currently, these four signs are monitored in a one-to-one manner, with poor integration. CN112263228A proposes a non-contact health sign detection system that uses a mirror to monitor the signs of patients using traditional vision algorithms. The functions of multiple devices are integrated into one device.

[0004] In recent years, deep learning algorithms have achieved remarkable results in various fields, especially in the fields of security, autonomous driving, and medical imaging. The deep learning algorithms of computer vision have the characteristics of high accuracy and good robustness compared with traditional vision algorithms, bringing the accuracy of vision algorithms to a new height. Summary of the Invention

[0005] Therefore, the present invention proposes a vital sign monitoring method based on deep learning, which not only gets rid of the limitation of the single function of monitoring devices but also ensures the accuracy of vital sign monitoring. The monitoring algorithm uses deep learning algorithms to monitor the three important sign parameters of body temperature, pulse, and respiration in real time, and analyzes the vital sign information of the above parameters through deep learning technology to timely detect dangerous fluctuations and change trends and give warnings of dangers.

[0006] The technical solution adopted by the present invention is a vital sign monitoring device based on deep learning, including: a computing and processing module, an infrared sensing module, an image acquisition module, a user interaction module, and a communication and transmission module. The body temperature information is obtained through the infrared sensing module, and the user's facial image is obtained through the image acquisition module. The user's facial image and thoracic activity image are obtained from the facial image, and the above information is transmitted to the computing and processing module. The functions of each module are as follows: the computing and processing module obtains the body temperature information by obtaining infrared parameters, obtains the heart rate information through the facial image and corresponding algorithms, and obtains the breathing information through the thoracic image; the computing module combines the above information and analyzes the vital sign information in the time series through a deep learning algorithm to predict the upcoming danger. The user interaction module can modify the network parameters of deep learning, adjust the corresponding weights for different diseases, and focus on predicting various risks that are likely to occur in different diseases. The communication and transmission module is responsible for transmitting the infrared sensing module and image information to the computing module and transmitting the above vital sign information in real time.

[0007] The image acquisition module consists of an RBG camera. The RBG camera is installed directly above and on the side of the subject to obtain clear and complete upper body image videos and thoracic movement videos. At the same time, an infrared temperature measurement device is installed above the RBG camera to obtain the body surface temperature and ambient temperature; the collected image and temperature information are transmitted into the computing and processing module to analyze the current body temperature, pulse, and breathing vital signs, and the information is stored in the computer and displayed on the monitor in a visual form for analysis and diagnosis.

[0008] The monitoring method of the present invention includes using an RBG camera to obtain the upper body image of the user and using the infrared sensing module for thermal imaging at the same time;

[0009] The facial area and the chest and abdomen area are detected through a deep learning object detection algorithm.

[0010] The thermal imaging map of the infrared spectrum is matched with the forehead part of the detected facial image, and the average temperature of the forehead area is taken as the body surface temperature of the current frame. The temperature information is stored, and the difference between the current temperature and the temperature of the previous frame is calculated and divided by the time interval between the two frames to obtain the temperature change rate between the two frames for subsequent processing.

[0011] Considering that the absorption of light by human muscles, skin, and bones is basically constant under light, while the absorption of light by the arterial blood vessels under the skin changes constantly due to the change of blood volume, the method of obtaining the heart rate through time-domain signal analysis is to detect the pixel brightness change waveform of the measured skin position to obtain the pulse wave signal, use the independent component analysis method to separate the photoplethysmogram signal from the facial video, and transform the signal to the frequency domain through fast Fourier transform to extract the heart rate from it. The specific steps of the heart rate detection algorithm are as follows:

[0012] 1) Using deep learning algorithms, face detection is performed on the facial images obtained by the image acquisition module to obtain the facial information of each frame of the image. After obtaining the precise facial position of each frame, it is used for the next step of facial signal extraction.

[0013] 2) Taking the average brightness of the facial video as the observation value, the average brightness values of the RGB three channels at different times can be obtained respectively, and three facial observation signal graphs are drawn. Independent component analysis method is used to perform blind source separation on the three-channel signals, solve the three source signals, and separate the photoplethysmogram signal to avoid the interference of other physiological signals and noise signals in the video on heart rate detection.

[0014] 3) The source signal is converted into a frequency-domain signal through fast Fourier transform, and the image is described as a complex exponential sum of different amplitudes, frequencies, and phases. Through experiments, it is found that among the three source signals, the frequency corresponding to the maximum amplitude of one of the source signals is the heart rate signal that needs to be detected in real time.

[0015] During breathing, the chest cavity will fluctuate with inhalation and exhalation. The breathing state can be evaluated by observing the ups and downs of the chest cavity through computer vision. The specific process of the algorithm is as follows:

[0016] 1) The side image of the human body is obtained through a camera, and the chest position and the binary image of the chest area of each frame of the chest image are obtained by using the Mask-RCNN algorithm.

[0017] 2) Through the binary image of the chest cavity, calculate the number of pixels in the connected area of the chest cavity part, and count the chest cavity area at different times, and draw a signal graph.

[0018] 3) Through pre-calibration, obtain how much actual area each pixel represents (mm / pixel), and obtain the signal graph of the chest cavity area changing with time.

[0019] 4) The fluctuation of breathing can be directly seen from the graph, and at the same time, the breathing signal can also be used as an important vital sign signal for subsequent vital sign monitoring.

[0020] Therefore, the above-mentioned vital signs are combined into a one-dimensional vector, which includes: current heart rate speed, heart rate intensity, heart rate rhythm, breathing frequency, breathing depth, breathing pattern, inhalation rhythm, body temperature, body temperature change rate, etc. Feature fusion is performed through one-dimensional convolution, and then the features are sent into the transformer network for classification of the current state.

[0021] In the process of feature extraction, the input features are designed according to the sampling frequency of the current RGB camera. For example, the image of the current frame can be processed to obtain vital sign information, and then combined with the one-dimensional vectors obtained from the previous 199 vital signs to generate a two-dimensional tensor. The two-dimensional tensor is subjected to one-dimensional convolution, which can fuse the features in different time periods to extract high-level semantic features, and then input into the transformer for further feature fusion and output of the prediction results.

[0022] The transformer consists of an encoder and a decoder. The encoder encodes 50 one-dimensional vectors step by step. The finally output feature vector synthesizes the feature information of the 50 vectors and inputs the feature vector into the decoder to output the predicted values corresponding to 50 states. This algorithm encodes 50 vectors in the time series, comprehensively considers the information in different time periods, and improves the prediction accuracy.

[0023] Compared with the prior art, while retaining the advantages of highly integrated non-contact vital sign monitoring, the accuracy of vital sign monitoring is further improved by using a deep neural network. The deep neural network has excellent adaptive characteristics and can self-learn through different data sets. Therefore, it can be applied to the vital sign monitoring of each department. Just according to the user's disease condition, let the neural network learn on the data set of this type of disease, and the neural network can learn how to diagnose and monitor this type of disease, avoiding the cumbersome work of manual monitoring and manual design of monitoring algorithms. Brief Description of the Drawings

[0024] Figure 1 Composition of the vital sign monitoring system device.

[0025] Figure 2 Schematic diagram of the scene layout of the vital sign monitoring device.

[0026] Figure 3 Overall flowchart of the vital sign monitoring technology.

[0027] Figure 4 Flowchart of the respiration monitoring technology solution.

[0028] Figure 5 Flowchart of the heart rate monitoring technology solution. Detailed Implementation Manner

[0029] The device of the present invention includes: a computing and processing module, an infrared sensing module, an image acquisition module, a user interaction module, a communication and transmission module, as Figure 1As shown in the figure. Among them, the main component of the image acquisition module is an RGB camera. The camera is placed directly above the user to obtain a clear and complete upper body image video. At the same time, an infrared temperature measurement device is placed above the camera to obtain the body surface temperature and ambient temperature of the user. The collected image and temperature information are transmitted to the calculation and processing module to analyze the current body temperature, pulse, and respiratory vital signs, and the information is stored in the computer and displayed on the monitor in a visual form for further analysis and diagnosis by doctors. When an emergency occurs in the vital signs, the communication and transmission module will send an alarm message to the on-duty doctors and nurses to prompt the doctor to perform an emergency diagnosis and care. The camera layout is as Figure 2 shown.

[0030] The monitoring method of the present invention is as Figure 3 shown, including

[0031] Using a camera to obtain the upper body image of the user, and at the same time using an infrared sensing module for thermal imaging;

[0032] Detecting the facial area and chest and abdomen area through a deep learning object detection algorithm.

[0033] Matching the thermal imaging map of the infrared spectrum with the forehead part of the detected facial image, taking the average temperature of the forehead area as the body surface temperature of the current frame, storing the temperature information, and calculating the difference between the current temperature and the temperature of the previous frame, dividing by the time interval between the two frames to obtain the temperature change rate between the two frames for subsequent processing.

[0034] Considering that the absorption of light by human muscles, skin, and bones is basically constant under light, while the absorption of light by the arterial blood vessels under the skin changes constantly due to the change of blood volume, the method of obtaining the heart rate through time-domain signal analysis is to detect the pixel brightness change waveform at the measured skin position and obtain the pulse wave signal through independent component analysis and Fourier transform.

[0035] The object detection algorithm uses the Mask-RCNN algorithm. Compared with ordinary object detection algorithms, while outputting the position coordinates of the facial area and chest area, it can also output the Mask (mask) of the entire image for region segmentation of the entire image. By calculating the area size of the chest area in the segmented image and performing time series analysis, information such as the respiratory frequency and depth can be obtained. The process is as Figure 4 shown.

[0036] Specifically, the skin color change caused by blood volume change is very weak and difficult to observe directly. Therefore, the Euler magnification algorithm is adopted to magnify the color change of the face in the video, making the color change of the face visible to the human eye. When magnifying the color change, only Gaussian pyramid decomposition is performed on the video image during the spatial domain processing. During the time domain filtering process, an ideal band-pass filter is selected. Since the frequencies of the heart rate and respiratory rate are different, the passband frequencies are also different. The standard heart rate of adults is between 60 and 100 beats per minute. When the heart rate exceeds 100 beats per minute, it is tachycardia; when the heart rate is below 60 beats per minute, it is bradycardia. When the heart rate of an average person is between 40 and 50 beats per minute, symptoms such as chest tightness and dizziness will occur; when the heart rate is between 35 and 40 beats per minute, the blood supply to the heart and brain organs will be affected, and there will be life-threatening risks. Therefore, the passband frequency is selected to be 0.58 - 2 Hz, corresponding to a heart rate of 34 - 120 beats per minute.

[0037] Specifically, during the heart rate detection process, the image has three RGB channels, and each channel is affected by noise to a different degree. Therefore, the Mask-RCNN algorithm is first used to detect the face position, and then the average brightness values of the R, G, and B channels in the face area of the video sequence are calculated respectively. Then, the photoplethysmogram signal is separated from the three color signals by using independent component analysis. Finally, experiments prove that the heart rate signal of the G channel has the best effect, and the heart rate is obtained by transforming the pulse wave signal to the frequency domain using the fast Fourier transform. The process is as Figure 5 shown.

[0038] Each of the above-mentioned heart rate, respiration, and temperature information has a corresponding threshold. When the heart rate, respiratory rate or depth, or temperature exceeds the threshold, an alarm message will be sent to the medical staff to notify the medical staff to conduct inspections and rescues.

[0039] Considering that the condition of critically ill patients deteriorates very rapidly in case of emergencies, usually when an emergency is triggered, there is a high probability that they cannot be rescued. Before the condition deteriorates sharply, clues can often be seen from the previous vital signs. Therefore, the vital sign information in the normal state can be analyzed and predicted. The vital sign information in the time series is analyzed through a neural network, and the health status of the user is predicted by observing some change trends and abnormal vital sign signs in the past time period. For this purpose, the above-mentioned vital signs are combined into a one-dimensional vector, feature fusion is performed through one-dimensional convolution, and then the features are sent into the transformer network for classification of the current state.

[0040] Specifically, during the feature extraction process, the input features are designed according to the sampling frequency of the current camera. For example, the image of the current frame can be processed to obtain vital sign information, which is then combined with the one-dimensional vectors obtained from the previous 199 vital signs to generate a two-dimensional tensor. The two-dimensional tensor is then subjected to one-dimensional convolution, which can fuse features from different time periods to extract high-level semantic features, and then input into the transformer for further feature fusion and output of the prediction results.

[0041] Specifically, the transformer consists of an encoder and a decoder. The encoder gradually encodes 50 one-dimensional vectors, and the finally output feature vector synthesizes the feature information of the 50 vectors and inputs the feature vector into the decoder to output the predicted values corresponding to 50 states. This algorithm encodes 50 vectors in the time series, comprehensively considers the information in different time periods, and improves the prediction accuracy.

[0042] Specifically, the above-mentioned one-dimensional vital sign vectors can be designed according to the required included features, such as heart rate speed, heart rate intensity, heart rate rhythm, respiratory rate, respiratory depth, respiratory pattern, inspiratory rhythm, body temperature, body temperature change rate, etc.

[0043] In this embodiment, the cameras are arranged on the side and above the patient to monitor the patient from multiple angles and improve the accuracy of vital sign monitoring. The front camera is used to capture the upper body image of the front, and the side camera is mainly used to observe the movement changes of the chest and abdomen during breathing. For the obtained video clips, they can be adjusted according to the performance of the computing device. Devices with superior graphics processing unit (GPU) performance can analyze frame by frame, and devices with general performance can perform frame skipping analysis. The images extracted from the video are subjected to subsequent processing.

[0044] In this example, the Mask-RCNN network is used for the object detection algorithm. Based on the faster-RCNN object detection algorithm, the Mask network is added to the Mask-RCNN, enabling it to handle not only object detection problems but also semantic segmentation problems.

[0045] In this example, when using the Mask-RCNN network, the network needs to be trained first. To improve the training efficiency, the pre-trained weights of the coco dataset are used. These weights are trained on the coco dataset, avoiding the disadvantages of getting stuck in local extreme points and slow training speed caused by randomly initializing the weights for training.

[0046] In this example, training is required on a self-built dataset. The self-built dataset contains 1500 images of humans lying flat. Use labelme to perform rectangular box annotation on the human face area and chest area in the pictures, generate a.json file, which records the position information and categories of each marked point. Then, this json file needs to be converted into a mask file in png format. Organize the 1500 image data into a form recognized by the model, including four folders. The folder cv2_mas stores the generated png format label files, the folder json stores the json files generated by labelme, the folder labelme_json stores the.yaml files converted from json files, and the folder pic stores the original images after size normalization.

[0047] In this example, the images are input into the Mask-RCNN network in sequence. To improve efficiency, multiple images are subjected to object detection and semantic segmentation at once. Set the batch-size to 16, that is, train 16 images at a time. Set the initial learning rate to 0.001, and add the learning rate exponential decay algorithm to make the learning rate smaller in the later stage of training to avoid parameter oscillation. The Mask-RCNN network outputs the position information and classification information of the rectangular boxes in the face and chest areas, and outputs a segmentation mask image of the same size as the input image.

[0048] In this example, the optimizer Adam is added to Mask-RCNN, introducing the first-order & second-order momentum. The gradient is updated through the accumulated first-order and second-order momentum, making the parameters more stable when Mask-RCNN finishes training an image of a batch-size.

[0049] In this example, use a thermal imager to convert the invisible infrared energy emitted by an object into a visible thermal image, and perform position matching with the face segmentation part output by Mask-RCNN. Take the average temperature at the forehead of the face as the current body surface temperature, and record the current forehead temperature, the temperature difference between the forehead and the background, and the current temperature change rate. The change rate is calculated as the current temperature minus the temperature of the previous frame and then divided by the time difference between two frames. A total of three parameters are recorded.

[0050] In this example, calculate the area of the chest segmentation image output by Mask-RCNN. Obtain the respiratory rate by calculating the time difference between the peak times of the changing chest area. Record the maximum area and the minimum area as the depths of exhalation and inhalation, and record the area of the current frame. A total of four parameters are recorded.

[0051] In this example, the original image is subjected to the Euler magnification algorithm to amplify the effect of skin color changes caused by blood volume changes. The average value of the G channel of the current image is calculated. Heart rate parameters are obtained through information such as the heart rate frequency, heart rate intensity, and RGB pixel change rate of the Fourier spectrogram. Information such as the speed, rhythm, strength, and weakness of the heart rate is used as important vital sign parameters.

[0052] In this example, the above 11 parameters are used as a one-dimensional vector, and a 200*11 matrix is constructed. The vector in the last row of this matrix is the vector of the current frame. For each row going forward, the frame number goes back 20 seconds. Therefore, this matrix stores vital sign information for approximately one hour and can also be adjusted at intervals according to needs to analyze vital sign information for different time lengths.

[0053] In this example, the above 200*11 matrix is subjected to one-dimensional convolution to extract high-level semantic features and input into the transformer network. This network can fuse semantic information from different time periods for state analysis in the time series. Finally, a probability distribution is output. This probability distribution describes the probability values of different states, which are divided into eight states: sepsis, stress ulcer, acute pulmonary edema, acute respiratory distress syndrome, acute kidney injury, superior mesenteric artery syndrome, heart failure, and electrolyte disorder. The prediction in this example is the probability distribution designed for severely burned patients. The present invention includes but is not limited to this symptom, and the final output layer can be designed according to the actual situation of different patients.

Claims

1. A vital sign monitoring device based on deep learning, characterized in that, Including: A calculation and processing module, an infrared sensing module, an image acquisition module, a user interaction module, and a communication and transmission module; Obtain body temperature information through the infrared sensing module, obtain the user's facial image through the image acquisition module, obtain the user's facial image and chest activity image from the facial image, and transmit the above information to the calculation and processing module; The functions of each module are as follows: The calculation and processing module obtains body temperature information by obtaining infrared parameters, obtains heart rate information through the facial image and corresponding algorithms, and obtains breathing information through the chest image; The calculation module combines the above information and analyzes the physical sign information in the time series through deep learning algorithms to predict the upcoming dangers; The user interaction module can modify the network parameters of deep learning, adjust the corresponding weights for different diseases, and predict various risks that are likely to occur in different diseases. The communication and transmission module is responsible for transmitting the infrared sensing module and image information to the calculation module and transmitting the physical sign information in real time; The monitoring method of the device includes using an RBG camera to obtain the upper body image of the user, and at the same time using the infrared sensing module for thermal imaging; Detect the facial area and the chest and abdomen area through deep learning object detection algorithms; Match the thermal imaging map of the infrared spectrum with the forehead part of the detected facial image, take the average temperature of the forehead area as the body surface temperature of the current frame, store the temperature information, and calculate the difference between the current temperature and the temperature of the previous frame, divide it by the time interval between the two frames to obtain the temperature change rate between the two frames for subsequent processing; The method of obtaining the heart rate through time-domain signal analysis is to detect the pixel brightness change waveform of the measured skin position to obtain the pulse wave signal, use the independent component analysis method to separate the photoplethysmogram signal from the facial video, and transform the signal to the frequency domain through fast Fourier transform to extract the heart rate; The specific steps of the heart rate detection algorithm are as follows: Use deep learning algorithms to perform face detection on the facial images obtained by the image acquisition module to obtain the facial information of each frame of the image; After obtaining the accurate facial position of each frame, it is used for the next step of facial signal extraction; Take the average brightness of the facial video as the observation value, obtain the average brightness values of the RGB three channels at different times respectively, and draw three facial observation signal graphs. Use the independent component analysis method to perform blind source separation on the three-channel signals, solve the three source signals, and separate the photoplethysmogram signal to avoid interference from other physiological signals and noise signals in the video on heart rate detection; Convert the source signal to a frequency-domain signal through fast Fourier transform, and describe the image as a complex exponential sum of different amplitudes, frequencies, and phases; Through experiments, it is found that among the three source signals, the frequency corresponding to the maximum amplitude of one of the source signals is the heart rate signal that needs to be detected in real time.

2. The vital sign monitoring device based on deep learning according to claim 1, wherein The image acquisition module consists of an RGB camera. The RGB camera is placed directly above and on the side of the subject to obtain clear and complete upper body image videos and chest fluctuation videos. At the same time, an infrared temperature measurement device is placed above the RGB camera to obtain the body surface temperature and the ambient temperature. The collected images and temperature information are transmitted to the calculation and processing module to analyze the current body temperature, pulse, and respiratory vital signs, and the information is stored in the computer and displayed on the monitor in a visual form.

3. The vital sign monitoring device based on deep learning according to claim 1, characterized in that, During breathing, the chest fluctuates with inhalation and exhalation. The breathing state is evaluated by observing the chest fluctuations through computer vision. The specific process is as follows: The side image of the human body is obtained through the camera, and the chest position and the binary image of the chest area of each frame of the chest image are obtained using the Mask-RCNN algorithm. Based on the binary image of the chest, calculate the number of pixels in the connected area of the chest part, and count the chest area at different times to draw a signal graph. Through pre-calibration, obtain the actual area represented by each pixel in mm / pixel, and obtain the signal graph of the chest area changing with time. The breathing fluctuation is seen from the signal graph, and at the same time, the breathing signal is used as an important vital sign signal for subsequent vital sign monitoring.

4. A vital sign monitoring device based on deep learning according to claim 1, characterized in that, During the feature extraction process, according to the sampling frequency of the current RGB camera, the input features are designed. The image of the current frame is processed to obtain the vital sign information, and then combined with the one-dimensional vector obtained from the previous 199 vital signs to generate a two-dimensional tensor. The two-dimensional tensor is subjected to one-dimensional convolution. One-dimensional convolution can fuse the features in different time periods to extract high-level semantic features, and then input them into the transformer for further feature fusion and output the prediction result. The transformer consists of an encoder and a decoder. The encoder encodes 50 one-dimensional vectors step by step. The finally output feature vector synthesizes the feature information of the 50 vectors and inputs the feature vector into the decoder to output the predicted values corresponding to 50 states. This algorithm encodes 50 vectors in the time series, comprehensively considers the information in different time periods, and improves the prediction accuracy.

Citation Information

Patent Citations

  • Mirror and non-contact health sign detection system

    CN112263228A

  • Low-cost physical sign monitoring method and system based on deep learning

    CN115512178A

  • A method for stabilizing vital sign measurements using parametric facial appearance models via remote sensors

    EP2960862A1