Suffocation risk assessment system and method based on infant milk regurgitation detection
The infant regurgitation detection system, which utilizes dual-spectrum collaborative detection and deep learning, solves the problem of accurately identifying infants with obstructed mouths and noses or liquid coverings. It achieves high-precision and rapid suffocation risk assessment, reduces false alarm rates, and provides 24/7 monitoring.
Patent Information
- Application Number
- CN202511003496.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-11
AI Technical Summary
Existing infant monitoring technologies cannot effectively identify mouth and nose obstruction or liquid coverage, leading to inaccurate suffocation risk assessments. They also cannot work effectively at night or in low-light conditions, and existing systems have high false alarm rates and slow response times.
Employing a dual-spectrum collaborative detection mechanism that combines visible light and thermal imaging, sub-pixel-level registration and dynamic temporal modeling are achieved through multimodal fusion and deep learning. Visible light is used to provide high-precision mouth and nose localization, while thermal imaging captures temperature changes. A dual-branch network is constructed to assess the risk of suffocation.
It significantly improves the accuracy and reliability of monitoring, reduces the false alarm rate by more than 40%, achieves all-weather monitoring, has an average early warning response time of less than 0.3 seconds, and provides all-weather protection without blind spots.
Smart Images

Figure CN120932361A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring technology, and in particular to a suffocation risk assessment system and method based on infant regurgitation detection. Background Technology
[0002] Spitting up is a common physiological phenomenon in newborns to two-year-old infants, especially during sleep. If spit-up is not cleared promptly, it can be aspirated into the lungs, causing choking, suffocation, or even serious respiratory infections, endangering life. Clinical data shows that accidental injuries caused by vomit aspiration account for a significant proportion of unnatural infant deaths, especially at night or when caregivers are not paying close attention. Since infants lack the ability to roll over or clear foreign objects from their mouth and nose, real-time monitoring of the mouth and nose area during sleep is crucial. However, achieving this goal faces numerous technical and application challenges. Existing infant monitoring technologies have several limitations: 1. Contact-based monitoring devices, such as breathing belts and chest / abdominal sensors, while capable of acquiring physiological parameters like heart rate and respiratory rate, require direct contact with the infant's body, affecting comfort and being prone to falling off, making them unsuitable for prolonged use. Furthermore, these devices cannot detect whether the mouth and nose are obstructed or covered by liquid. 2. Visible light camera-based visual monitoring systems rely on ambient lighting conditions, exhibiting poor recognition performance at night or in low-light environments. They also have low accuracy when facing light-colored coverings (such as thin gauze or milk stains), easily leading to missed detections. 3. Sound detection systems analyze crying or coughing sounds for anomaly detection, but suffer from significant environmental noise interference and only trigger alarms after an abnormal event occurs, lacking early warning capabilities and exhibiting delayed response. 4. Single-modal sensing systems: Most products use a single sensor (e.g., only a camera or only a temperature sensor), resulting in limited data dimensions and a lack of multi-source information fusion mechanisms, leading to high false alarm rates, poor stability, and difficulty adapting to complex monitoring scenarios.
[0003] In addition, infant asphyxia is one of the leading causes of accidental infant death, especially incidents occurring during sleep. According to statistics from the World Health Organization, tens of thousands of infants worldwide die each year from sleep-related asphyxia accidents. These accidents often result from blankets, pillows, or other soft objects accidentally covering the mouth and nose, obstructing breathing. Infant care is both important and meticulous, requiring parents to constantly monitor their baby's physical condition and behavioral changes. Because newborns' hearts and lungs are not yet fully developed, clothing, blankets, and other coverings can pose a suffocation risk, making constant monitoring of the baby's condition particularly crucial. However, parents or caregivers cannot provide continuous 24 / 7 care. Therefore, how to effectively monitor an infant's condition over extended periods and issue timely warnings in cases of potential asphyxia has become a pressing issue. Achieving this goal faces numerous technical and application challenges in practice.
[0004] Chinese patent CN119007237A discloses a method, device, equipment, and storage medium for real-time identification of dangerous sleeping positions in infants and toddlers. The method includes: acquiring at least one set of visible light images and thermal infrared images of the sleeping state of a human being to be tested; performing target detection on the visible light images to detect first position information corresponding to the infant's head; mapping the first position information to the real-time infrared images to output second position information of the infant in the infrared images; analyzing the thermal infrared data within the infant's head area based on the second position information, and issuing a safety alarm when the infant is identified as being in a dangerous sleeping position; extracting features from pixels within the infant's head area based on the target position information, and outputting target feature information related to the infant's sleeping position; calculating a comprehensive sleeping position judgment value based on the target feature information, and identifying the dangerous sleeping position using the comprehensive sleeping position judgment value; and mapping the first position information to the real-time infrared images of the infant's head. Feature extraction is performed on pixels within a certain region; the area of the infant's head region in the thermal infrared image is calculated, and the rate of change of the infant's head region area between multiple frames of real-time thermal infrared images is obtained; the target feature information is determined based on the statistical feature information and the rate of change of area; a comprehensive sleeping posture judgment value is calculated based on the target feature information, and the dangerous sleeping posture is identified through the comprehensive sleeping posture judgment value; it mainly judges whether there is an occlusion problem based on the rate of change of thermal imaging area, but this also has some possible false alarm problems. For example, if the infant's arm covers the head or the blanket covers the face and forehead, it will be judged as occlusion, but this kind of occlusion will not have any impact on the infant. This scheme is too sensitive and is prone to false alarms; judging whether the infant is in a suffocation risk due to occlusion is mainly based on whether the mouth and nose are covered; therefore, it is important to accurately judge whether the infant is actually covered, and reducing false alarms will reduce the burden on caregivers.
[0005] Based on this, this application proposes a suffocation risk assessment system and method based on infant spitting up detection to solve the problems existing in the prior art. Summary of the Invention
[0006] The purpose of this invention is to provide a suffocation risk assessment system and method based on infant regurgitation detection.
[0007] To achieve the above objectives, the present invention provides the following solution:
[0008] This invention provides a method for assessing the risk of suffocation based on infant regurgitation detection, comprising the following steps:
[0009] S1. Acquire visible light and thermal images of the same monitoring scene;
[0010] S2. Perform extrinsic and intrinsic parameter corrections on the acquired visible light and thermal imaging images, and complete the viewing angle registration to obtain the registration matrix H;
[0011] S3. In the visible light image, the human body detection model is used to locate the infant human body target and obtain the infant bounding box; within the infant bounding box, the coordinates of the infant's mouth and nose region are obtained through face detection and facial key point detection;
[0012] S4. Based on the registration matrix H obtained in step S2 and the mouth and nose region coordinates obtained in step S3, map the mouth and nose region coordinates onto the thermal imaging image and extract the corresponding mouth and nose thermal sub-region ROI.
[0013] S5. Perform continuous frame statistics and time series feature construction on the ROI region to obtain the temperature distribution vector sequence F. t Temperature-color temporal characteristics G t ; F t Substituting the data into a pre-trained infant occlusion risk assessment model, the occlusion risk probability is obtained, and G is... t Substitute the pre-trained infant spitting up assessment model to obtain the probability of spitting up. Integrate the probability of occlusion risk and the probability of spitting up to obtain the choking risk probability P.
[0014] S6. If the probability exceeds the corresponding threshold, an alarm signal is triggered.
[0015] Preferably, in step S1, the visible light image is acquired using two coaxially mounted visible light cameras with an optical center deviation of less than 0.5 mm and a field-of-view matching error of less than 1 degree. The visible light cameras use a 2-megapixel CMOS sensor, support 1080P@30fps, and are equipped with an 850nm invisible thermal imaging fill light. The thermal imaging image is acquired using a thermal imaging camera with a resolution of not less than 320×240 and a thermal sensitivity of <50mK. A hardware trigger signal ensures that the time deviation between the acquisition of the visible light image and the thermal imaging image is less than 1m.
[0016] Preferably, step S2 includes:
[0017] S21. Perform image distortion correction on the images acquired by the visible light camera and the thermal imaging camera to obtain distorted visible light images and distorted thermal imaging images;
[0018] S22. Perform feature extraction and calibration parameter calculation on the distortion-free visible light image and the distortion-free thermal imaging image to obtain the registration matrix H from the visible light image to the thermal imaging image.
[0019] Preferably, step S3 includes:
[0020] S31. Construct and train a human detection model, a facial localization model, and a facial landmark localization model;
[0021] S32. Use the human detection model YOLOv8n as the detection model to locate the bounding box of the baby's body, detect whether the baby appears, and if it does, crop the image of that region.
[0022] S33. Use a facial localization model to detect the infant's facial region within the captured infant bounding box, and use a facial keypoint localization model to locate the coordinates of the infant's mouth and nose within the facial region.
[0023] Preferably, step S4 includes:
[0024] S41. Map the visible light mouth and nose region coordinates back to the thermal imaging region coordinates according to the registration matrix H;
[0025] S42. Extract the thermal imaging area image based on the coordinates of the thermal imaging area.
[0026] Preferably, step S5 includes:
[0027] S51. Construct and train an infant suffocation risk assessment model and an infant regurgitation assessment model;
[0028] S52. Acquire temporal continuous frame features of the oral and nasal thermal sub-region F t This includes inter-frame mean temperature difference, local binarization pattern temporal texture descriptor, and thermal diffusion gradient; it also collects temporal continuous frame features G of the mouth and nose region. t This includes the average temperature across consecutive frames and the temperature drop gradient, as well as the increments in brightness and saturation of the visible color gamut.
[0029] S53. F t Substituting the data into the infant occlusion risk assessment model, we obtain the occlusion risk probability, and then assign G... t Substitute the data into the infant spitting-up assessment model to obtain the probability of spitting-up occurring;
[0030] S54. Integrate the probability of suffocation risk and the probability of vomiting to obtain the probability of choking risk P.
[0031] Preferably, the infant occlusion risk assessment model adopts a two-branch CNN-LSTM model, and the infant regurgitation assessment model adopts a two-branch CNN-Transformer model.
[0032] Preferably, the training method of the infant occlusion risk assessment model is as follows: collect and label the oral and nasal thermal subregion sequence data, with labels including three categories: occluded, partially occluded, and uncovered; use cross-entropy and focus loss function to jointly train the model; freeze the network weights and derive the inference model when the validation AUC reaches 0.95 or higher.
[0033] Preferably, the training method of the infant spitting up assessment model is as follows: collect and label the sequence data of the mouth and nose area, with labels including two categories: spitting up and not spitting up; use the focus loss function to jointly train the model; freeze the network weights and derive the inference model when the validation set F1 reaches 0.92 or higher.
[0034] This invention also provides a choking risk assessment system based on infant spitting up detection, comprising:
[0035] The image acquisition module includes a visible light camera and a thermal imaging camera, used to acquire an infant video stream and decompose the video stream into multiple frames of images, including visible light images and thermal imaging images;
[0036] The edge computing unit includes a calibration and registration module, a detection and positioning module, a risk assessment module, and an alarm module.
[0037] The correction and registration module is used to perform image distortion correction on images acquired by the visible light camera and the thermal imaging camera to obtain distorted visible light images and distorted thermal imaging images; and to perform feature extraction and calibration parameter calculation on the distorted visible light images and distorted thermal imaging images to obtain the registration matrix H from the visible light image to the thermal imaging image.
[0038] The detection and localization module consists of multiple deep learning models. First, it locates the infant's body area, then it locates the infant's face area through the body area, and finally it locates the coordinates of the mouth and nose area through the infant's face area.
[0039] The risk assessment module is used to collect temporal continuous frame features F of the oral and nasal thermal sub-region. t This includes inter-frame mean temperature difference, local binarization pattern temporal texture descriptor, and thermal diffusion gradient; it also collects temporal continuous frame features G of the mouth and nose region. t This includes the average temperature of consecutive frames and the temperature drop gradient, as well as the increments in brightness and saturation of the visible color gamut; F tSubstituting the data into the infant occlusion risk assessment model, we obtain the occlusion risk probability, and then assign G... t Substitute the data into the infant regurgitation assessment model to obtain the probability of regurgitation; integrate the probability of occlusion risk and the probability of regurgitation to obtain the suffocation risk probability P.
[0040] The alarm module is used to issue an early warning when the probability of suffocation exceeds a certain threshold, reminding caregivers to intervene.
[0041] The present invention achieves the following beneficial technical effects compared to the prior art:
[0042] This invention provides a suffocation risk assessment system and method based on infant regurgitation detection. Through multimodal fusion and deep learning collaborative optimization, it significantly improves the accuracy and reliability of monitoring. The dual-spectrum collaborative detection mechanism combines visible light and thermal imaging, and uses sub-pixel-level registration to achieve all-weather monitoring: visible light provides high-precision mouth and nose positioning, while thermal imaging captures respiratory thermal field anomalies caused by sudden temperature drops and occlusion, completely solving the problems of failure in low-light environments and detection of transparent objects. Dynamic temporal modeling analyzes temperature gradients and HSV color gamut mutations through continuous frame analysis, and fuses spatiotemporal features through a dual-branch network, improving the accuracy of regurgitation detection to 98%, achieving an AUC > 0.95 for suffocation risk assessment, and reducing the false alarm rate by more than 40%. The system-level optimized design includes coaxial dual-camera hardware synchronization, dynamic threshold adjustment, and multi-level alarm verification, achieving a 0.3-second response speed, providing contactless and full-coverage safety protection for infants and young children.
[0043] The regurgitation risk prediction method combines high-resolution morphological information from visible light images with body temperature distribution characteristics from thermal infrared images. Employing a multimodal data fusion approach, it comprehensively optimizes the detection of infants' mouth and nose occlusion, significantly improving detection accuracy. Thermal infrared images can accurately capture the temperature distribution characteristics of the infant's mouth and nose area. When regurgitation occurs, abnormal changes in the temperature field can effectively distinguish between regurgitation and non-regurgitation states, solving the problem of difficulty in liquid detection using traditional purely visual methods. Through the complementary information from visible light and thermal infrared images, a more comprehensive monitoring dimension is provided: visible light images provide high-precision spatial positioning information and extract incremental information on brightness and saturation from the HSV color gamut, ensuring the accuracy of mouth and nose area detection; thermal infrared images... The system provides reliable temperature characteristics, unaffected by lighting conditions, and the fusion of these two technologies significantly improves the system's robustness. An innovative spatiotemporal temperature feature analysis method judges the state of the mouth and nose through continuous multi-frame temperature change trends, enabling the identification of some critical situations such as vomiting and achieving early warning. Compared to single-frame analysis, this significantly reduces the false negative rate. A cascaded detection architecture is employed, first locating the infant as a whole, then precisely locating the mouth and nose area, and finally analyzing thermodynamic characteristics. This layered processing effectively eliminates interference from other heat sources and improves the specificity of the detection. The dynamic risk assessment mechanism comprehensively considers factors such as vomiting and its duration, achieving graded warnings through probabilistic output. This ensures timely alerts for high-risk situations while avoiding excessive interference with caregivers. In summary, this invention, through technological innovations such as multimodal data fusion, precise coordinate mapping, temporal feature analysis, and intelligent risk assessment, achieves a comprehensive improvement in the accuracy, real-time performance, and reliability of infant vomiting detection, providing an effective technical means to prevent infant vomiting and choking. The system performs exceptionally well in various real-world application scenarios: it maintains a detection accuracy rate of over 98% even in complete darkness; and its average warning response time is less than 0.3 seconds, significantly better than existing solutions. These advantages enable this invention to truly provide infants and young children with all-weather, comprehensive protection.
[0044] The occlusion risk prediction method combines high-resolution morphological information from visible light images with body temperature distribution characteristics from thermal infrared images. Employing a multimodal data fusion approach, it comprehensively optimizes the detection of occlusion in infants and young children, significantly improving detection accuracy. Thermal infrared images can accurately capture the temperature distribution characteristics of the infant's mouth and nose area. When occlusion occurs, abnormal changes in the temperature field can effectively distinguish between occluded and uncovered states, solving the problem of traditional pure visual methods' difficulty in detecting transparent or light-colored occlusions. Through the complementary information from visible light and thermal infrared images, a more comprehensive monitoring dimension is provided: visible light images provide high-precision spatial positioning information, ensuring the accuracy of mouth and nose area detection; thermal infrared images provide reliable... Temperature characteristics, unaffected by lighting conditions, and their fusion significantly enhance the system's robustness. An innovative spatiotemporal temperature feature analysis method uses continuous multi-frame temperature change trends to determine respiratory status, identifying critical situations such as partial occlusion and enabling early warning. This significantly reduces the false negative rate compared to single-frame analysis. A cascaded detection architecture first locates the infant as a whole, then precisely locates the mouth and nose area, and finally analyzes thermodynamic characteristics. This layered processing effectively eliminates interference from other heat sources, improving detection specificity. A dynamic risk assessment mechanism comprehensively considers factors such as occlusion degree and duration, achieving tiered warnings through probabilistic output. This ensures timely alerts for high-risk situations while avoiding excessive interference with caregivers. In summary, this invention, through technological innovations such as multimodal data fusion, precise coordinate mapping, temporal feature analysis, and intelligent risk assessment, achieves a comprehensive improvement in the accuracy, real-time performance, and reliability of infant mouth and nose occlusion detection, providing an effective technical means to prevent infant suffocation risks. The system performs exceptionally well in various real-world application scenarios: maintaining a detection accuracy rate of over 98% even in complete darkness; achieving a 95% recognition rate for coverings such as thin gauze that are difficult to detect using traditional methods; and an average warning response time of less than 0.3 seconds, significantly superior to existing solutions. These advantages enable this invention to truly provide infants and young children with all-weather, comprehensive protection, effectively reducing the risk of accidental suffocation. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 The flowchart of the suffocation risk assessment method based on infant regurgitation detection provided by the present invention is shown below. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] The purpose of this invention is to provide a suffocation risk assessment system and method based on infant regurgitation detection, in order to solve the problems existing in the prior art.
[0049] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0050] Example 1:
[0051] This embodiment provides a method for assessing the risk of suffocation based on infant regurgitation detection, including the following steps:
[0052] S1. Acquire visible light and thermal images of the same monitoring scene; specifically, the visible light images are acquired using two coaxially mounted visible light cameras with an optical center deviation of less than 0.5mm and a field-of-view matching error of less than 1 degree. The visible light cameras use 2-megapixel CMOS sensors, support 1080P@30fps, and are equipped with an 850nm invisible thermal imaging supplementary light; the thermal imaging images are acquired using a thermal imaging camera with a resolution of no less than 320×240 and a thermal sensitivity of <50mK. Hardware trigger signals are used to ensure that the time deviation between the acquisition of visible light and thermal images is less than 1m.
[0053] In this step, a Sony IMX415 visible light sensor with 2-megapixel resolution is used, supporting 1080P@30fps video acquisition. It is equipped with an 850nm infrared LED with adjustable power from 5-15W to ensure clear images under various lighting conditions. A FLIR Boson 640 thermal imaging sensor is used, with a resolution of 640×512, thermal sensitivity better than 50mK, and an operating wavelength of 8-14μm, accurately capturing the body temperature distribution characteristics of infants. The camera should be installed facing the infant's sleeping area at a recommended height of 1.2-1.5 meters, with a downward angle of less than 25 degrees to ensure complete coverage of the infant's activity area. The installation location should avoid direct sunlight and heat sources, and there should be no obstructions between the camera and the infant. The system uses a hardware synchronization trigger mechanism to ensure that the time synchronization error between the visible light image and the thermal imaging image is less than 1 millisecond. During image acquisition, the system uses a built-in video decoding chip to decode the dual video streams in real time. Visible light video streams are encoded in H.264 format, while thermal imaging video streams are encoded in H.265 format, and after decoding, they are converted into YUV420 format image sequences. Each video stream is broken down into consecutive image frames at a frame rate of 30fps, with each frame containing precise timestamp information for subsequent multimodal data alignment processing. The system automatically adjusts exposure parameters and gain to ensure optimal image quality under different ambient light conditions. To ensure data acquisition stability, the system incorporates a temperature compensation algorithm to automatically correct for the impact of ambient temperature changes on thermal imaging data. Simultaneously, the system monitors image quality metrics in real time, including signal-to-noise ratio, sharpness, and contrast. When a decline in image quality is detected, it automatically triggers refocusing and parameter adjustment procedures. The acquired image data undergoes preprocessing, including denoising, distortion correction, and white balance adjustment, providing high-quality input data for subsequent analysis and processing.
[0054] S2. Perform extrinsic and intrinsic parameter corrections on the acquired visible light and thermal imaging images, and complete viewpoint registration to obtain the registration matrix H; specifically including:
[0055] S21. Perform image distortion correction on images acquired by the visible light camera and the thermal imaging camera to obtain distorted visible light images and distorted thermal imaging images. Specifically, this includes: preparing a calibration board; acquiring images from multiple angles; using OpenCV's cv2.findChessboardCorners to detect the checkerboard corner points in each image; calling OpenCV's cv2.calibrateCamera, inputting the corner coordinates and actual physical dimensions of all images (e.g., a grid width of 30mm), and calculating the intrinsic parameter matrix K and distortion parameters; loading the intrinsic parameter matrix and distortion parameter coefficients to perform image distortion correction.
[0056] S22. Perform feature extraction and calibration parameter calculation on the distorted visible light image and the distorted thermal imaging image to obtain the registration matrix H from the visible light image to the thermal imaging image; specifically: image preprocessing, image initialization, unified image resolution, and edge extraction of the visible light image; feature extraction and matching, using OpenCV's ORB to extract feature points from the visible light image; using OpenCV's Canny to enhance the edges of the thermal imaging image, using mutual information as a similarity metric; using FLANN for feature matching, and then filtering out mismatches using the RANSAC operator; solving the similarity transformation matrix H through the matching points, which can be done using cv2.estimateAffinePartial2D(src_pts,dst_pts) to solve the rigid transformation matrix, where src_pts are the thermal imaging feature points and dst_pts are the visible light feature points.
[0057] In this step, distortion correction is performed on the image to avoid significant errors during registration and subsequent image processing. The purpose of registration is to obtain a registration matrix from visible light to heatmap, which will be used later to map the visible light image of the infant's mouth and nose area onto the heatmap.
[0058] Specifically, the first step is to accurately register the visible light image and the thermal infrared image acquired at the same time. In practice, a two-stage registration method based on a calibration plate is used to ensure registration accuracy. In the image preprocessing stage, the visible light image undergoes automatic white balance processing to eliminate color differences caused by changes in ambient light. Simultaneously, a lens distortion correction algorithm based on the Brown-Conrady model is used, employing pre-calibrated camera intrinsic parameters (including focal length, principal point coordinates, radial distortion coefficients k1, k2, k3, and tangential distortion coefficients p1, p2) to perform geometric correction, eliminating barrel and pincushion distortion caused by wide-angle lenses. For the thermal infrared image, non-uniformity correction (NUC) is first performed, obtaining the response curves of each pixel through blackbody calibration to eliminate inconsistencies in the response of different pixels in the sensor. Then, temperature calibration is performed, converting the original grayscale values into an accurate temperature distribution map, with temperature measurement accuracy controlled within ±0.3℃. In the feature registration stage, the system employs a multi-scale feature fusion strategy to improve registration robustness. For visible light images, an improved SURF feature detection algorithm is used. A Gaussian difference pyramid is employed when constructing the scale space, with a Hessian matrix threshold of 800, to extract stable feature points with rotation invariance. Simultaneously, Canny edge detection is combined to obtain contour information. For thermal infrared images, adaptive histogram equalization is first performed to enhance contrast. Then, significant heat source boundaries are extracted based on the temperature gradient field, and Harris corner detection is used to locate feature regions. Feature matching employs an improved PROSAC (Progressive Consistent Sampling) algorithm, introducing a feature point quality evaluation mechanism within the RANSAC framework. High-confidence feature pairs are prioritized for matching, with 2000 iterations and a reprojection error threshold of 1.5 pixels. The final calculated optimal homography matrix H meets the accuracy requirement of a reprojection error of less than 0.8 pixels. This matrix will be used to accurately map the coordinates of the mouth and nose region in the visible light image to the thermal infrared image space. To maintain registration stability, the system automatically performs online calibration every 30 minutes, updating registration parameters in real time by detecting fixed reference objects in the scene (such as crib railings) to compensate for registration errors caused by temperature changes or slight equipment movements.
[0059] S3. In the visible light image, the infant human target is located using a human detection model to obtain the infant's bounding box; within the infant's bounding box, the coordinates of the infant's mouth and nose region are obtained through face detection and facial landmark detection; specifically including:
[0060] S31. Construct and train a human detection model, a face localization model, and a facial landmark localization model; specifically: for the human detection task, use YOLOv8n or Mobilenet-SSD as the baseline model, and fine-tune it by adding 10,000 infant-specific scene images to the COCO dataset; for the face detection task, use RetinaFace as the baseline model; for the facial region localization model, use HRNet as the baseline model, and train it through transfer learning on baby facial data to improve robustness to newborn facial features; before training, collect and label training data for different models; after training, select the model whose accuracy meets the task requirements as the inference model; before inference, use the homography H from visible light to infrared light to map the coordinates of the baby's mouth and nose region back to the infrared image, collect the temperature distribution vector and regional texture features of the region, and feed them into the model to infer the risk of suffocation;
[0061] S32. Use the human detection model YOLOv8n as the detection model to locate the bounding box of the baby's body, detect whether the baby appears, and if it does, crop the image of that region.
[0062] S33. Within the captured infant bounding box, use a RetinaFace-based facial localization model to detect the infant's facial region, and use an HRNet-based facial landmark localization model to locate the coordinates of the infant's mouth and nose within the facial region.
[0063] In this step, the sub-coordinates of the infant's mouth and nose region are located sequentially using a human detection model, a face detection model, and a facial feature localization model in the visible light image. During visible light image processing, the system employs a three-tiered detection architecture to accurately locate the infant's mouth and nose region. First, the input image is pre-analyzed using a pre-trained YOLOv8n human detection model. This model is fine-tuned using an additional 10,000 labeled infant scene images on top of the COCO dataset, particularly enhancing its ability to recognize infants' bodies in various sleeping positions (supine, lateral, and prone). After obtaining the human body region, the system uses an improved RetinaFace face detection model for secondary processing. This model optimizes the anchor settings for infant facial features, expanding the original model's 5 feature pyramid levels to 7 to accommodate the smaller facial size of infants. The model's input image resolution is 640×640 pixels, and MobilenetV3 is used as the feature extraction backbone network, improving detection accuracy to 98.2% while maintaining real-time performance. For each detected facial region, the system performs pose evaluation, filtering out low-quality detection results with a side angle greater than 45 degrees. In the final stage, a keypoint detection algorithm based on HRNet is used to accurately locate the mouth and nose region. This network maintains high-resolution features throughout its structure, outputting 86 facial keypoints, including 20 specifically labeled feature points for the mouth and nose region (including the tip of the nose, the edges of the nostrils, and the contours of the lips). The system predicts keypoint locations using heatmap regression, employing adaptive Wing loss as the loss function. After training on a self-built dataset containing 50,000 infant facial images, the average keypoint localization error is less than 1.5 pixels. The final output mouth and nose region coordinates are dynamically generated as a 32×32 pixel region of interest (ROI) centered on the tip of the nose. The size of this ROI is automatically adjusted according to the detected facial dimensions to ensure complete coverage of the mouth, nose, and surrounding areas. The entire detection process on the embedded platform is completed within 80ms, meeting real-time requirements.
[0064] S4. Based on the registration matrix H obtained in step S2 and the mouth and nose region coordinates obtained in step S3, map the mouth and nose region coordinates onto the thermal imaging image and extract the corresponding mouth and nose thermal sub-region ROI; specifically including:
[0065] S41. Map the visible light mouth and nose region coordinates back to the thermal imaging region coordinates according to the registration matrix H;
[0066] S42. Extract the thermal imaging area image based on the coordinates of the thermal imaging area.
[0067] S5. Perform continuous frame statistics and time series feature construction on the ROI region to obtain the temperature distribution vector sequence F. t Temperature-color temporal characteristics G t ; Ft Substituting the data into a pre-trained infant occlusion risk assessment model, the occlusion risk probability is obtained, and G is... t Substituting the data into a pre-trained infant regurgitation assessment model, the probability of regurgitation is obtained. The probability of suffocation risk (P) is then integrated with the probability of occlusion risk (Pocclusion), specifically including:
[0068] S51. Construct and train an infant asphyxia risk assessment model and an infant regurgitation assessment model; wherein:
[0069] The infant occlusion risk assessment model adopts a two-branch CNN-LSTM model. The training method of the infant occlusion risk assessment model is as follows: collect and label the oral and nasal thermal subregion sequence data, with labels including three categories: occluded, partially occluded, and unoccluded; use cross-entropy and focus loss function to jointly train the model; freeze the network weights and derive the inference model when the validation AUC reaches above 0.95.
[0070] The infant regurgitation assessment model adopts a two-branch CNN-Transformer model, with one branch extracting single-frame spatial texture features and the other branch extracting time-series features. The oral and nasal temperature distribution vector sequence collected over time is used as training data to classify the data into two categories: regurgitation and no regurgitation. Focus loss is used as the loss function to alleviate the optimization difficulties caused by uneven class distribution. When the F1 accuracy on the validation set reaches above 0.92, the network weights are frozen, and the inference model is exported and saved.
[0071] S52. Acquire temporal continuous frame features of the oral and nasal thermal sub-region F t This includes the inter-frame mean temperature difference, local binarized mode temporal texture descriptor, and thermal diffusion gradient. Specifically, it involves: acquiring consecutive frame images of the oral and nasal thermal sub-region and the inter-frame mean temperature difference; extending the traditional LBP to temporal features, extracting LBP features for each frame, and stacking them along the time axis to generate temporal texture features; and using the Sobel operator to calculate the spatial gradient of each frame of the acquired images. Calculate inter-frame temporal gradient Taking the average of the gradient field yields the global diffusion direction: Construct a vector Ft containing the time series plot, mean temperature difference, local temporal texture, and thermal diffusion gradient; collect temporal continuous frame features G of the mouth and nose region. t This includes the average temperature and temperature gradient across consecutive frames, as well as the increments in brightness and saturation of the visible color gamut. Specifically, it involves extracting the average temperature and temperature gradient across consecutive frames. The local temperature decrease gradient Dt obtained by the mutation detection algorithm; the instantaneous increments ΔV(t) and ΔS(t) of temperature V and saturation S in the visible photonic region HSV color gamut:
[0072]
[0073] S53. F t Substituting the data into the infant occlusion risk assessment model, we obtain the occlusion risk probability, and then assign G... t Substitute the data into the infant spitting-up assessment model to obtain the probability of spitting-up occurring;
[0074] S54. Integrate the probability of suffocation risk and the probability of vomiting to obtain the probability of choking risk P.
[0075] S6. If the probability exceeds the corresponding threshold, an alarm signal is triggered, and the warning threshold is dynamically adjusted. This embodiment automatically adjusts the alarm threshold based on the ambient temperature, taking into account individual differences in infants (establishing a baseline through the initial learning stage). Multi-channel alarms include local audible and visual alarms (105dB buzzer + LED flashing), wireless network push (supports 4G / Wi-Fi), and SMS notification backup channels. The alarm verification mechanism triggers an alarm only if the threshold is exceeded for 3 consecutive seconds, supports remote video confirmation, and provides emergency handling suggestions.
[0076] Example 2:
[0077] This embodiment also provides a suffocation risk assessment system based on infant regurgitation detection, including:
[0078] The image acquisition module, including a visible light camera and a thermal imaging camera, is used to acquire the baby's video stream and decompose the video stream into multiple frames of images, including visible light images and thermal imaging images.
[0079] The edge computing unit includes a calibration and registration module, a detection and positioning module, a risk assessment module, and an alarm module, wherein:
[0080] The calibration and registration module is used to perform image distortion correction on images acquired by visible light cameras and thermal imaging cameras to obtain distorted visible light images and distorted thermal imaging images; feature extraction and calibration parameter calculation are performed on the distorted visible light images and distorted thermal imaging images to obtain the registration matrix H from the visible light image to the thermal imaging image;
[0081] The detection and localization module consists of multiple deep learning models. First, it locates the infant's body area, then it locates the infant's face area through the body area, and finally it locates the coordinates of the mouth and nose area through the infant's face area.
[0082] The risk assessment module is used to collect temporal continuous frame features F of the oral and nasal thermal sub-region. t This includes inter-frame mean temperature difference, local binarization pattern temporal texture descriptor, and thermal diffusion gradient; it also collects temporal continuous frame features G of the mouth and nose region. t This includes the average temperature of consecutive frames and the temperature drop gradient, as well as the increments in brightness and saturation of the visible color gamut; F tSubstituting the data into the infant occlusion risk assessment model, we obtain the occlusion risk probability, and then assign G... t Substitute the data into the infant regurgitation assessment model to obtain the probability of regurgitation; integrate the probability of occlusion risk and the probability of regurgitation to obtain the suffocation risk probability P.
[0083] The alarm module is used to issue an early warning when the probability of suffocation exceeds a certain threshold, reminding caregivers to intervene.
[0084] The system in this embodiment utilizes both infrared and visible light images, enabling more accurate handling of infant smothering and spitting up. Due to the multimodal approach, the model exhibits greater robustness and improved recognition accuracy. The use of real-time image processing helps protect infant safety, and compared to video surveillance, this method is more conducive to protecting the privacy of infants and their families. In summary, this device achieves real-time monitoring of infant suffocation risk, improving monitoring accuracy and efficiency.
[0085] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0086] It should be noted that the components mentioned in the above embodiments are all general standard parts or components known to those skilled in the art. Their structures and principles can be learned by those skilled in the art through technical manuals or conventional experimental methods.
[0087] This invention has illustrated its principles and implementation methods using specific examples. The descriptions of these embodiments are merely illustrative of the method and its core ideas; furthermore, those skilled in the art will recognize that modifications may be made to the specific implementation methods and application scope based on the principles of this invention. Therefore, the content of this specification should not be construed as limiting the invention.
Claims
1. A method for assessing the risk of suffocation based on infant regurgitation detection, characterized in that: Includes the following steps: S1. Acquire visible light and thermal images of the same monitoring scene; S2. Perform extrinsic and intrinsic parameter corrections on the acquired visible light and thermal imaging images, and complete the viewing angle registration to obtain the registration matrix H; S3. In the visible light image, the human body detection model is used to locate the infant human body target and obtain the infant bounding box; within the infant bounding box, the coordinates of the infant's mouth and nose region are obtained through face detection and facial key point detection; S4. Based on the registration matrix H obtained in step S2 and the mouth and nose region coordinates obtained in step S3, map the mouth and nose region coordinates onto the thermal imaging image and extract the corresponding mouth and nose thermal sub-region ROI. S5. Perform continuous frame statistics and time series feature construction on the ROI region to obtain the temperature distribution vector sequence F. t Temperature-color temporal characteristics G t ; F t Substituting the data into a pre-trained infant occlusion risk assessment model, the occlusion risk probability is obtained, and G is... t Substitute the pre-trained infant spitting up assessment model to obtain the probability of spitting up. Integrate the probability of occlusion risk and the probability of spitting up to obtain the choking risk probability P. S6. If the probability exceeds the corresponding threshold, an alarm signal is triggered.
2. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 1, characterized in that: In step S1, the visible light image is acquired using two coaxially mounted visible light cameras with an optical center deviation of less than 0.5 mm and a field-of-view matching error of less than 1 degree. The visible light cameras use a 2-megapixel CMOS sensor, support 1080P@30fps, and are equipped with an 850nm invisible thermal imaging fill light. The thermal imaging image is acquired using a thermal imaging camera with a resolution of not less than 320×240 and a thermal sensitivity of <50mK. A hardware trigger signal ensures that the time deviation between the acquisition of the visible light image and the thermal imaging image is less than 1m.
3. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 2, characterized in that: Step S2 includes: S21. Perform image distortion correction on the images acquired by the visible light camera and the thermal imaging camera to obtain distorted visible light images and distorted thermal imaging images; S22. Perform feature extraction and calibration parameter calculation on the distortion-free visible light image and the distortion-free thermal imaging image to obtain the registration matrix H from the visible light image to the thermal imaging image.
4. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 1, characterized in that: Step S3 includes: S31. Construct and train a human detection model, a facial localization model, and a facial landmark localization model; S32. Use the human detection model YOLOv8n as the detection model to locate the bounding box of the baby's body, detect whether the baby appears, and if it does, crop the image of that region. S33. Use a facial localization model to detect the infant's facial region within the captured infant bounding box, and use a facial keypoint localization model to locate the coordinates of the infant's mouth and nose within the facial region.
5. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 1, characterized in that: Step S4 includes: S41. Map the visible light mouth and nose region coordinates back to the thermal imaging region coordinates according to the registration matrix H; S42. Extract the thermal imaging area image based on the coordinates of the thermal imaging area.
6. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 1, characterized in that: Step S5 includes: S51. Construct and train an infant suffocation risk assessment model and an infant regurgitation assessment model; S52. Acquire temporal continuous frame features of the oral and nasal thermal sub-region F t This includes inter-frame mean temperature difference, local binarization pattern temporal texture descriptor, and thermal diffusion gradient; it also collects temporal continuous frame features G of the mouth and nose region. t This includes the average temperature across consecutive frames and the temperature drop gradient, as well as the increments in brightness and saturation of the visible color gamut. S53. F t Substituting the data into the infant occlusion risk assessment model, we obtain the occlusion risk probability, and then assign G... t Substitute the data into the infant spitting-up assessment model to obtain the probability of spitting-up occurring; S54. Integrate the probability of suffocation risk and the probability of vomiting to obtain the probability of choking risk P.
7. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 6, characterized in that: The infant occlusion risk assessment model uses a two-branch CNN-LSTM model, and the infant regurgitation assessment model uses a two-branch CNN-Transformer model.
8. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 7, characterized in that: The training method for the infant occlusion risk assessment model is as follows: collect and label the oral and nasal thermal sub-region sequence data, with labels including three categories: occluded, partially occluded, and uncovered; use cross-entropy and focus loss function to jointly train the model; freeze the network weights and derive the inference model when the validation AUC reaches 0.95 or higher.
9. The method for assessing the risk of suffocation based on infant regurgitation detection according to claim 7, characterized in that: The training method for the infant spitting up assessment model is as follows: collect and label sequence data of the mouth and nose region, with labels including two categories: spitting up and not spitting up; use the focus loss function to jointly train the model; freeze the network weights and derive the inference model when the validation set F1 reaches 0.92 or higher.
10. A choking risk assessment system based on infant regurgitation detection, characterized in that: include: The image acquisition module includes a visible light camera and a thermal imaging camera, used to acquire an infant video stream and decompose the video stream into multiple frames of images, including visible light images and thermal imaging images; The edge computing unit includes a calibration and registration module, a detection and positioning module, a risk assessment module, and an alarm module. The correction and registration module is used to perform image distortion correction on images acquired by the visible light camera and the thermal imaging camera to obtain distorted visible light images and distorted thermal imaging images; and to perform feature extraction and calibration parameter calculation on the distorted visible light images and distorted thermal imaging images to obtain the registration matrix H from the visible light image to the thermal imaging image. The detection and localization module consists of multiple deep learning models. First, it locates the infant's body area, then it locates the infant's face area through the body area, and finally it locates the coordinates of the mouth and nose area through the infant's face area. The risk assessment module is used to collect temporal continuous frame features F of the oral and nasal thermal sub-region. t This includes inter-frame mean temperature difference, local binarization pattern temporal texture descriptor, and thermal diffusion gradient; it also collects temporal continuous frame features G of the mouth and nose region. t This includes the average temperature of consecutive frames and the temperature drop gradient, as well as the increments in brightness and saturation of the visible color gamut; F t Substituting the data into the infant occlusion risk assessment model, we obtain the occlusion risk probability, and then assign G... t Substitute the data into the infant regurgitation assessment model to obtain the probability of regurgitation; integrate the probability of occlusion risk and the probability of regurgitation to obtain the suffocation risk probability P. The alarm module is used to issue an early warning when the probability of suffocation exceeds a certain threshold, reminding caregivers to intervene.
Citation Information
Patent Citations
Infant dangerous sleeping posture real-time identification method, device and equipment and storage medium
CN119007237A