Infant nursing method and device based on mouth and nose shielding detection and robot
By performing grayscale conversion and in-depth information processing on real-time video data in infant sleep scenarios, the oral and nose area is accurately positioned, and the accuracy of oral and nose occlusion detection detection of infant care robots in home scenarios is solved, and efficient safety monitoring is achieved.
Patent Information
- Application Number
- CN202510524620.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
The existing infant and child care robots cannot accurately perform oral and nose occlusion detection in family scenarios, and there are problems such as insufficient sensor sensitivity, poor environmental adaptability and high false alarm rate.
By obtaining real-time video data in the sleep care scenario of infants and young children, decompose it into multi-frame images, performing grayscale conversion and face detection, positioning the oral and nose area, combining depth information for occlusion detection, and issuing an alarm when occlusion is detected.
It improves the accuracy of oral and nose occlusion detection and environmental adaptability, reduces false alarms and missed reports, ensures the safety of infants and young children, and achieves active protection.
Smart Images

Figure CN120451538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infant care robots, and in particular to an infant care method, device and robot based on mouth and nose occlusion detection. Background Art
[0002] During the daily sleep care of infants and young children, ensuring their airways are unobstructed is crucial. Because infants and young children's physiological structures are not yet fully developed and their autonomous control abilities are limited, they are prone to respiratory problems caused by obstructions to their mouths and noses. Common sources of obstruction include soft items such as quilts, toys, and pillows, or obstructions caused by improper sleeping positions. These obstructions not only pose a risk of suffocation but also affect the infant's sleep quality and physical health. With the development of smart home technology, infant care robots have gradually become an important tool to assist in childcare. By integrating visual perception, sensor fusion, and intelligent judgment algorithms, infant care robots have the potential to achieve real-time monitoring and intelligent response to infants' mouth and nose obstructions, thereby significantly improving the safety and scientific nature of home care.
[0003] However, existing infant care robots still face numerous challenges in detecting mouth and nose occlusion, such as insufficient sensor sensitivity, poor environmental adaptability, and high false alarm rates. Therefore, developing high-precision, environmentally adaptable mouth and nose occlusion detection technology suitable for infant care robots is crucial for improving infant safety monitoring. This will not only help reduce false alarms and missed alerts, but also promote the transition of infant care robots from assistive care to active protection, creating a safer and more intelligent environment for infants and young children.
[0004] Existing Chinese patent CN116959071A discloses a method and terminal device for detecting mouth and nose occlusion in children, comprising: obtaining a facial image of the child to be tested; inputting the facial image of the child to be tested into a trained posture detection model to perform posture estimation and obtain a posture result; determining a corresponding trained key point occlusion estimation model based on the posture result to obtain an optimal key point occlusion estimation model; inputting the facial image of the child to be tested into the optimal key point occlusion estimation model to perform key point occlusion detection and obtain a key point occlusion detection result; and determining the mouth and nose occlusion result based on the key point occlusion detection result. The above patent uses a posture estimation model and a key point occlusion estimation model to gradually process the facial image of the child to be tested to detect mouth and nose occlusion. However, the posture estimation and key point occlusion estimation processes rely on different training models, which leads to complex model selection and training processes. Secondly, key point detection is sensitive to facial posture and lighting changes, affecting detection accuracy.
[0005] Therefore, how to accurately detect mouth and nose occlusion of infants and young children in home care scenarios is an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides an infant care method, device and robot based on mouth and nose occlusion detection, which is used to solve the problem in the prior art that it is impossible to accurately detect mouth and nose occlusion of infants in family care scenarios.
[0007] The technical solution adopted in the present invention is:
[0008] In a first aspect, the present invention provides a method for caring for infants and young children based on mouth and nose occlusion detection, the method comprising:
[0009] S1: Acquire real-time video data collected by a care robot in a baby sleep care scene, and decompose the real-time video data into multiple frames of real-time images;
[0010] S2: Convert each frame of real-time image into a grayscale image, perform infant face detection on the grayscale image, and obtain infant face area position information;
[0011] S3: Positioning the mouth and nose region of the infant's face region to obtain target position information of the mouth and nose region;
[0012] S4: Acquire depth information of the mouth and nose region based on the target position information of the mouth and nose region;
[0013] S5: Detecting occlusion of the infant's mouth and nose based on the mouth and nose area depth information and the mouth and nose area target position information, and outputting a detection result;
[0014] S6: According to the detection result, when it is detected that the mouth and nose of the infant are blocked, the care robot is controlled to issue an alarm message.
[0015] Preferably, the S3 includes:
[0016] S31: Enlarging the position information of the infant's face region to obtain an image of the infant's head region;
[0017] S32: Inputting the infant head region image into a pre-trained hair detection model, and outputting the infant hair region position information;
[0018] S33: Determine the target position information of the mouth and nose area based on the position information of the infant's face area and the position information of the infant's hair area.
[0019] Preferably, the S33 includes:
[0020] S331: Acquire first center position information of the infant's face region and second center position information of the infant's hair region based on the infant's face region position information and the infant's hair region position information;
[0021] S332: Connecting a first center point of the infant's face region and a second center point of the infant's hair region based on the first center position information and the second center position information, and outputting a connecting line direction as a direction of the mouth and nose region;
[0022] S333: Determine initial position information of the mouth and nose region based on the direction of the mouth and nose region and the position information of the infant's face region;
[0023] S334: Extending the width of the initial mouth and nose area according to the initial position information of the mouth and nose area to determine the target position information of the mouth and nose area.
[0024] Preferably, the S4 includes:
[0025] S41: Inputting the real-time image into a pre-trained depth calculation model to output real-time depth information;
[0026] S42: Determine the depth information of the mouth and nose area based on the target position information of the mouth and nose area and the real-time depth information.
[0027] Preferably, the S5 includes:
[0028] S51: Obtaining the width of the mouth and nose area according to the target position information of the mouth and nose area;
[0029] S52: Determine a depth value sequence of each column of pixel points in the mouth and nose region according to the width of the mouth and nose region and the depth information of the mouth and nose region;
[0030] S53: Processing the depth value sequence to determine whether it meets the mouth and nose occlusion requirements;
[0031] S54: If the depth value sequence meets the mouth and nose occlusion requirement, it is detected that the infant's mouth and nose are occluded, and a safety reminder is issued;
[0032] S55: If the depth value sequence does not meet the mouth and nose occlusion requirement, it is detected that the infant's mouth and nose are not occluded, and no safety reminder is issued.
[0033] Preferably, the S53 includes:
[0034] S531: performing averaging processing on each depth value sequence to determine a first depth mean value;
[0035] S532: performing difference calculation on the first depth mean values of adjacent columns to obtain depth differences;
[0036] S533: averaging the depth differences to determine a second depth mean;
[0037] S534: Obtain the maximum depth difference among the depth differences, and determine whether the mouth and nose occlusion requirement is met based on the maximum depth difference and the second depth average.
[0038] Preferably, the S534 includes:
[0039] S5341: Compare the maximum depth difference with the second depth mean, and obtain the number of maximum depth differences greater than the second depth mean;
[0040] S5342: Compare the number of the maximum depth differences with a preset threshold. If the number of the maximum depth differences is greater than the preset threshold, the mouth and nose occlusion requirement is met.
[0041] S5343: If the number of maximum depth differences is less than or equal to the preset threshold, the mouth and nose occlusion requirement is not met.
[0042] In a second aspect, the present invention provides an infant care device based on mouth and nose occlusion detection, the device comprising:
[0043] A real-time image acquisition module is used to acquire real-time video data collected by the care robot in the infant sleep care scene, and decompose the real-time video data into multiple frames of real-time images;
[0044] A face detection module is used to convert each frame of real-time image into a grayscale image, perform infant face detection on the grayscale image, and obtain the location information of the infant's face area;
[0045] An oropharyngeal region positioning module, configured to perform oropharyngeal region positioning on the infant's face region position information to obtain target oropharyngeal region position information;
[0046] A depth information acquisition module, configured to acquire depth information of the mouth and nose region based on the target position information of the mouth and nose region;
[0047] an occlusion detection module, configured to detect occlusion of the infant's mouth and nose based on the depth information of the mouth and nose area, and output a detection result;
[0048] The alarm module is used to control the care robot to issue an alarm message when it detects that the mouth and nose of the infant are blocked according to the detection result.
[0049] In a third aspect, an embodiment of the present invention further provides an infant care robot, comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect of the above embodiment.
[0050] In a fourth aspect, an embodiment of the present invention further provides a storage medium having computer program instructions stored thereon, which implements the method of the first aspect of the above-mentioned embodiment when the computer program instructions are executed by a processor.
[0051] In summary, the beneficial effects of the present invention are as follows:
[0052] The present invention provides an infant care method, device, and robot based on mouth and nose occlusion detection, comprising: obtaining real-time video data collected by a care robot in an infant sleep care scenario, and decomposing the real-time video data into multiple frames of real-time images; converting each frame of the real-time image into a grayscale image, performing infant face detection on the grayscale image, and obtaining position information of the infant's facial area; performing mouth and nose area positioning on the infant's facial area position information, and obtaining target position information of the mouth and nose area; obtaining depth information of the mouth and nose area based on the target position information of the mouth and nose area; detecting mouth and nose occlusion of the infant based on the mouth and nose area depth information and the mouth and nose area target position information, and outputting a detection result; and controlling the care robot to issue an alarm message when it is detected that the infant's mouth and nose are blocked based on the detection result. The present invention uses an infant face detection algorithm to accurately locate the facial area, ensuring the precise positioning of the mouth and nose areas, thereby improving the accuracy of occlusion detection. At the same time, the acquisition of depth information further enhances the ability to identify occlusion situations, and can maintain efficient detection under different ambient lighting and facial postures. By combining these steps, the present invention can provide timely and accurate mouth and nose occlusion detection, effectively protecting the safety of infants and young children. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work, and these are all within the scope of protection of the present invention.
[0054] Figure 1 Schematic diagram of the overall working process of the infant care method based on mouth and nose occlusion detection in Example 1 of the present invention;
[0055] Figure 2 This is a schematic diagram of the process of locating the mouth and nose area based on the facial area position information of the infant in Example 1 of the present invention;
[0056] Figure 3This is a schematic diagram of the process of determining the target position information of the mouth and nose area in Example 1 of the present invention;
[0057] Figure 4 This is a schematic diagram of the process of obtaining depth information of the mouth and nose area in Example 1 of the present invention;
[0058] Figure 5 This is a schematic diagram of the process of detecting mouth and nose occlusion of infants and young children in Example 1 of the present invention;
[0059] Figure 6 Schematic diagram of the process of processing the depth value sequence in embodiment 1 of the present invention;
[0060] Figure 7 Schematic diagram of the process of determining whether the mouth and nose occlusion requirements are met in Example 1 of the present invention;
[0061] Figure 8 This is a structural block diagram of an infant care device based on mouth and nose occlusion detection in Example 2 of the present invention;
[0062] Figure 9 This is a schematic structural diagram of the infant care robot in Example 3 of the present invention. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the orientation or position relationship indicated by the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further limitations, elements defined by the phrase "comprising..." do not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising the elements. The embodiments of the present invention and the features thereof may be combined with each other if there is no conflict, and all are within the scope of protection of the present invention.
[0064] Example 1
[0065] See Figure 1 Embodiment 1 of the present invention discloses a method for caring for infants and young children based on mouth and nose occlusion detection, the method comprising:
[0066] S1: Acquire real-time video data collected by a care robot in a sleep care scene for infants and young children, and decompose the real-time video data into multiple frames of real-time images;
[0067] Specifically, in the infant sleep care scenario, the real-time video data collected by the care robot is obtained. The care robot is an intelligent device designed specifically for caring for infants and young children. It usually integrates hardware modules such as cameras, sensors, microphones, speakers, etc., and can realize real-time monitoring of infants and young children, crying recognition, sleeping posture detection, environmental monitoring and other functions; it can use AI algorithms to identify the status of infants and young children and issue reminders when abnormalities occur, such as turning over to cover the mouth and nose, not moving for a long time, or crying for too long, so as to assist parents in scientific and safe care; the care robot usually has an Internet connection function and can be realized through a mobile phone app Remote viewing and interaction are now possible, and the video data is decomposed into multiple continuous image frames (i.e. each frame of the video) in chronological order; the decomposition is completed through the video decoding module, such as frame extraction operations on the real-time stream based on image processing tools such as OpenCV or FFmpeg. The main purpose of this process is to structure the continuous video stream into single-frame image data, which is convenient for subsequent AI vision tasks such as image analysis, target detection, and posture recognition; through frame decomposition, frame-by-frame recognition and analysis of key behaviors such as infant posture, turning over, crying, and mouth and nose occlusion can be achieved, thereby providing raw data support for intelligent judgment and abnormal alarm.
[0068] S2: Convert each frame of real-time image into a grayscale image, perform infant face detection on the grayscale image, and obtain infant face area position information;
[0069] Specifically, through color space conversion, the RGB (red, green, and blue) three-channel information in the color image is converted into a single grayscale value. This grayscale value is based on the weighted average of the brightness of each channel. For example, the formula Gray = 0.299*R + 0.587*G + 0.114*B is used to better match the human eye's sensitivity to different color brightness. The grayscale image simplifies the complexity of data processing and removes unnecessary color information, making subsequent infant face detection more efficient and robust. After obtaining the grayscale image, a pre-trained face detection algorithm, such as Haar features or a deep learning algorithm, is used. The deep learning algorithm includes an object detection algorithm trained on YoloV8s to locate the infant's facial area. This process involves scanning features in the image, such as the location of the eyes, nose, and mouth, and matching them with the face template in the object detection model trained on YoloV8s to determine the accurate location information of the facial area. This information provides an accurate basis for subsequent posture, breathing monitoring, and expression analysis, and also significantly improves the efficiency and reliability of infant face recognition in complex environments.
[0070] S3: Positioning the mouth and nose region of the infant's face region to obtain target position information of the mouth and nose region;
[0071] Specifically, when processing the facial area of infants and young children, the overall contour of the face is identified through image processing or deep learning technology to determine the boundaries and key points of the face. The key point positions include the position information of facial features such as the eyes, nose, and mouth, and this information is further used for refined positioning. The mouth and nose area is the focus area of the infant's face. By analyzing these key points, the area where the mouth and nose are located can be more accurately delineated. The target position information of the mouth and nose area obtained not only provides a basis for subsequent analysis, such as for detecting breathing or eating status, but can also be used to generate training data to improve performance in infant face processing scenarios.
[0072] In one embodiment, see Figure 2 , the S3 includes:
[0073] S31: Enlarging the position information of the infant's face region to obtain an image of the infant's head region;
[0074] Specifically, the location information of the infant's facial region is first obtained, represented as Face(x1, y1, w1, h1), where x1 and y1 are the coordinates of the upper left corner of the minimum bounding rectangle of the facial region, and w1 and h1 are the width and height of the minimum bounding rectangle of the facial region. To better cover and analyze the entire head region of the infant, the facial region is enlarged. The enlargement method is usually doubled, that is, the original width and height are expanded to twice, resulting in a larger rectangular region called Head(x2, y2, w2, h2). This can include more surrounding information, especially areas such as hair and forehead. After the enlargement process, the new region not only facilitates the subsequent model's detection of hair, but also ensures that it contains sufficient contextual information, thereby improving the relative positioning accuracy between regions in the image.
[0075] S32: Inputting the infant head region image into a pre-trained hair detection model, and outputting the infant hair region position information;
[0076] Specifically, in the enlarged head region image Head(x2, y2, w2, h2), where x2 and y2 are the coordinates of the upper left corner of the minimum bounding rectangle of the head region, and w2 and h2 are the width and height of the minimum bounding rectangle of the face region, the infant head region image is input into a pre-trained hair detection model for processing. The hair detection model here uses the YOLOv8s model, which is a lightweight object detection model that can effectively identify object boundaries. In this step, YOLOv8s scans the entire head region image and uses its deep learning feature extraction capabilities to identify and output the position coordinate information of the infant's hair region, namely Hair(x3, y3, w3, h3), where x3 and y3 are the coordinates of the upper left corner of the minimum bounding rectangle of the hair region, and w3 and h3 are the width and height of the minimum bounding rectangle of the face region. This position information marks the precise boundary of the hair, helps distinguish the hair region from other facial regions, and facilitates subsequent positioning and processing.
[0077] S33: Determine the target position information of the mouth and nose area based on the position information of the infant's face area and the position information of the infant's hair area.
[0078] Specifically, based on the infant's facial area position information Face (x1, y1, w1, h1) and hair area position information Hair (x3, y3, w3, h3) obtained in the previous steps, combined with the relative position and structural characteristics of these two areas, the target position information of the mouth and nose area is further derived. Specifically, the mouth and nose area is usually located in the lower half of the face and not in the hair area. By excluding the hair area and using the geometric features of the face, such as the relative position relationship between the eyes, nose and mouth, the mouth and nose area can be accurately located. This target position information will serve as an important basis for subsequent processing and will be used for tasks such as respiratory monitoring and facial expression analysis to ensure the accuracy and effectiveness of the analysis.
[0079] In one embodiment, see Figure 3 , the S33 includes:
[0080] S331: Acquire first center position information of the infant's face region and second center position information of the infant's hair region based on the infant's face region position information and the infant's hair region position information;
[0081] Specifically, in this step, the previously obtained infant facial area position information Face (x1, y1, w1, h1) and hair area position information Hair (x3, y3, w3, h3) are first used to calculate the center point positions of the two areas, namely, the center C1 of the hair area and the center C2 of the facial area. The obtained C1 and C2 represent the center positions of the hair and face, respectively, providing basic coordinate information for the subsequent determination of the mouth and nose areas.
[0082] S332: Connecting a first center point of the infant's face region and a second center point of the infant's hair region based on the first center position information and the second center position information, and outputting a connecting line direction as a direction of the mouth and nose region;
[0083] Specifically, based on the calculated C1 and C2, a line is formed connecting these two center points. The direction of this line is considered the primary direction of the nose and mouth region, as the nose is typically located at the center of the face and is directly related to the center points of the face and hair regions. By calculating the line connecting C1 to C2, the possible location of the nose and mouth region and its relative orientation can be determined, providing a clear guide for further regional positioning.
[0084] S333: Determine initial position information of the mouth and nose region based on the direction of the mouth and nose region and the position information of the infant's face region;
[0085] Specifically, based on the determined direction of the mouth and nose area, the position information of the infant's facial area will be combined to start determining the initial position information of the mouth and nose area. According to the previously calculated connection direction, the initial position of the mouth and nose area can be defined in the lower half of the facial area. Since the nose position is close to the center of the face, the lower half of the face is selected as the preliminary range M_N(x4, y4, w4, h4) of the mouth and nose area, where x4 and y4 are the coordinates of the upper left corner of the minimum circumscribed rectangle of the mouth and nose area, and w4 and h4 are the width and height of the minimum circumscribed rectangle of the facial area, ensuring that it contains the main mouth and nose structures and providing an accurate reference for subsequent processing.
[0086] S334: Based on the initial position information of the mouth and nose area, extend the width of the initial mouth and nose area to determine the target position information of the mouth and nose area.
[0087] Specifically, after obtaining the initial position information of the mouth and nose area, the width is extended to ensure that the mouth and nose area can better cover all the features of the area. Specifically, the height is kept unchanged, and the width is extended to the left and right according to the previous calculation until it approaches the boundary of the Head (x2, y2, w2, h2) area. In this process, the actual boundary of the head area is taken into account to ensure that the final mouth and nose area M_N_new (x4, y4, w4, h4) not only conforms to physical reality, but also can effectively adapt to subsequent monitoring and analysis tasks, providing the final basis for judging the target position information of the mouth and nose area.
[0088] S4: Acquire depth information of the mouth and nose region based on the target position information of the mouth and nose region;
[0089] Specifically, based on the previously determined target position information M_N_new(x4,y4,w4,h4) of the mouth and nose area, the depth information of the area is further obtained. The acquisition of depth information usually relies on methods such as stereo vision technology, depth sensors or multi-view image analysis. These technologies capture the three-dimensional structure of the mouth and nose area and provide the distance and spatial position data of the area relative to the camera or observer. Specifically, the depth value of the corresponding mouth and nose area is extracted from the depth image to generate a depth map. These depth values can reflect the shape, convexity and concavity of the mouth and nose area and its relative position in the overall facial structure. Such depth information not only helps to improve the accuracy of facial feature recognition, but can also be used in infant health monitoring. For example, by analyzing the depth changes of the mouth and nose area during breathing, the physiological state of infants and young children can be monitored in real time. Therefore, through the fusion of depth information, the mouth and nose area can be understood and analyzed more comprehensively, providing richer basic data for subsequent processing and decision-making.
[0090] In one embodiment, see Figure 4 , said S4 includes:
[0091] S41: Inputting the real-time image into a pre-trained depth calculation model to output real-time depth information;
[0092] Specifically, a real-time image containing an infant's face is input into a pre-trained depth calculation model. The depth calculation model is based on deep learning methods, such as neural networks or convolutional neural networks (CNN), and has been trained with a large amount of image data containing depth information. The input image is usually an RGB image or an infrared image. The depth calculation model analyzes each pixel in the image and predicts its depth value relative to the camera. The depth calculation model analyzes information such as light, shadow, and edge features of objects in the scene, and outputs a depth map Depth(x,y) of the same size as the input image, where the value of each pixel represents the depth information of the point. The depth map can reflect the three-dimensional structure of the infant's face and its surrounding environment in real time, providing basic data for subsequent depth analysis.
[0093] S42: Determine the depth information of the mouth and nose area based on the target position information of the mouth and nose area and the real-time depth information.
[0094] Specifically, after obtaining the depth information Depth(x,y) of the entire real-time image, the corresponding regional depth field D_M_N(x,y) is extracted from the depth map based on the previously determined target position information M_N_new(x4,y4,w4,h4) of the mouth and nose area. By matching the coordinates of the mouth and nose area in the image with the depth data in the depth map, the depth value of each pixel in the mouth and nose area is accurately calculated. This process not only provides the depth distribution of the mouth and nose area, but also reveals the three-dimensional structural characteristics of the area. For example, by analyzing the depth gradient of D_M_N(x,y), the concave and convex changes and spatial position of the mouth and nose can be determined. Ultimately, this depth information helps to more comprehensively understand the facial features of infants and young children, and provides support for health monitoring or image processing, such as using depth changes for further analysis and judgment in respiratory monitoring or facial gesture recognition.
[0095] S5: Detecting occlusion of the infant's mouth and nose based on the mouth and nose area depth information and the mouth and nose area target position information, and outputting the detection result.
[0096] Specifically, infant and young child mouth and nose occlusion detection is performed based on the previously acquired mouth and nose area depth information D_M_N(x,y) and the mouth and nose area target position information M_N_new(x4,y4,w4,h4). First, by comparing the depth information of the mouth and nose area with the expected depth value, it is determined whether there is an abnormal depth change in the area. For example, if the depth value of a certain part suddenly becomes shallow, it means that an object (such as a hand, blanket, etc.) is blocking the mouth and nose area. Occlusion usually causes the mouth and nose area in the depth map to become flat or lack obvious concave and convex changes. The nature and severity of the occlusion can also be further confirmed by analyzing the size, shape, and duration of the occluded area. Once abnormal occlusion is detected, the detection result is output, including whether there is occlusion, the location of the occlusion, and its coverage. This result can be used to issue an alarm in real time to remind caregivers of the possible risk of respiratory obstruction, especially in infant sleep monitoring, to ensure timely intervention and protect the safety of infants and young children.
[0097] In one embodiment, see Figure 5 , the S5 includes:
[0098] S51: Obtaining the width of the mouth and nose area according to the target position information of the mouth and nose area;
[0099] Specifically, the width of the nose and mouth region is first obtained based on the previously determined target position information M_N_new (x4, y4, w4, h4). Specifically, the w value in M_N_new, i.e., the width parameter of the nose and mouth region, can be used to determine the horizontal pixel range of the region in the image. This width information is critical foundational data for subsequent analysis. The depth changes of each column of pixels in the nose and mouth region will be analyzed column by column based on this width. Determining the width ensures that the depth analysis is limited to the nose and mouth region, thereby avoiding the influence of interference information and improving the accuracy of occlusion detection.
[0100] S52: Determine a depth value sequence of each column of pixel points in the mouth and nose region according to the width of the mouth and nose region and the depth information of the mouth and nose region;
[0101] Specifically, based on the determined width of the nose and mouth region, the depth value sequence for each column of pixels is extracted from the corresponding depth information D_M_N(x,y). Specifically, the nose and mouth region is divided into w columns, each containing several pixels. The depth values of these pixels are averaged to obtain the average depth value for each column. This processed average depth value sequence represents the depth distribution of the nose and mouth region at different lateral positions. By sequentially arranging the depth values of each column, a longitudinal depth curve is generated that reflects the depth variation of the nose and mouth region, providing data support for subsequent detection of occlusion features.
[0102] S53: Processing the depth value sequence to determine whether it meets the mouth and nose occlusion requirements;
[0103] Specifically, after obtaining the average depth value sequence of each column of pixels in the mouth and nose area, these depth values are analyzed to determine whether they meet the characteristics of mouth and nose occlusion. First, the average depth difference Delta_depth between two consecutive columns is calculated, and the changing trend of the depth value is analyzed. Under normal circumstances, due to the hollow structure between the head and the bed, the depth value of the mouth and nose area usually shows an obvious "convex" shape change, which means that the average depth difference is larger in the middle area and gradually decreases at the edge. However, when there is mouth and nose occlusion, obstructions such as textiles will make the depth change tend to be gentle, showing a "slope" type of continuous change, and the depth difference Delta_depth is small. By analyzing the depth value sequence, the smoothness of the depth change can be judged, and then whether there is occlusion. Once a depth change that meets the occlusion characteristics is detected, the occlusion judgment result will be output to ensure the safety of the mouth and nose area.
[0104] In one embodiment, see Figure 6 , the S53 includes:
[0105] S531: performing averaging processing on each depth value sequence to determine a first depth mean value;
[0106] Specifically, the depth value sequence of each column is averaged to determine the first depth mean of the column. Specifically, the corresponding depth information D_M_N(x,y) is extracted from each column of pixels, and then these depth values are added and divided by the number of pixels to obtain the average depth value of each column. By performing this process on each column of the entire mouth and nose area, the system generates a mean sequence reflecting the longitudinal depth change. This first depth mean provides basic data for subsequent judgments and can describe the overall depth trend of the mouth and nose area at different lateral positions.
[0107] S532: performing difference calculation on the first depth mean values of adjacent columns to obtain depth differences;
[0108] Specifically, the difference calculation is performed on the first depth means of adjacent columns to obtain the depth difference between every two columns. The depth means of two adjacent columns are taken out in turn from the mean sequence, and the depth mean of the previous column is subtracted from the depth mean of the latter column to obtain the depth difference between the columns. These depth differences can reflect the lateral depth changes in the mouth and nose area. If there is a large difference, it usually indicates a hollowing phenomenon between the head and the bed surface, while a smaller difference may mean that an obstruction smoothly covers the mouth and nose area.
[0109] S533: averaging the depth differences to determine a second depth mean;
[0110] Specifically, after obtaining the depth differences between all columns, these depth differences are averaged to determine a second depth mean, or the average level of depth differences. Specifically, all depth differences are added together and divided by the number of differences. This second depth mean provides a reference for measuring the overall depth variation trend of the nose and mouth area. A larger second depth mean typically indicates a more pronounced protrusion in the head structure, while a smaller second depth mean may indicate the presence of an occluder.
[0111] S534: Obtain the maximum depth difference among the depth differences, and determine whether the mouth and nose occlusion requirement is met based on the maximum depth difference and the second depth average.
[0112] Specifically, the maximum depth difference is extracted from all depth differences as the most significant depth change representative. Based on the maximum depth difference and the second depth mean, further judgment will be made as to whether the infant is in an occlusion state, and the corresponding detection result will be output.
[0113] In one embodiment, see Figure 7 , the S534 includes:
[0114] S5341: Compare the maximum depth difference with the second depth mean, and obtain the number of maximum depth differences greater than the second depth mean;
[0115] Specifically, the maximum depth difference is compared with the second depth mean one by one to determine how many depth differences are greater than the second depth mean. The depth difference sequence is traversed, and those depth differences greater than the second depth mean are screened out and counted. The purpose of this step is to find the most significant depth difference in the depth change of the mouth and nose area, and to judge whether there is occlusion based on these significant changes. This comparison method can effectively measure whether the depth field of the mouth and nose area presents a "convex" feature or whether the depth change is smooth due to the occlusion.
[0116] S5342: Compare the number of the maximum depth differences with a preset threshold. If the number of the maximum depth differences is greater than the preset threshold, the mouth and nose occlusion requirement is met.
[0117] Specifically, the maximum depth difference obtained in the previous step is compared with a preset threshold. The preset threshold is usually a reasonable value set based on actual conditions. For example, the preset threshold can be set to 3 times the second depth mean to distinguish normal depth changes from depth changes caused by occlusion. If the maximum depth difference greater than 3 times the second depth mean exceeds the threshold, the depth field is considered to meet the unobstructed characteristic, that is, the depth change is significant enough to indicate that the hollow area between the head and the bed surface is not obstructed. This comparison mechanism ensures that the system can sensitively and accurately detect depth changes in the mouth and nose area.
[0118] S5343: If the number of maximum depth differences is less than or equal to the preset threshold, the mouth and nose occlusion requirement is not met.
[0119] Specifically, if the maximum depth difference is less than or equal to a preset threshold, the depth change in the mouth and nose area is considered unobstructed. In other words, the depth change is relatively gentle because the mouth and nose area is covered by an object such as textile, resulting in no obvious "convex" change in the depth field. This judgment is based on the maximum depth difference, and comparing it with the preset threshold can help the system effectively distinguish between occlusion and unobstructed states, further improving detection accuracy.
[0120] S54: If the depth value sequence meets the mouth and nose occlusion requirement, it is detected that the infant's mouth and nose are occluded, and a safety reminder is issued;
[0121] Specifically, if the depth value sequence matches the characteristics of mouth and nose occlusion, the infant's mouth and nose will be detected and a safety alert will be issued immediately. Based on the results of the previous step, the possibility of occlusion is determined and the corresponding warning mechanism is triggered, such as sending an alarm to the guardian or emitting an audible reminder. This step is crucial because it directly affects the safety of the infant, and the system's response speed and accuracy are particularly critical.
[0122] S55: If the depth value sequence does not meet the mouth and nose occlusion requirement, it is detected that the infant's mouth and nose are not occluded, and no safety reminder is issued.
[0123] Specifically, if the depth value sequence does not meet the mouth and nose occlusion requirements, the infant's mouth and nose are considered unobstructed, and no safety alert is issued. In this case, the depth variation in the mouth and nose area is considered normal, and there is sufficient space between the head and the bed surface, indicating that the infant is breathing smoothly. This step ensures that the system does not issue false alarms, thereby avoiding unnecessary stress or interference for guardians and ensuring the reliability and effectiveness of the system in actual use.
[0124] S6: According to the detection result, when it is detected that the mouth and nose of the infant are blocked, the care robot is controlled to issue an alarm message.
[0125] Specifically, if the pre-processed image processing and analysis results indicate that an infant's mouth or nose is obstructed by an object (such as a blanket, stuffed toy, or arm), the robot's alarm mechanism is immediately triggered. This alarm can be sent locally via a variety of means, such as a sound alarm, light prompt, or an emergency notification sent to the parent's mobile app via the internet. This early warning of suffocation risks improves the safety of infants and young children at night or when they are unattended, effectively preventing accidents.
[0126] In one embodiment, after S6, the method further includes:
[0127] S71: When it is detected that the mouth and nose of the infant are blocked, a control instruction of the robotic arm of the care robot is obtained;
[0128] Specifically, if the infant's mouth and nose are detected to be obstructed, the care robot's response process will be triggered. At this time, the system's central control unit generates a control strategy for adjusting the shooting angle and issues corresponding robotic arm control instructions. These robotic arm control instructions may include parameters such as the initial rotation direction, angle range, rotation step value, and detection priority (such as prioritizing restoring mouth and nose visibility or searching the chest area), ensuring that subsequent actions are targeted and efficient.
[0129] S72: According to the robotic arm control instruction and the preset rotation strategy, the current posture of the infant is adjusted by the robotic arm to obtain an image of the infant after the posture adjustment;
[0130] Specifically, after receiving the robotic arm control command, the care robot's robotic arm dynamically adjusts the infant's position according to a preset rotation strategy. This strategy predetermines the rotation axis (such as yaw rotation around the vertical axis and pitch angle adjustment) based on the infant's current posture and the known environment layout, gradually modifying the camera's viewing angle to enable observation and image capture from different angles. With each rotation, the camera captures a new image frame, which serves as input for subsequent recognition and analysis, ensuring real-time feedback and a high response rate.
[0131] S73: Performing mouth and nose front detection and chest cavity detection on the infant image to obtain mouth and nose front detection results and chest cavity detection results;
[0132] Specifically, for the adjusted image frames, a dual-branch recognition network is used to perform mouth and nose frontal detection and chest area detection. Mouth and nose frontal detection is based on facial key point positioning and facial orientation estimation to determine whether the infant's mouth and nose are clearly visible in the current image, with particular attention paid to ensuring that the mouth and nose are unobstructed, complete, and at an appropriate angle. Chest detection identifies the edge contours and dynamic changes of the infant's chest (such as weak, fluctuating breathing signals) to determine whether the infant's chest has been successfully located. The detection results are assigned a confidence score to guide subsequent decision-making.
[0133] S74: performing feedback adjustment on the preset rotation strategy based on the mouth and nose front detection results and the chest cavity detection results, controlling the robotic arm to adjust the posture of the infant based on the adjusted rotation strategy, and determining the adjusted posture of the infant;
[0134] Specifically, based on the results of the current round of detection, if the front of the mouth and nose and the chest area have not been detected simultaneously, the rotation strategy will be adjusted in a feedback manner. The feedback mechanism includes: modifying the rotation direction (such as changing to reverse rotation or switching to another rotation axis); reducing or expanding the angle step value to improve search efficiency or accuracy; setting the current direction as an invalid direction and excluding it to prevent repeated invalid detection. The adjusted strategy will update the control instructions of the robotic arm, so that it continues to move according to the new rotation strategy, thereby gradually approaching the optimal viewing angle and finally determining the adjusted infant posture.
[0135] S75: According to the adjusted infant posture, when the front of the mouth and nose and the chest cavity are successfully detected, the robotic arm stops adjusting the infant posture.
[0136] Specifically, when the front of the infant's mouth and nose is successfully detected in one or more consecutive frames of images, and the chest area is located, and the corresponding confidence level meets the preset threshold, the current position of the robotic arm is determined to be the optimal detection position. At this point, the posture adjustment process of the current robotic arm is terminated, and the current angle and infant posture are recorded as the reference detection posture for initialization configuration in subsequent monitoring tasks. Through this process, it is ensured that the appropriate monitoring perspective can be automatically restored in the event that the infant rotates or is blocked, thereby improving the intelligence, adaptability and safety of the overall detection system.
[0137] Example 2
[0138] See Figure 8 Embodiment 2 of the present invention further provides an infant care device based on mouth and nose occlusion detection, the device comprising:
[0139] A real-time image acquisition module is used to acquire real-time video data collected by the care robot in the infant sleep care scene, and decompose the real-time video data into multiple frames of real-time images;
[0140] A face detection module is used to convert each frame of real-time image into a grayscale image, perform infant face detection on the grayscale image, and obtain the location information of the infant's face area;
[0141] An oropharyngeal region positioning module, configured to perform oropharyngeal region positioning on the infant's face region position information to obtain target oropharyngeal region position information;
[0142] A depth information acquisition module, configured to acquire depth information of the mouth and nose region based on the target position information of the mouth and nose region;
[0143] an occlusion detection module, configured to detect occlusion of the infant's mouth and nose based on the depth information of the mouth and nose area, and output a detection result;
[0144] The alarm module is used to control the care robot to issue an alarm message when it detects that the mouth and nose of the infant are blocked according to the detection result.
[0145] Specifically, an infant care device based on mouth and nose occlusion detection provided by an embodiment of the present invention is adopted, and the device includes: a real-time image acquisition module, used to acquire real-time video data collected by a care robot in an infant sleep care scene, and decompose the real-time video data into multiple frames of real-time images; a face detection module, used to convert each frame of real-time image into a grayscale image, perform infant face detection on the grayscale image, and obtain infant face area position information; a mouth and nose area positioning module, used to perform mouth and nose area positioning on the infant face area position information, and obtain mouth and nose area target position information; a depth information acquisition module, used to obtain mouth and nose area depth information based on the mouth and nose area target position information; an occlusion detection module, used to detect infant mouth and nose occlusion based on the mouth and nose area depth information, and output a detection result; an alarm module, used to control the care robot to issue an alarm message when it is detected that the infant's mouth and nose are blocked based on the detection result. This device uses an infant face detection algorithm to accurately locate the facial area, ensuring the precise positioning of the mouth and nose area, thereby improving the accuracy of occlusion detection. At the same time, the acquisition of depth information further enhances the ability to identify occlusion situations, and can maintain efficient detection under different ambient lighting and facial postures. By combining these steps, the present invention can provide timely and accurate mouth and nose occlusion detection, effectively protecting the safety of infants and young children.
[0146] Example 3
[0147] In addition, combined Figure 1 The infant care method based on mouth and nose occlusion detection of the first embodiment of the present invention can be implemented by an infant care robot. Figure 9 A schematic diagram of the hardware structure of the infant care robot provided in Example 3 of the present invention is shown.
[0148] The infant care robot may include a processor and a memory storing computer program instructions.
[0149] Specifically, the processor may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits for implementing the embodiments of the present invention.
[0150] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include a removable or non-removable (or fixed) medium. Where appropriate, the memory may be inside or outside the data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0151] The processor reads and executes computer program instructions stored in the memory to implement any one of the infant care methods based on mouth and nose occlusion detection in the above embodiments.
[0152] In one example, the infant care robot may further include a communication interface and a bus. Figure 9 As shown, the processor, memory, and communication interface are connected via a bus and communicate with each other.
[0153] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.
[0154] Bus comprises hardware, software or both, couples the parts of described equipment together.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.
[0155] Example 4
[0156] In addition, in conjunction with the infant care method based on mouth and nose occlusion detection in the first embodiment above, the fourth embodiment of the present invention may also be implemented by providing a computer-readable storage medium. The computer-readable storage medium stores computer program instructions; when executed by a processor, the computer program instructions implement any of the infant care methods based on mouth and nose occlusion detection in the above embodiments.
[0157] In summary, the embodiments of the present invention provide an infant care method, device, and robot based on mouth and nose occlusion detection.
[0158] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0159] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0160] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant location, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0161] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.
[0162] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.
Claims
1. A method for caring for infants and young children based on mouth and nose occlusion detection, characterized in that: The method comprises: S1: Acquire real-time video data collected by a care robot in a sleep care scene for infants and young children, and decompose the real-time video data into multiple frames of real-time images; S2: Convert each frame of real-time image into a grayscale image, perform infant face detection on the grayscale image, and obtain infant face area position information; S3: Positioning the mouth and nose region of the infant's face region to obtain target position information of the mouth and nose region; S4: Acquire depth information of the mouth and nose region based on the target position information of the mouth and nose region; S5: Detecting occlusion of the infant's mouth and nose based on the mouth and nose area depth information and the mouth and nose area target position information, and outputting a detection result; S6: According to the detection result, when it is detected that the mouth and nose of the infant are blocked, the care robot is controlled to issue an alarm message.
2. The infant care method based on mouth and nose occlusion detection according to claim 1, characterized in that: The S3 includes: S31: Enlarging the position information of the infant's face region to obtain an image of the infant's head region; S32: Inputting the infant head region image into a pre-trained hair detection model, and outputting the infant hair region position information; S33: Determine the target position information of the mouth and nose area based on the position information of the infant's face area and the position information of the infant's hair area.
3. The infant care method based on mouth and nose occlusion detection according to claim 2, characterized in that: The S33 includes: S331: Acquire first center position information of the infant's face region and second center position information of the infant's hair region based on the infant's face region position information and the infant's hair region position information; S332: Connecting a first center point of the infant's face region and a second center point of the infant's hair region based on the first center position information and the second center position information, and outputting a connecting line direction as a direction of the mouth and nose region; S333: Determine initial position information of the mouth and nose region based on the direction of the mouth and nose region and the position information of the infant's face region; S334: Extending the width of the initial mouth and nose area according to the initial position information of the mouth and nose area to determine the target position information of the mouth and nose area.
4. The infant care method based on mouth and nose occlusion detection according to claim 1, characterized in that: The S4 includes: S41: Inputting the real-time image into a pre-trained depth calculation model to output real-time depth information; S42: Determine the depth information of the mouth and nose area based on the target position information of the mouth and nose area and the real-time depth information.
5. The infant care method based on mouth and nose occlusion detection according to any one of claims 1 to 4, characterized in that: The S5 includes: S51: Obtaining the width of the mouth and nose area according to the target position information of the mouth and nose area; S52: Determine a depth value sequence of each column of pixel points in the mouth and nose region according to the width of the mouth and nose region and the depth information of the mouth and nose region; S53: Processing the depth value sequence to determine whether it meets the mouth and nose occlusion requirements; S54: If the depth value sequence meets the mouth and nose occlusion requirement, it is detected that the infant's mouth and nose are occluded, and a safety reminder is issued; S55: If the depth value sequence does not meet the mouth and nose occlusion requirement, it is detected that the infant's mouth and nose are not occluded, and no safety reminder is issued.
6. The infant care method based on mouth and nose occlusion detection according to claim 5, characterized in that: The S53 includes: S531: performing averaging processing on each depth value sequence to determine a first depth mean value; S532: performing difference calculation on the first depth mean values of adjacent columns to obtain depth differences; S533: averaging the depth differences to determine a second depth mean; S534: Obtain the maximum depth difference among the depth differences, and determine whether the mouth and nose occlusion requirement is met based on the maximum depth difference and the second depth average.
7. The infant care method based on mouth and nose occlusion detection according to claim 6, characterized in that: The S534 includes: S5341: Compare the maximum depth difference with the second depth mean, and obtain the number of maximum depth differences greater than the second depth mean; S5342: Compare the number of the maximum depth differences with a preset threshold. If the number of the maximum depth differences is greater than the preset threshold, the mouth and nose occlusion requirement is met. S5343: If the number of maximum depth differences is less than or equal to the preset threshold, the mouth and nose occlusion requirement is not met.
8. An infant care device based on mouth and nose occlusion detection, characterized in that: The device comprises: A real-time image acquisition module is used to acquire real-time video data collected by the care robot in the infant sleep care scene, and decompose the real-time video data into multiple frames of real-time images; A face detection module is used to convert each frame of real-time image into a grayscale image, perform infant face detection on the grayscale image, and obtain the location information of the infant's face area; An oropharyngeal region positioning module, configured to perform oropharyngeal region positioning on the infant's face region position information to obtain target oropharyngeal region position information; A depth information acquisition module, configured to acquire depth information of the mouth and nose region based on the target position information of the mouth and nose region; an occlusion detection module, configured to detect occlusion of the infant's mouth and nose based on the depth information of the mouth and nose area, and output a detection result; The alarm module is used to control the care robot to issue an alarm message when it detects that the mouth and nose of the infant are blocked according to the detection result.
9. An infant care robot, characterized in that: include: At least one processor, at least one memory, and computer program instructions stored in the memory, which implement the method according to any one of claims 1 to 7 when the computer program instructions are executed by the processor.
10. A storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Children mouth and nose shielding detection method and terminal equipment
CN116959071A