A method for detecting the posture of an infant and a computer device

By using camera equipment or wearable devices to detect and record the posture of infants and young children, the problem of parents having difficulty adjusting their infants' sleeping posture is solved. This enables scientific posture monitoring and reminders, prevents poor head shape, and improves safety and health.

CN122454471APending Publication Date: 2026-07-24HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-01-22
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Parents often struggle to adjust their infants' sleeping positions scientifically and reasonably, leading to poor head shape. Existing technology is unable to effectively monitor and remind them to adjust their posture.

Method used

The system detects infants' postures using cameras or wearable devices, categorizes and records the duration of each posture, and issues reminders when the posture exceeds a preset duration, guiding caregivers to adjust their posture.

Benefits of technology

It enables scientific monitoring and reminders of infants' postures, helping caregivers adjust their postures appropriately, preventing poor head shape, and improving safety and health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454471A_ABST
    Figure CN122454471A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of infant posture detection method and computer equipment, can be applied computer vision technical field, this method includes: first, according to the video stream of real-time shooting of shooting device determination infant current posture, and in the case where determining that current posture belongs to safe posture (such as, sleep on back, left side sleep, right side sleep, lie on back, left side lie, right side lie), the duration of current posture is counted, when the duration reaches the first preset duration, it indicates that infant maintains the current posture too long, at this time, output first target instruction, for reminding guardian to adjust the posture of infant, so that guardian can clearly know the duration of various postures of infant, to facilitate guardian to scientifically and reasonably adjust the posture of infant, so that the duration of various postures is as reasonable and balanced as possible, and sleep out ideal head shape.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method for detecting infant posture and a computer device. Background Technology

[0002] Infants and young children are among the family members who require the most special care, and their health needs to be given particular attention.

[0003] In infancy, a child's skull is not yet fully ossified, especially in infants under 6 months old, whose skulls are relatively soft and highly malleable. As the child grows, the bones gradually calcify, and the skull gradually hardens. Therefore, in the first few months after birth, if a part of the head bears the weight of the entire head for an extended period, such as when the child is always lying on their back, the delicate occipital bone will be subjected to prolonged pressure, and its normal curvature will gradually be flattened. As the child grows and the bone calcifies, a flat head will result.

[0004] Generally speaking, side-lying with alternating sides is a safe and ideal posture for infants and toddlers, and it also results in a beautiful head shape. Additionally, given the active nature of infants and toddlers, pillows can be correctly selected and used to stabilize their bodies and heads during sleep. However, parents often find it difficult to know the duration of each sleeping position and thus cannot scientifically and reasonably adjust their infants' postures. Summary of the Invention

[0005] This application provides a method and computer device for detecting infant postures. The device categorizes the infant's current posture and records the duration of each posture when it falls within a safe posture (e.g., lying on their back, lying on their left side, lying on their right side, lying on their back, lying on their left side, lying on their right side, etc.). If the current posture is maintained for too long (i.e., reaching a first preset duration), a first target instruction is issued to remind the caregiver to adjust the infant's posture. This allows the caregiver to clearly know the duration of each infant's posture, facilitating the caregiver to scientifically and reasonably adjust the infant's posture, ensuring that the duration of each posture is as reasonable and balanced as possible, thus promoting an ideal head shape.

[0006] Based on this, the embodiments of this application provide the following technical solutions:

[0007] Firstly, this application provides a method for detecting an infant's posture. The method specifically includes: First, determining the infant's current posture based on a captured video stream, wherein the video stream is captured by one or more pre-deployed camera devices (e.g., webcams). For example, the camera device can be installed where the infant can clearly be captured (e.g., near a crib) and / or in areas where the infant frequently moves (e.g., on a living room sofa or carpet). Next, determining whether the infant's current posture is a safe or dangerous posture. If the current posture is determined to be a safe posture, the duration of the current posture is further calculated. This can be achieved directly through the camera device or using a wearable device (e.g., a wearable device with a gyroscope or pressure detection). This application does not limit the specific method used for duration calculation. Finally, determining whether the duration of the current posture reaches a first preset duration. If so, the current posture is considered to have been maintained for too long, and a first target instruction is output to remind the infant to adjust their posture. For example, the first target instruction can be to trigger the guardian's mobile phone, computer, or other target devices to vibrate or ring as a reminder, or the first target instruction can be sent to the guardian's mobile phone, computer, or other target devices in the form of text or voice as a reminder. For another example, the first target instruction can be to trigger the camera device to flash lights or make a sound as a reminder, or the first target instruction can be to trigger the display device associated with the camera device to display a reminder. This application does not limit the specifics of these methods.

[0008] In the above embodiments of this application, the current posture of the infant is classified, and the duration of each posture is recorded when it is a safe posture (e.g., lying on the back, lying on the left side, lying on the right side, lying on the back, lying on the left side, lying on the right side, etc.). When the duration reaches a first preset duration, it indicates that the infant has maintained the current posture for too long. At this time, a first target instruction is output to remind the guardian to adjust the infant's posture, so that the guardian can clearly know the duration of the infant's various postures, which makes it easier for the guardian to scientifically and reasonably adjust the infant's posture, so that the duration of each posture is as reasonable and balanced as possible, and to achieve an ideal head shape.

[0009] In one possible implementation of the first aspect, where the infant's posture includes the infant's sleeping position, the infant's current posture refers to the current sleeping position. In this case, the infant's safe posture may include, but is not limited to: sleeping on their back, sleeping on their left side, and sleeping on their right side.

[0010] In the above embodiments of this application, the safe sleeping postures are specifically described. By clearly classifying the sleeping postures of infants and young children, scientific guidance is provided for adjusting the sleeping postures of infants and young children.

[0011] In one possible implementation of the first aspect, determining the infant's current posture based on the video stream can be achieved as follows: First, a first detection is performed on the infant's state based on the video stream to obtain a first detection result. It is then determined whether the first detection result satisfies a first condition. If so, the infant's state is determined to be a sleep detection state, whereby the sleep detection state triggers a second detection of the infant's state. Next, a second detection is performed on the infant's state based on the video stream to obtain a second detection result. It is then determined whether the second detection result satisfies a second condition. If so, the infant's state is determined to be a sleep state (i.e., the infant is asleep). Finally, the infant's current sleeping posture is determined based on this sleep state.

[0012] In the above embodiments of this application, two detection steps, a first detection and a second detection, are used to determine whether an infant is in a sleep state. The first detection is a coarse detection, used by the user to determine whether the infant has begun to fall asleep. The second detection is a fine detection, used to further determine whether the infant has truly fallen asleep. This is because when an infant is just about to fall asleep, they are in light sleep, and their body will have more unconscious activities (such as unconscious turning over, startle reflex, etc.). At this time, there is a relatively high probability of failing to fall asleep. Therefore, at this stage, a coarse detection process is used to determine whether the infant is more likely to "successfully fall asleep" or "fail to fall asleep" (that is, by judging whether the result of the first detection meets the first condition; if it does, it is considered that the infant is more likely to fall asleep successfully). This avoids directly entering the more refined second detection (because the second detection requires more information and consumes more computing resources), thus improving detection efficiency and real-time performance.

[0013] In one possible implementation of the first aspect, the initial detection of the infant's state based on the video stream can be achieved as follows: First, a first sliding window of a specified length is initialized for preliminary detection of the infant's current state. This first sliding window slides forward as the video stream updates in real time. Since the length of the first sliding window is fixed, the number of target image frames included in the video stream within the first sliding window is also fixed, which is n frames. Based on this, n frames of images in the video stream at each moment can be determined, and these n target images are the images included in the first sliding window at the current moment. Then, the first detection of the infant's state in the n target images included in the first sliding window is performed based on a trained detector.

[0014] In the above embodiments of this application, it is specifically described that the first detection of the infant's state based on the video stream is to detect the infant's state in each frame of the target image in the first sliding window, so that the detection has comprehensive coverage.

[0015] In one possible implementation of the first aspect, the method for performing the first detection of the infant's state in n frames of target images based on a trained detector can be as follows: First, the head position of the infant in the n frames of target images is detected using the trained detector, resulting in n detection boxes, which can be called first detection boxes, with one first detection box corresponding to each frame of target images. Then, the state of the infant in the i-th frame of target images is determined based on the first detection box corresponding to the i-th frame of target images and the first detection box corresponding to the (i-1)-th frame (i.e., the previous frame) of target images, where n ≥ i ≥ 2.

[0016] In the above embodiments of this application, the process of first detecting the state of an infant based on a video stream is specifically described. It is a detector based on deep learning. It combines the detection boxes of the infant's head position in the current frame and the previous frame to determine whether the infant is "stationary" or "moving" in the current frame. For each target image, it combines the target images of historical frames to determine the state of the infant in the current target image, which has high continuity and high real-time performance.

[0017] In one possible implementation of the first aspect, one way to determine the state of the infant in the target image i based on the first detection box corresponding to the target image i and the first detection box corresponding to the target image i-1 can be: to determine the state of the infant in the target image i (e.g., whether it is in a static state or a moving state) by calculating the intersection of union (IoU) value between the first detection box corresponding to the target image i and the first detection box corresponding to the target image i-1.

[0018] In the above embodiments of this application, it is specifically described that the state of the infant in each frame of the target image is determined by calculating the IoU value of the first detection box corresponding to the two frames of target images. The calculation method is simple and time-saving.

[0019] In one possible implementation of the first aspect, the method for determining whether the first detection result satisfies the first condition may include, but is not limited to: (1) determining that the first detection result satisfies the first condition when the ratio between the number of first target images and the number of second target images reaches a first preset threshold. Wherein, the first target image and the second target image belong to the n-frame target images, the first target image is a target image where the infant's state is static, and the second target image is a target image where the infant's state is dynamic. The static state is used to characterize that the infant is in a static state, and the dynamic state is used to characterize that the infant is in a dynamic state. (2) determining that the first detection result satisfies the first condition when the difference between the number of first target images and the number of second target images reaches a second preset threshold.

[0020] In the above embodiments of this application, several forms of determining that the first detection result meets the first condition are specifically described, which have flexibility and wide adaptability.

[0021] In one possible implementation of the first aspect, the second detection of the infant's state based on the video stream can be implemented as follows: First, a second sliding window of a specified length is initialized for further detection of the infant's current state. This second sliding window of the specified length slides forward as the video stream is updated in real time. Since the length of the second sliding window is fixed, the number of target image frames in the video stream included in the second sliding window is also fixed, which is m frames. Based on this, m frames of images in the video stream at each moment can be determined, and these m target images are the images included in the second sliding window at the current moment. Then, the second detection of the infant's state in the m target images included in the second sliding window is performed based on a trained detector.

[0022] In the above embodiments of this application, it is specifically described that the second detection of the infant's state based on the video stream is to detect the infant's state in each frame of the target image in the second sliding window, so that the detection has comprehensive coverage.

[0023] In one possible implementation of the first aspect, the second detection of the infant's state in m frames of target images based on a trained detector can be achieved as follows: First, the human body position of the infant in the m frames of target images is detected using the trained detector, resulting in m detection boxes, which can be called second detection boxes, with one second detection box corresponding to each frame of target images. Then, the state of the infant in the j-th frame of target images is determined based on the second detection box corresponding to the j-th frame of target images and the second detection box corresponding to the (j-1)-th frame (i.e., the previous frame) of target images, where m ≥ j ≥ 2.

[0024] In the above embodiments of this application, the process of performing a second detection of the infant's state based on the video stream is specifically described. It is a detector based on deep learning, which combines the detection boxes of the infant's body position in the current frame and the previous frame to determine whether the infant is "stationary" or "moving" in the current frame. For each target image, the target images of historical frames are combined to determine the infant's state in the current target image, which has high continuity and high real-time performance.

[0025] In one possible implementation of the first aspect, one way to determine the state of the infant in the target image j based on the second detection box corresponding to the target image j and the second detection box corresponding to the target image j-1 can be: to determine the state of the infant in the target image j (e.g., whether it is a static state or a moving state) by calculating the IoU value between the second detection box corresponding to the target image j and the second detection box corresponding to the target image j-1.

[0026] In the above embodiments of this application, it is specifically described that the state of the infant in each frame of the target image is determined by calculating the IoU value of the second detection box corresponding to the two frames of target images. The calculation method is simple and time-saving.

[0027] In one possible implementation of the first aspect, the method for determining whether the second detection result satisfies the second condition may include, but is not limited to: (1) determining that the second detection result satisfies the second condition when the ratio between the number of third target images and the number of fourth target images reaches a third preset threshold. Wherein, the third target image and the fourth target image belong to m-frame target images, the third target image is a target image where the infant's state is static, and the fourth target image is a target image where the infant's state is moving. Static state is used to characterize that the infant is in a static state, and moving state is used to characterize that the infant is in a moving state. (2) determining that the second detection result satisfies the second condition when the difference between the number of third target images and the number of fourth target images reaches a fourth preset threshold.

[0028] In the above embodiments of this application, several forms of determining that the second detection result meets the second condition are specifically described, which have flexibility and wide adaptability.

[0029] In one possible implementation of the first aspect, determining the infant's current sleeping position based on the sleep state can be as follows: First, an expandable window is initialized (i.e., the length of the window can be adaptively adjusted). When it is determined that the infant is in a sleeping state, the current sleeping position of the infant in the current frame target image of the video stream is identified according to the trained classifier. The current frame target image belongs to k frame target images, which are images included in the preset expandable window. The state of the infant in the k frame target images is all in a sleeping state.

[0030] In the above embodiments of this application, when an infant is in a sleep state, an expandable window is initialized, and the infant's sleeping position (e.g., prone, supine, left side, right side) in each frame of the target image in the expandable window is identified based on a deep learning classifier. The algorithm based on the deep learning classifier is simple to deploy and has a short inference process.

[0031] In one possible implementation of the first aspect, after determining that the infant's state is a sleep state, the method may further include: continuing to perform a second detection on the infant's state based on the video stream, and the detection result obtained at this time may be called a third detection result. When the third detection result does not meet the second condition (e.g., the ratio of the number of "moving" frames to the number of "still" frames of the infant exceeds a preset threshold), it means that the infant may be awake (but not necessarily actually awake), and at this time the infant's state is determined to enter the coarse detection sleep state.

[0032] In the above embodiments of this application, after the infant enters a sleep state, the infant's state needs to be continuously monitored through a second detection. When the infant's state does not meet the second condition, it means that the infant may have woken up (but not necessarily actually woken up), and at this time the infant's state enters the coarse detection sleep state. This achieves real-time detection of the infant's state, allowing for timely adjustments to the infant's current state, resulting in high real-time performance.

[0033] In one possible implementation of the first aspect, after determining that the infant's state is a sleep detection state, the method may further include: continuing to perform a first detection on the infant's state based on the video stream, the detection result obtained at this time may be called a fourth detection result, when the fourth detection result does not meet the first condition (e.g., the difference between the number of "moving state" frames and the number of "still state" frames of the infant exceeds a preset threshold), then it is determined that the infant has really woken up, and at this time the infant's state is determined to be a moving state.

[0034] In the above embodiments of this application, after the infant's state is reset to the sleep detection state, in order to further determine whether the infant is truly awake, it is necessary to continuously monitor the infant's state through the first detection. When the infant's state continues not to meet the first condition, it is determined that the infant has truly woken up, and at this time the infant's state is in a dynamic state. This can improve the detection accuracy.

[0035] In one possible implementation of the first aspect, when the infant's posture includes a lying position, the infant's current posture refers to the current lying position. In this case, the infant's safe posture may include, but is not limited to: supine, left-side lying, and right-side lying.

[0036] In the above embodiments of this application, the safe postures in the lying position are specifically described. By clearly classifying the lying positions of infants and young children, scientific guidance is provided for adjusting the lying positions of infants and young children.

[0037] In one possible implementation of the first aspect, after determining the infant's current posture based on the video stream, the method may further include: if it is determined that the current posture is a dangerous posture and the duration of the dangerous posture reaches a second preset duration, outputting a second target instruction, the second target instruction being used to remind the infant of a safety risk (e.g., suffocation risk).

[0038] In the above embodiments of this application, it is specifically described that when the infant's current posture is a dangerous posture and it lasts for a certain period of time, the guardian is reminded that there is a safety risk to the infant (e.g., the risk of suffocation when sleeping on their stomach), thereby improving safety.

[0039] In one possible implementation of the first aspect, dangerous postures include: lying prone, or prone position.

[0040] In the above embodiments of this application, dangerous postures are clearly defined to facilitate accurate identification.

[0041] In one possible implementation of the first aspect, after calculating the duration of the current posture, the method may further include: displaying the duration on a target device (e.g., a guardian's mobile phone, computer, etc.).

[0042] In the above embodiments of this application, by displaying the duration of the current posture on the target device, it is convenient for the guardian to check the infant's posture in real time.

[0043] In one possible implementation of the first aspect, after calculating the duration of the current posture, the method may further include: automatically generating a posture adjustment strategy based on the historical posture statistics of the infant in the historical video stream and the duration of the current posture. The posture adjustment strategy may be included in the first target instruction and sent to the guardian's mobile phone or computer or other target device, so that the guardian can directly adjust the infant's posture based on the posture adjustment strategy.

[0044] In the above embodiments of this application, a scientific and reasonable posture adjustment strategy is automatically generated by combining the statistics of historical video streams and the current video stream, thereby improving the user experience.

[0045] A second aspect of this application provides a computer device having the function of implementing the method of the first aspect or any possible implementation thereof. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.

[0046] A third aspect provides a computer device that may include a memory, a processor, and a bus system, wherein the memory is used to store a computer program (also referred to as a program or computer-readable instructions), and the processor is used to invoke the program stored in the memory to execute the method of the first aspect of the present application or any possible implementation thereof.

[0047] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0048] The fifth aspect of this application provides a computer program or a computer program product containing instructions that, when the computer program or computer program product is run on a computer, causes the computer to perform the method described in the first aspect or any possible implementation of the first aspect.

[0049] A sixth aspect of this application provides a chip including at least one processor and at least one interface circuit coupled to the processor. The interface circuit performs transceiver functions and sends instructions to the at least one processor. The at least one processor runs a computer program or instructions, having the functionality to implement the methods described in the first aspect or any possible implementation of the first aspect. This functionality can be implemented in hardware, software, or a combination of hardware and software, including one or more modules corresponding to the described functions. Furthermore, the interface circuit is used to communicate with other modules outside the chip.

[0050] In some implementations of this application, some of the one or more processors may implement some steps of the above method through dedicated hardware. For example, the processing involving neural network models may be implemented by a dedicated neural network processor or graphics processor.

[0051] The method provided in this application embodiment can be implemented by a single chip or by multiple chips working together. Attached Figure Description

[0052] Figure 1 A schematic diagram of the main framework of artificial intelligence provided in the embodiments of this application;

[0053] Figure 2 A system architecture diagram of the task processing system provided in this application embodiment;

[0054] Figure 3 A flowchart illustrating an infant posture detection method provided in an embodiment of this application;

[0055] Figure 4 A schematic diagram of a sliding window provided in an embodiment of this application;

[0056] Figure 5 An example implementation diagram of the second target instruction provided in the embodiments of this application;

[0057] Figure 6 Another implementation example diagram of the second target instruction provided in the embodiments of this application;

[0058] Figure 7 An example implementation diagram of the first target instruction provided in the embodiments of this application;

[0059] Figure 8 Another implementation example diagram of the first target instruction provided in the embodiments of this application;

[0060] Figure 9 A schematic diagram of a computer device provided in an embodiment of this application;

[0061] Figure 10 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0062] This application provides a method and computer device for detecting infant postures. The method classifies the infant's current posture and records the duration of each posture when it falls within a safe posture (e.g., lying on their back, lying on their left side, lying on their right side, supine, lying on their left side, lying on their right side, etc.). If the current posture is maintained for too long (i.e., reaching a first preset duration), a first target instruction is issued to remind the caregiver to adjust the infant's posture. This allows the caregiver to clearly know the duration of each infant's posture, facilitating scientific and reasonable adjustments to the infant's posture, ensuring that the duration of each posture is as reasonable and balanced as possible, leading to an ideal head shape.

[0063] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0064] To better understand the solutions of the embodiments of this application, the relevant terms and concepts that may be involved in the embodiments of this application will be introduced below. It should be understood that the explanation of the relevant concepts may be limited due to the specific circumstances of the embodiments of this application, but it does not mean that this application can only be limited to that specific situation. The specific circumstances of different embodiments may also differ, and no specific limitation is made here.

[0065] (1) Neural Network

[0066] A neural network can be composed of neural units, specifically understood as a neural network with input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Neural networks with many hidden layers are called deep neural networks (DNNs). The function of each layer in a neural network can be expressed mathematically. To describe it physically, each layer in a neural network can be understood as transforming the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations are: 1. Dimensionality increase / decrease; 2. Magnification / scaling; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are... Operation 4 is completed using "+b", and operation 5 is implemented using "a()". The term "space" is used here because the objects being classified are not individual things, but a class of things; space refers to the set of all individuals within this class of things. Here, W is the weight matrix of each layer of the neural network, where each value represents the weight of a neuron in that layer. This matrix W determines the spatial transformation from the input space to the output space, as described above; that is, the W of each layer of the neural network controls how the space is transformed. The purpose of training the neural network is to ultimately obtain the weight matrices of all layers of the trained neural network. Therefore, the training process of a neural network is essentially learning how to control spatial transformation, more specifically, learning the weight matrix.

[0067] It should be noted that in the embodiments of this application, the detectors and classifiers used for machine learning tasks (such as active learning, supervised learning, unsupervised learning, semi-supervised learning, etc.) are essentially neural networks.

[0068] (2) Loss Function

[0069] During neural network training, to ensure the output closely approximates the desired predicted value, we compare the network's current prediction with the target value. Based on the difference, we update the weight matrix of each layer (usually pre-configuring parameters before the initial update). For example, if the predicted value is too high, the weight matrix is ​​adjusted to predict a lower value, and this process continues until the neural network accurately predicts the target value. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are crucial equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (Loss) indicates a greater difference, and training the neural network becomes a process of minimizing this loss.

[0070] During the training of a neural network, the back propagation (BP) algorithm can be used to correct the parameters in the initial neural network model, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates error loss. This error loss information is then propagated back to update the parameters in the initial neural network model, leading to convergence of the error loss. The back propagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the neural network model, such as the weight matrix.

[0071] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0072] First, the overall workflow of the artificial intelligence system is described; please refer to [link / reference]. Figure 1 , Figure 1 The diagram illustrates a structural framework for artificial intelligence (AI). The framework is further elaborated below along two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that AI brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed through technological means) to the industrial ecosystem of the system.

[0073] (1) Infrastructure

[0074] Infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. This communication occurs through sensors; computing power is provided by intelligent chips (hardware acceleration chips such as CPUs, NPUs, GPUs, ASICs, and FPGAs); and the basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.

[0075] (2) Data

[0076] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0077] (3) Data processing

[0078] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.

[0079] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training of data by symbolizing and formalizing it.

[0080] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.

[0081] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.

[0082] (4) General ability

[0083] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0084] (5) Smart Products and Industry Applications

[0085] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They encapsulate overall artificial intelligence solutions, productize intelligent information decision-making, and realize practical applications. Their application areas mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, autonomous driving, and smart cities.

[0086] The detector and classifier described in this application embodiment can be applied to the detection of infant postures. Specifically, in conjunction with... Figure 1 In this application embodiment, the data in the infrastructure acquisition dataset can be multiple image or video data (also called training data or training samples, multiple training data constitute a training set) acquired through camera devices (e.g., cameras, video recorders, etc.). The detector and classifier are trained using the training set to obtain the trained detector and classifier, respectively. For example, the detector in this application embodiment can be a YOLO (You Only Look Once) series detector, and the classifier can be a classifier based on the Contrastive Language-Image Pretraining (CLIP) framework. Specifically, this application does not limit the type of detector and classifier used.

[0087] The architecture of the task processing system will be described next; please refer to [link / reference]. Figure 2 , Figure 2 This is a system architecture diagram of a task processing system provided in an embodiment of this application. Figure 2 In this application, the task processing system 200 includes an execution device 210, a training device 220, a database 230, a client device 240, a data storage system 250, and a data acquisition device 260. The execution device 210 includes a computing module 211. The data acquisition device 260 acquires the large-scale open-source dataset (i.e., training set) required by the user and stores it in the database 230. The training device 220 trains the detector 201 and classifier 202 required by this application based on the training set maintained in the database 230. The trained detector 201 and classifier 202 are then applied on the execution device 210. It is important to note that since the training objectives of detector 201 and classifier 202 are different, the training sets maintained in database 230 can include two sets: training set 1 and training set 2. Training set 1 is used to train detector 201 so that the trained detector 201 can accurately identify targets in the image (such as the position of an infant's head, the position of an infant's body, etc.). Training set 2 is used to train classifier 202 so that the trained classifier 202 can accurately classify the pose of an infant in the image.

[0088] The execution device 210 can access data, code, etc., in the data storage system 250, and can also store data, instructions, etc., in the data storage system 250. The data storage system 250 can be located within the execution device 210, or it can be an external memory relative to the execution device 210.

[0089] The trained detector 201 and classifier 202, trained by training device 220, can be applied to different systems or devices (i.e., execution device 210), specifically edge devices or end-device devices, such as mobile phones, tablets, laptops, and camera devices (e.g., webcams). Figure 2 In this embodiment, the execution device 210 is equipped with an I / O interface 212 for data interaction with external devices. The "user" can input data to the I / O interface 212 through the client device 240. For example, the client device 240 can be a camera device. The video stream captured by the camera device is input to the computing module 211 of the execution device 210. The computing module 211 detects the input video stream and obtains detection and classification results. Based on these detection and classification results, the infant's posture is detected. Furthermore, in some embodiments of this application, the client device 240 can also be integrated into the execution device 210. For example, when the execution device 210 is a camera device, the video stream can be directly captured by the camera device, and the computing module 211 within the camera device can then perform detection, classification, and other processes on the video stream to detect the infant's posture. The product form of the execution device 210 and the client device 240 is not limited here.

[0090] It is worth noting that Figure 2 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in Figure 2 In this context, the data storage system 250 is an external memory relative to the execution device 210; however, in other cases, the data storage system 250 may be placed within the execution device 210. Figure 2 In this context, the client device 240 is an external device relative to the execution device 210. In other cases, the client device 240 may also be integrated into the execution device 210.

[0091] It should also be noted that the training of the detector 201 and classifier 202 described in this application embodiment can be implemented on the cloud side. For example, a training set can be obtained by a cloud-side training device 220 (which can be set on one or more servers or virtual machines), and the detector 201 and classifier 202 can be trained based on the training set to obtain the trained detector 201 and classifier 202. Then, the trained detector 201 and classifier 202 are sent to the execution device 210 for application, for example, to the execution device 210 for infant posture detection. Figure 2 The corresponding system architecture describes a system where the training device 220 trains the detector 201 and classifier 202, and the trained detector 201 and classifier 202 are then sent to the execution device 210 for use. The training of the detector 201 and classifier 202 described in the above embodiment can also be implemented on the terminal side, meaning the training device 220 can be located on the terminal side. For example, a terminal device (e.g., a camera) can acquire a training set and train the detector 201 and classifier 202 based on the training set, resulting in the trained detector 201 and classifier 202. This trained detector 201 and classifier 202 can then be used directly on the terminal device or sent by the terminal device to other devices for use. Specifically, this application does not limit the device (cloud side or terminal side) on which the detector 201 and classifier 202 are trained or applied.

[0092] The following describes the infant posture detection method provided in the embodiments of this application. Please refer to [link / reference]. Figure 3 , Figure 3 A flowchart illustrating an infant posture detection method provided in this application embodiment may specifically include:

[0093] 301. Determine the infant's current posture based on the video stream, which is captured by pre-deployed camera equipment.

[0094] First, the infant's current posture is determined based on the captured video stream, which is obtained from pre-deployed camera equipment (e.g., a webcam). For example, the camera can be installed where it can clearly capture the infant's sleeping area (e.g., near the crib) and / or frequently used areas (e.g., on the living room sofa or carpet), ensuring the infant doesn't occupy too small a proportion of the frame; ideally, the camera should be positioned 1 to 2 meters away from the sleeping area. It's important to note that in this embodiment, the video stream is captured in real-time and updated continuously, not a fixed segment of video.

[0095] It should also be noted that in this embodiment, the camera device can be a single unit. For example, this single camera device can be deployed where the infant sleeps, because infants sleep for long periods, and this allows for the collection of as much data as possible about their sleeping postures (maintaining one posture for a long time during sleep has a greater impact). Alternatively, there can be multiple cameras, which can be deployed in various areas where the infant appears, making the collected posture information more comprehensive. Specifically, this application does not limit the number of camera devices.

[0096] It should be noted that in some embodiments of this application, the infant's posture includes, but is limited to, the infant's sleeping posture and the infant's lying posture. The infant's sleeping posture refers to the posture of the infant while asleep, and the infant's lying posture refers to the posture of the infant while awake. In the embodiments of this application, the method for determining the infant's current posture will vary slightly depending on the type of infant's posture. For ease of understanding, the following explanation uses the infant's sleeping posture and the infant's lying posture as examples to illustrate the method for determining the infant's current posture:

[0097] I. Infant and toddler postures: Situations where infants and toddlers sleep in the following positions

[0098] When "infant posture" refers to an infant's sleeping position, it means the infant is currently in that position. In this context, safe postures for infants include, but are not limited to: sleeping on their back, sleeping on their left side, and sleeping on their right side. Dangerous postures for infants include, but are not limited to: sleeping on their stomach, also known as prone sleeping.

[0099] At this point, one approach to determining the infant's current posture based on the video stream could be as follows: First, perform a first detection on the infant's state based on the video stream, obtain a first detection result, and determine whether the first detection result meets a first condition. If so, determine that the infant's state is a sleep detection state, whereby the sleep detection state triggers a second detection on the infant's state. Next, perform a second detection on the infant's state based on the video stream, obtain a second detection result, and determine whether the second detection result meets a second condition. If so, determine that the infant's state is a sleep state (i.e., the infant is asleep). Finally, determine the infant's current sleeping posture based on this sleep state.

[0100] In the method described above for determining the current posture of an infant in this application, two detection steps are used: a first detection and a second detection, to determine whether the infant is asleep. The first detection is a coarse detection, used by the user to determine whether the infant has begun to fall asleep. The second detection is a fine detection, used to further determine whether the infant has truly fallen asleep. This is because when an infant is just about to fall asleep, they are in light sleep, and their body will have more unconscious movements (such as unconscious turning over, startle reflex, etc.). At this time, there is a relatively high probability of failing to fall asleep. Therefore, at this stage, a coarse detection process is used to determine whether the infant is more likely to "successfully fall asleep" or "fail to fall asleep" (determined by whether the result of the first detection meets the first condition; if it does, it is considered that the infant is more likely to have successfully fallen asleep). This avoids directly entering the more fine second detection (because the second detection requires more information and consumes more computing resources), thus improving detection efficiency and real-time performance.

[0101] To further understand the two detection processes described above, they will be explained below:

[0102] A. The process of first detecting the condition of infants and young children based on video streams.

[0103] After obtaining the real-time video stream of the infant through the camera device, a first sliding window (which can be denoted as sliding window 1) of a specified length of N is initialized to initially detect the current state of the infant. For example, the current state of the infant can be initialized as "not asleep".

[0104] Specifically, the sliding window 1 of a specified length N slides forward as the video stream is updated in real time. Since the length of the sliding window 1 is fixed at N, the number of frames of the target image in the video stream included in the sliding window 1 is also fixed, let's say n frames. In this way, the first detection can be performed on the state of the infant in the n target images included in the sliding window 1 based on the trained detector.

[0105] In some embodiments of this application, the head position of the infant in the n frames of target images can be detected using a trained detector to obtain n detection boxes, which can be called first detection boxes, with one first detection box corresponding to each frame of target images. Then, the state of the infant in the i-th frame of target images is determined based on the first detection box corresponding to the i-th frame and the first detection box corresponding to the (i-1)-th frame (i.e., the previous frame), where n ≥ i ≥ 2. For example, the state of the infant in the i-th frame of target images (e.g., whether it is stationary or moving) can be determined by calculating the IoU value between the first detection box corresponding to the i-th frame and the first detection box corresponding to the (i-1)-th frame.

[0106] It should be noted that in this embodiment, the target image included in the sliding window 1 has a fixed number of frames, but the target image included in the sliding window 1 is updated in real time as the sliding window 1 slides.

[0107] For a clearer understanding of this point, please refer to the following document. Figure 4 The example shown is in Figure 4 In this application, it is assumed that the sliding window 1 includes 50 frames of the target image, and that the sliding window 1 slides at a certain rate, which can be customized and is not limited in this application. Figure 4 This is for illustrative purposes only. Assume that at time t1, the target images of the video stream included in sliding window 1 are frames 11 to 60 (for illustrative purposes only). At time t2, sliding window 1 slides once along the sliding direction, and at this time, the target images included in sliding window 1 are frames 13 to 62 (for illustrative purposes only). And so on, at time tx, sliding window 1 continues to slide once along the sliding direction, and at this time, the target images included in sliding window 1 are frames k to k+49 (for illustrative purposes only). Figure 4 It is known that the number of target image frames contained in the sliding window 1 is always 50, but as the window slides, the target images it contains are updated in real time. Therefore, when describing the processing of n frames of target images in the following sections of this application, it refers to processing the current n frames of target images contained in the sliding window 1 at the time being discussed (e.g., time t1, time t2), which will not be elaborated further.

[0108] It should also be noted that this embodiment discusses the case of a single infant in the video stream. Therefore, one frame of the target image in sliding window 1 corresponds to one first detection box. If the video stream includes two or more infants, then one frame of the target image in sliding window 1 corresponds to a group of first detection boxes. The number of first detection boxes in a group corresponds to the number of infants. For one infant, one frame of the target image still corresponds to one first detection box. Taking a video stream with three infants as an example: based on the trained detector, the head positions of the infants in the n frames of the target image are detected, resulting in n groups of first detection boxes. Each group of first detection boxes includes three first detection boxes (let's call them detection box 1, detection box 2, and detection box 3, corresponding to infant 1, infant 2, and infant 3 respectively). One frame of the target image corresponds to one group of first detection boxes. Specifically, one frame of the target image corresponds to detection box 1, detection box 2, and detection box 3. For one infant (infant 1), one frame of the target image still corresponds to one first detection box, i.e., detection box 1.

[0109] In summary, the state of the infant in the n target images included in the sliding window 1 at each time moment can be obtained through the above method. The state of the infant in each target image of the sliding window 1 at each time moment is the first detection result at that time moment.

[0110] Still with Figure 4 For example, at time t1, the state of the infant in each frame of the target image from frame 11 to frame 60 can be calculated using the above method. Assuming that statistical analysis shows the infant's state in frames 11 to 40 (30 frames in total) is "moving" (representing the infant is in motion), and the infant's state in frames 41 to 60 (20 frames in total) is "stationary" (representing the infant is stationary), this statistical result is the first detection result at time t1. Here, the target image where the infant's state is "stationary" is defined as the first target image, and the target image where the infant's state is "moving" is defined as the second target image. Therefore, the first detection result at time t1 can be simply described as: 30 frames of moving state + 20 frames of stationary state.

[0111] Similarly, at time t2, since only frames 61 and 62 are newly added target images to sliding window 1, at time t2, it is only necessary to calculate the state of the infants in frames 61 and 62 (assuming they are both "static") based on the above method. Since the length of sliding window 1 is fixed, the states of the infants in frames 11 and 12 also need to be removed. Therefore, the first detection result at time t2 is: the first detection result at time t1, excluding the states of the infants in frames 11 and 12, and adding the states of the infants in frames 61 and 62. From the first detection result at time t1, it can be seen that the states of the infants in frames 11 and 12 are both "moving". Therefore, the first detection result at time t2 can be summarized as: 28 moving frames + 22 static frames.

[0112] By following this pattern, we can obtain the first detection result corresponding to sliding window 1 at each time step.

[0113] Therefore, in some embodiments of this application, the methods for determining whether the first detection result meets the first condition may include, but are not limited to:

[0114] (1) If the ratio between the number of first target images and the number of second target images reaches a first preset threshold, the first detection result is determined to satisfy the first condition.

[0115] In this case, at any time, when the ratio between the number of "static" frames (which can be denoted as n1) and the number of "moving" frames (which can be denoted as n2) in sliding window 1 reaches a preset first threshold (which can be denoted as p1), that is, when the following equation (1) is satisfied, it is determined that the first detection result satisfies the first condition.

[0116]

[0117] (2) If the difference between the number of first target images and the number of second target images reaches a second preset threshold, the first detection result is determined to meet the first condition.

[0118] In this case, at any time, when the difference between the number of "static" frames (i.e., n1) and the number of "moving" frames (i.e., n2) in sliding window 1 reaches the second preset threshold (which can be denoted as q1), i.e. when the following equation (2) is satisfied, it is determined that the first detection result satisfies the first condition.

[0119] n1-n2≥q1, n1+n2=n (2)

[0120] B. The process of conducting a second detection of the infant's condition based on the video stream.

[0121] Once the infant's state is determined to be in a sleep state after the first coarse detection, a second sliding window (which can be denoted as sliding window 2) with a specified length of M is initialized to further detect the infant's current state.

[0122] Specifically, the sliding window 2 of a specified length M slides forward as the video stream is updated in real time. Since the length of the sliding window 2 is fixed at M, the number of frames of the target image in the video stream included in the sliding window 2 is also fixed, let's say m frames. In this way, the second detection can be performed on the state of the infant in the m target images included in the sliding window 2 based on the trained detector.

[0123] In some embodiments of this application, the position of the infant's body in the m-frame target images can be detected using a trained detector, resulting in m detection boxes, which can be called second detection boxes. One second detection box corresponds to one target image. Then, the state of the infant in the j-th frame target image is determined based on the second detection box corresponding to the j-th frame target image and the second detection box corresponding to the (j-1)-th frame (i.e., the previous frame), where m ≥ j ≥ 2. For example, the state of the infant in the j-th frame target image (e.g., whether it is stationary or moving) can be determined by calculating the IoU value between the second detection box corresponding to the j-th frame target image and the second detection box corresponding to the (j-1)-th frame target image.

[0124] Similar to sliding window 1, in this embodiment, the target image included in sliding window 2 has a fixed number of frames, but the target image included in sliding window 2 is updated in real time as sliding window 2 slides. See the above for details. Figure 4 The example of sliding window 1 will not be elaborated here.

[0125] Furthermore, it should be noted that this embodiment discusses the case where there is only one infant in the video stream. Therefore, one frame of the target image in sliding window 2 corresponds to one second detection box. If the video stream includes two or more infants, then one frame of the target image in sliding window 2 corresponds to a group of second detection boxes. The number of second detection boxes in a group corresponds to the number of infants. For one infant, one frame of the target image still corresponds to one second detection box. For details, please refer to the above description of sliding window 1, which will not be repeated here.

[0126] In summary, the state of the infant in the m target images included in the sliding window 2 at each time moment can be obtained in the above manner. The state of the infant in each target image of the sliding window 2 at each time moment is the second detection result at that time moment. The statistical method is similar to that of the first detection result. For details, please refer to the above description of the first detection result. It will not be repeated here.

[0127] Therefore, in some embodiments of this application, the methods for determining whether the second detection result meets the second condition may include, but are not limited to:

[0128] (1) When the ratio between the number of third target images and the number of fourth target images reaches the third preset threshold, the second detection result is determined to meet the second condition.

[0129] Among them, the third target image and the fourth target image belong to the m-frame target images. The third target image is the target image in which the infant is in a "static state", and the fourth target image is the target image in which the infant is in a "moving state".

[0130] In this case, at any time, when the ratio between the number of "static" frames (which can be denoted as m1) and the number of "moving" frames (which can be denoted as m2) in the sliding window 2 reaches the preset third threshold (which can be denoted as p2), that is, when the following equation (3) is satisfied, it is determined that the second detection result satisfies the second condition.

[0131]

[0132] (2) When the difference between the number of third target images and the number of fourth target images reaches the fourth preset threshold, the second detection result is determined to meet the second condition.

[0133] In this case, at any time, when the difference between the number of "static" frames (i.e., m1) and the number of "moving" frames (i.e., m2) in sliding window 2 reaches the preset fourth threshold (which can be denoted as q2), i.e. when the following equation (4) is satisfied, it is determined that the second detection result satisfies the second condition.

[0134] m1-m2≥q2, m1+m2=m (4)

[0135] It should be noted that in some implementations of the application, after the infant enters a sleep state, it is necessary to continue to perform a second detection on the infant's state based on the video stream. The detection result obtained at this time can be called the third detection result. When the third detection result does not meet the second condition (e.g., the ratio of the number of "moving" frames to the number of "still" frames of the infant exceeds a preset threshold), it means that the infant may have woken up (but not necessarily actually woken up). At this time, it is determined that the infant's state has entered the coarse detection sleep state.

[0136] It should also be noted that in some embodiments of this application, after the infant's state changes from sleep state to sleep detection state, in order to further determine whether the infant is truly awake, it is necessary to continue to perform a first detection on the infant's state based on the video stream. The detection result obtained at this time can be called the fourth detection result. When the fourth detection result does not meet the first condition (e.g., the difference between the number of "moving" frames and the number of "still" frames of the infant exceeds a preset threshold), it is determined that the infant has truly woken up, and the infant's state is determined to be moving.

[0137] Once the second precise detection confirms that the infant is asleep, the infant's current sleeping position can be determined based on the sleep state.

[0138] Specifically, an expandable window (i.e., the length of which can be adaptively adjusted) can be initialized, denoted as window Q. When it is determined that the infant is asleep, firstly, the current sleeping posture of the infant in the target image of the current frame of the video stream is identified according to the trained classifier. For example, an adaptive expanded detection box can be generated based on the first detection box determined by the detector. The expanded detection box is obtained by expanding the length and width of the first detection box according to a preset method, such as by expanding the length and width of the first detection box according to a preset ratio (e.g., the length is expanded to 1.5 times and the width to 2 times). The purpose is to include not only the infant's head position information, but also the environmental information around the head (e.g., pillow, bedside toys). Based on more environmental information, the accuracy of classification can be improved. Then, the classifier can identify the infant's current sleeping position based on the bounding box of the target image in each current frame, and record the current sleeping position of each frame in the window Q. Assuming that a total of k target images are counted in the entire infant's sleep state, then the window Q includes the k target images.

[0139] It's important to note that the length of window Q is determined by the infant's current sleep duration. Window Q begins when the infant begins to fall asleep and ends when they wake up. This allows us to track the duration of each sleep posture within window Q. It's also crucial to understand that because prone sleeping poses a suffocation risk, if the prone sleeping duration exceeds the preset time, a target command will be sent to alert the parents immediately, not necessarily at the end of window Q.

[0140] II. Infant / toddler posture: When the infant / toddler is lying down

[0141] When "infant posture" refers to an infant's lying position, it means the infant is currently in that position. In this context, safe postures for infants include, but are not limited to: lying on their back, lying on their left side, and lying on their right side. Dangerous postures for infants include, but are not limited to: lying prone, also known as prone lying.

[0142] At this point, determining the infant's current state based on the video stream does not require sleep detection. Instead, the infant's lying posture can be directly identified using the trained classifier. This identification process is similar to the classifier's sleep posture identification described above, and will not be repeated here.

[0143] 302. If the current posture is determined to be a safe posture, calculate the duration of the current posture.

[0144] Next, it is determined whether the infant's current posture is a safe posture or a dangerous posture. If it is determined that the current posture is a safe posture, the duration of the current posture is further recorded.

[0145] It should be noted that in some embodiments of this application, the duration of the infant's current posture can be directly recorded using a camera device, or a wearable device (such as a wearable device with a gyroscope or pressure detection) can be used to record the duration of the infant's current posture. This application does not limit the method of recording the duration.

[0146] It should also be noted that in some embodiments of this application, if it is determined that the current posture is a dangerous posture (e.g., lying on one's stomach), and the duration of the dangerous posture reaches a second preset duration, a second target instruction is output. The second target instruction is used to remind the infant that there is a safety risk (e.g., suffocation risk).

[0147] As an example, such as Figure 5 As shown, the second target instruction can be a text prompt message sent by the camera device to the guardian's mobile phone, such as... Figure 5 The message "xxx slept for 2 minutes, please check xxx's status promptly" can be displayed as a text message (e.g., ...). Figure 5 As shown in (a) above, it can also be displayed through a specific application (APP) deployed on the guardian's mobile phone (such as...). Figure 5 (as shown in (b)); as another example, such as Figure 6 As shown, the second target instruction can be a voice prompt sent by the camera device to the guardian's mobile phone, such as... Figure 6 The voice prompt shown reads, "xxx has been sleeping face down for 2 minutes, please check xxx's status promptly." As another example, this second objective instruction could also trigger a specific vibration on the guardian's phone to remind them to check the infant's status. This application does not specify the exact form of the second objective instruction.

[0148] It should also be noted that, in some embodiments of this application, the duration of the current posture can be displayed on the display interface of the target device (e.g., the guardian's mobile phone, computer, etc.) to facilitate the guardian to view the infant's posture in real time.

[0149] 303. If the duration reaches the first preset duration, output the first target instruction, which is used to remind the infant to adjust his posture.

[0150] Finally, it is determined whether the duration of the current posture has reached the first preset duration. If so, the current posture is considered to have been maintained for too long, and a first target instruction is output to remind the infant to adjust their posture. For example, the first target instruction can be triggered by the guardian's mobile phone, computer, or other target devices to vibrate or ring, or it can be sent to the guardian's mobile phone, computer, or other target devices in the form of text or voice. This application does not limit the specifics.

[0151] It should be noted that, in some embodiments of this application, a posture adjustment strategy can be automatically generated based on the historical posture statistics of the infant in the historical video stream and the duration of the current posture. The posture adjustment strategy can be included in the first target instruction and sent to the guardian's mobile phone or computer or other target device, so that the guardian can directly adjust the infant's posture based on the posture adjustment strategy.

[0152] As an example, such as Figure 7 As shown, the first target instruction can be a text prompt message sent by the camera device to the guardian's mobile phone, such as... Figure 7 The message "xxx has accumulated 12 hours of sleeping on their back, 20 hours of sleeping on their left side, and 10 hours of sleeping on their right side. It is recommended to increase xxx's right-side sleeping time" can be displayed as a text message (e.g., ...). Figure 7 As shown in (a) above, it can also be displayed through a specific app deployed on the guardian's mobile phone (such as...). Figure 7 (as shown in (b)); as another example, such as Figure 8 As shown, the first target instruction can be a voice prompt sent by the camera device to the guardian's mobile phone, such as... Figure 8 The voice prompt shown reads, "xxx has accumulated 12 hours of sleeping on their back, 20 hours on their left side, and 10 hours on their right side. It is recommended to increase xxx's right-side sleeping time." As another example, this first objective instruction could also trigger a specific vibration on the guardian's mobile phone to prompt the guardian to adjust the infant's posture in a timely manner. This application does not specifically limit the form of the first objective instruction.

[0153] Based on the above embodiments, in order to better implement the above solutions of this application, related equipment for implementing the above solutions is also provided below. See details. Figure 9 , Figure 9This is a schematic diagram of a computer device provided in an embodiment of this application. The computer device 900 may specifically include: a determination module 901, a statistics module 902, and a reminder module 903. The determination module 901 is used to determine the current posture of an infant based on a video stream captured by a pre-deployed camera device. The statistics module 902 is used to calculate the duration of the current posture when it is determined to be a safe posture. The reminder module 903 is used to output a first target instruction when the duration reaches a first preset duration. The first target instruction is used to remind the infant to adjust their posture.

[0154] In one possible design, where the infant posture includes the infant sleeping position, the current posture includes the current sleeping position, and the safe posture includes at least one of the following: sleeping on the back, sleeping on the left side, or sleeping on the right side.

[0155] In one possible design, the determining module 901 is specifically used for: performing a first detection on the infant's state based on the video stream, obtaining a first detection result, and determining the infant's state as a sleep detection state if the first detection result satisfies a first condition, the sleep detection state being used to trigger a second detection on the infant's state; performing a second detection on the infant's state based on the video stream, obtaining a second detection result, and determining the infant's state as a sleep state if the second detection result satisfies a second condition, the sleep state being used to characterize the infant being in a sleep state; and determining the infant's current sleeping posture based on the sleep state.

[0156] In one possible design, the determining module 901 is further configured to: determine n target images in the video stream, the n target images being images included in a preset first sliding window; and perform a first detection on the state of the infant in the n target images based on a trained detector.

[0157] In one possible design, the determination module 901 is further configured to: detect the head position of the infant in the n-frame target images based on the trained detector, and obtain n first detection boxes, one first detection box corresponding to one frame of target images; determine the state of the infant in the i-th frame target image based on the first detection box corresponding to the i-th frame target image and the first detection box corresponding to the (i-1)-th frame target image, where n≥i≥2.

[0158] In one possible design, the determining module 901 is further configured to: determine the state of the infant in the target image i based on the IoU value between the first detection box corresponding to the target image i and the first detection box corresponding to the target image i-1.

[0159] In one possible design, determining that the first detection result satisfies the first condition includes: determining that the first detection result satisfies the first condition when the ratio between the number of first target images and the number of second target images reaches a first preset threshold, wherein the first target image and the second target image belong to the n-frame target images, the first target image is a target image in which the infant is in a static state, and the second target image is a target image in which the infant is in a moving state, wherein the static state is used to characterize that the infant is in a static state, and the moving state is used to characterize that the infant is in a moving state; or, determining that the first detection result satisfies the first condition when the difference between the number of first target images and the number of second target images reaches a second preset threshold.

[0160] In one possible design, the determining module 901 is further configured to: determine m target images in the video stream, wherein the m target images are images included in a preset second sliding window; and perform a second detection on the state of the infant in the m target images according to a trained detector.

[0161] In one possible design, the determination module 901 is further configured to: detect the human body position of the infant in the m-frame target image based on the trained detector, and obtain m second detection boxes, one second detection box corresponding to one frame of target image; determine the state of the infant in the j-frame target image based on the second detection box corresponding to the j-th frame target image and the second detection box corresponding to the (j-1)-th frame target image, where m≥j≥2.

[0162] In one possible design, the determining module 901 is further configured to: determine the state of the infant in the target image j based on the IoU value between the second detection box corresponding to the target image j and the second detection box corresponding to the target image j-1.

[0163] In one possible design, determining that the second detection result satisfies the second condition includes: when the ratio between the number of third target images and the number of fourth target images reaches a third preset threshold, determining that the second detection result satisfies the second condition, wherein the third target image and the fourth target image belong to the m-frame target images, the third target image is a target image in which the infant is in a static state, and the fourth target image is a target image in which the infant is in a moving state, wherein the static state is used to characterize that the infant is in a static state, and the moving state is used to characterize that the infant is in a moving state; or, when the difference between the number of third target images and the number of fourth target images reaches a fourth preset threshold, determining that the second detection result satisfies the second condition.

[0164] In one possible design, the determining module 901 is further configured to: identify the current sleeping position of the infant in the current frame target image according to the trained classifier, wherein the current frame target image belongs to k frame target images, the k frame target images are images included in a preset expandable window, and the state of the infant in the k frame target images is the sleeping state.

[0165] In one possible design, the determining module 901 is further configured to: after determining that the infant's state is a sleep state, continue to perform the second detection on the infant's state according to the video stream to obtain a third detection result, and determine that the infant's state is the sleep detection state if the third detection result does not meet the second condition.

[0166] In one possible design, the determining module 901 is further configured to: after determining that the infant's state is the sleep detection state, continue to perform the first detection on the infant's state according to the video stream to obtain the fourth detection result, and if it is determined that the fourth detection result does not meet the first condition, determine that the infant's state is the motion state, which is used to characterize that the infant is in a motion state.

[0167] In one possible design, where the infant posture includes the infant lying position, the current posture includes the current lying position, and the safe posture includes at least one of the following: supine, left-side lying, and right-side lying.

[0168] In one possible design, the reminder module 903 is further configured to: output a second target instruction when it is determined that the current posture is a dangerous posture and the duration of the dangerous posture reaches a second preset duration, the second target instruction being used to remind the infant that there is a safety risk.

[0169] In one possible design, the dangerous posture includes: lying face down, or prone.

[0170] In one possible design, the statistics module 902 is also used to: display the duration of the current posture on the target device after calculating the duration.

[0171] In one possible design, the statistics module 902 is further configured to: after calculating the duration of the current posture, generate a posture adjustment strategy based on the historical posture statistics of the infant in the historical video stream and the duration of the current posture, the posture adjustment strategy being included in the first target instruction.

[0172] It should be noted that the information interaction and execution process between the modules / units in the computer device 900 are based on the same concept as the method embodiments described above in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0173] Next, we will introduce another computer device provided in the embodiments of this application. Please refer to [link / reference]. Figure 10 , Figure 10 This is a schematic diagram of a computer device provided in an embodiment of this application. The computer device 1000 may be equipped with... Figure 9 The computer device 900 described in the corresponding embodiment is used to implement Figure 9 Corresponding to the functionality of the computer device 900 in the embodiment, specifically, the computer device 1000 is implemented by one or more servers. The computer device 1000 can vary significantly due to differences in configuration or performance, and may include one or more central processing units (CPUs) 1022 and memory 1032, and one or more storage media 1030 (e.g., one or more mass storage devices) for storing application programs 1042 or data 1044. The memory 1032 and storage media 1030 can be temporary or persistent storage. The program stored in the storage media 1030 may include one or more modules (not shown in the figure), each module including a series of instruction operations on the computer device 1000. Furthermore, the CPU 1022 may be configured to communicate with the storage media 1030 and execute the series of instruction operations in the storage media 1030 on the computer device 1000.

[0174] Computer device 1000 may also include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0175] In this embodiment, the central processing unit 1022 is used to execute... Figure 3The steps in the corresponding embodiment are as follows. For example, the central processing unit 1022 can be used to: first, determine the infant's current posture based on the video stream captured in real time by the shooting device, and if the current posture is determined to be a safe posture (e.g., lying on one's back, lying on one's left side, lying on one's right side, lying on one's back, lying on one's left side, lying on one's right side), count the duration of the current posture, and when the duration reaches a first preset duration, it indicates that the infant has maintained the current posture for too long. At this time, a first target instruction is output to remind the guardian to adjust the infant's posture, so that the guardian can clearly know the duration of the various postures of the infant, which is convenient for the guardian to scientifically and reasonably adjust the infant's posture, so that the duration of various postures is as reasonable and balanced as possible, and to achieve an ideal head shape.

[0176] It should be noted that the specific manner in which the central processing unit 1022 executes the above steps is different from that described in this application. Figure 3 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in the above embodiments of this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0177] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0178] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0179] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0180] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A method for detecting infant posture, characterized in that, include: The infant's current posture is determined based on a video stream captured by pre-deployed camera equipment; If it is determined that the current posture is a safe posture, the duration of the current posture is recorded. If the duration reaches a first preset duration, a first target instruction is output, which is used to remind the infant to adjust his posture.

2. The method according to claim 1, characterized in that, When the infant posture includes the infant sleeping position, the current posture includes the current sleeping position, and the safe posture includes at least any one of the following: Sleep on your back, sleep on your left side, or sleep on your right side.

3. The method according to claim 2, characterized in that, Determining the infant's current posture based on the video stream includes: The state of the infant is first detected based on the video stream to obtain a first detection result. If the first detection result meets a first condition, the state of the infant is determined to be a sleep detection state. The sleep detection state is used to trigger a second detection of the state of the infant. The state of the infant is detected by the video stream to obtain a second detection result. If the second detection result meets the second condition, the state of the infant is determined to be a sleep state. The sleep state is used to characterize that the infant is in a sleep state. The infant's current sleeping position is determined based on the sleep state.

4. The method according to claim 3, characterized in that, The first detection of the infant's state based on the video stream includes: Determine n target images in the video stream, wherein the n target images are the images included in a preset first sliding window; The state of the infant in the n target images is first detected based on the trained detector.

5. The method according to claim 4, characterized in that, The first detection of the state of the infant in the n frames of target images based on the trained detector includes: The head position of the infant in the n target images is detected by the trained detector to obtain n first detection boxes, with one first detection box corresponding to one target image. The state of the infant in the target image i is determined based on the first detection box corresponding to the target image i and the first detection box corresponding to the target image i-1, where n≥i≥2.

6. The method according to claim 5, characterized in that, Determining the state of the infant in the target image i based on the first detection box corresponding to the target image i-th frame and the first detection box corresponding to the target image i-1-th frame includes: The state of the infant in the target image i is determined based on the intersection-union ratio (IoU) between the first detection box corresponding to the target image i and the first detection box corresponding to the target image i-1.

7. The method according to any one of claims 4-6, characterized in that, The determination that the first detection result meets the first condition includes: When the ratio between the number of first target images and the number of second target images reaches a first preset threshold, it is determined that the first detection result satisfies the first condition, the first target image and the second target image belong to the n-frame target images, the first target image is a target image in which the state of the infant is static, and the second target image is a target image in which the state of the infant is dynamic. The static state is used to characterize that the infant is in a static state, and the dynamic state is used to characterize that the infant is in a dynamic state. or, If the difference between the number of the first target images and the number of the second target images reaches a second preset threshold, the first detection result is determined to satisfy the first condition.

8. The method according to any one of claims 3-7, characterized in that, The second detection of the infant's state based on the video stream includes: Determine m target images in the video stream, wherein the m target images are images included in a preset second sliding window; A second detection is performed on the state of the infant in the m-frame target images based on the trained detector.

9. The method according to claim 8, characterized in that, The second detection of the infant's state in the m-frame target images based on the trained detector includes: The human body position of the infant in the m target images is detected by the trained detector to obtain m second detection boxes, with one second detection box corresponding to one target image. The state of the infant in the target image j is determined based on the second detection box corresponding to the target image j and the second detection box corresponding to the target image j-1, where m≥j≥2.

10. The method according to claim 9, characterized in that, The step of determining the state of the infant in the target image j based on the second detection box corresponding to the target image j-1 and the second detection box corresponding to the target image j-1 includes: The state of the infant in the target image j is determined based on the IoU value between the second detection box corresponding to the target image j and the second detection box corresponding to the target image j-1.

11. The method according to any one of claims 8-10, characterized in that, The determination that the second detection result meets the second condition includes: When the ratio between the number of third target images and the number of fourth target images reaches a third preset threshold, the second detection result is determined to satisfy the second condition. The third target image and the fourth target image belong to the m-frame target images. The third target image is a target image in which the state of the infant is static, and the fourth target image is a target image in which the state of the infant is moving. The static state is used to characterize that the infant is in a static state, and the moving state is used to characterize that the infant is in a moving state. or, If the difference between the number of the third target images and the number of the fourth target images reaches a fourth preset threshold, the second detection result is determined to satisfy the second condition.

12. The method according to any one of claims 3-11, characterized in that, Determining the infant's current sleeping position based on the sleep state includes: The current sleeping position of the infant in the current frame target image is identified according to the trained classifier. The current frame target image belongs to k frame target images, which are images included in a preset expandable window. The state of the infant in the k frame target images is the sleeping state.

13. The method according to any one of claims 3-12, characterized in that, After determining that the infant is in a sleeping state, the method further includes: The second detection is performed on the infant's state based on the video stream to obtain a third detection result. If the third detection result does not meet the second condition, the infant's state is determined to be the sleep detection state.

14. The method according to claim 13, characterized in that, After determining that the infant's state is the sleep detection state, the method further includes: The first detection is performed on the state of the infant based on the video stream to obtain a fourth detection result. If the fourth detection result does not meet the first condition, the state of the infant is determined to be in motion. The motion state is used to characterize that the infant is in motion.

15. The method according to any one of claims 1-14, characterized in that, When the infant posture includes a lying position, the current posture includes a current lying position, and the safe posture includes at least one of the following: Lying on your back, lying on your left side, lying on your right side.

16. The method according to any one of claims 1-15, characterized in that, After determining the infant's current posture based on the video stream, the method further includes: If it is determined that the current posture is a dangerous posture and the duration of the dangerous posture reaches a second preset duration, a second target instruction is output. The second target instruction is used to remind the infant that there is a safety risk.

17. The method according to claim 16, characterized in that, The dangerous postures include: Sleeping on one's stomach, or lying prone.

18. The method according to any one of claims 1-17, characterized in that, After calculating the duration of the current pose, the method further includes: The duration is displayed on the target device.

19. The method according to any one of claims 1-18, characterized in that, After calculating the duration of the current pose, the method further includes: Based on the historical posture statistics of infants and toddlers in the historical video stream and the duration of the current posture, a posture adjustment strategy is generated, which is included in the first target instruction.

20. A computer device, characterized in that, include: A determination module is used to determine the current posture of an infant based on a video stream captured by a pre-deployed camera device; The statistics module is used to calculate the duration of the current posture when it is determined that the current posture is a safe posture. The reminder module is used to output a first target instruction when the duration reaches a first preset duration. The first target instruction is used to remind the infant to adjust his posture.

21. The device according to claim 20, characterized in that, When the infant posture includes the infant sleeping position, the current posture includes the current sleeping position, and the safe posture includes at least any one of the following: Sleep on your back, sleep on your left side, or sleep on your right side.

22. The device according to claim 21, characterized in that, The determining module is specifically used for: The state of the infant is first detected based on the video stream to obtain a first detection result. If the first detection result meets a first condition, the state of the infant is determined to be a sleep detection state. The sleep detection state is used to trigger a second detection of the state of the infant. The state of the infant is detected by the video stream to obtain a second detection result. If the second detection result meets the second condition, the state of the infant is determined to be a sleep state. The sleep state is used to characterize that the infant is in a sleep state. The infant's current sleeping position is determined based on the sleep state.

23. The device according to claim 22, characterized in that, The determining module is further configured to: Determine n target images in the video stream, wherein the n target images are the images included in a preset first sliding window; The state of the infant in the n target images is first detected based on the trained detector.

24. The device according to claim 23, characterized in that, The determining module is further configured to: The head position of the infant in the n target images is detected by the trained detector to obtain n first detection boxes, with one first detection box corresponding to one target image. The state of the infant in the target image i is determined based on the first detection box corresponding to the target image i and the first detection box corresponding to the target image i-1, where n≥i≥2.

25. The device according to claim 24, characterized in that, The determining module is further configured to: The state of the infant in the target image i is determined based on the IoU value between the first detection box corresponding to the target image i and the first detection box corresponding to the target image i-1.

26. The device according to any one of claims 23-25, characterized in that, The determination that the first detection result meets the first condition includes: When the ratio between the number of first target images and the number of second target images reaches a first preset threshold, it is determined that the first detection result satisfies the first condition, the first target image and the second target image belong to the n-frame target images, the first target image is a target image in which the state of the infant is static, and the second target image is a target image in which the state of the infant is dynamic. The static state is used to characterize that the infant is in a static state, and the dynamic state is used to characterize that the infant is in a dynamic state. or, If the difference between the number of the first target images and the number of the second target images reaches a second preset threshold, the first detection result is determined to satisfy the first condition.

27. The device according to any one of claims 22-26, characterized in that, The determining module is further configured to: Determine m target images in the video stream, wherein the m target images are images included in a preset second sliding window; A second detection is performed on the state of the infant in the m-frame target images based on the trained detector.

28. The device according to claim 27, characterized in that, The determining module is further configured to: The human body position of the infant in the m target images is detected by the trained detector to obtain m second detection boxes, with one second detection box corresponding to one target image. The state of the infant in the target image j is determined based on the second detection box corresponding to the target image j and the second detection box corresponding to the target image j-1, where m≥j≥2.

29. The device according to claim 28, characterized in that, The determining module is further configured to: The state of the infant in the target image j is determined based on the IoU value between the second detection box corresponding to the target image j and the second detection box corresponding to the target image j-1.

30. The device according to any one of claims 27-29, characterized in that, The determination that the second detection result meets the second condition includes: When the ratio between the number of third target images and the number of fourth target images reaches a third preset threshold, the second detection result is determined to satisfy the second condition. The third target image and the fourth target image belong to the m-frame target images. The third target image is a target image in which the state of the infant is static, and the fourth target image is a target image in which the state of the infant is moving. The static state is used to characterize that the infant is in a static state, and the moving state is used to characterize that the infant is in a moving state. or, If the difference between the number of the third target images and the number of the fourth target images reaches a fourth preset threshold, the second detection result is determined to satisfy the second condition.

31. The device according to any one of claims 22-30, characterized in that, The determining module is further configured to: The current sleeping position of the infant in the current frame target image is identified according to the trained classifier. The current frame target image belongs to k frame target images, which are images included in a preset expandable window. The state of the infant in the k frame target images is the sleeping state.

32. The device according to any one of claims 22-31, characterized in that, The determining module is further configured to: After determining that the infant is in a sleep state, the second detection is performed on the infant's state based on the video stream to obtain a third detection result. If the third detection result does not meet the second condition, the infant's state is determined to be the sleep detection state.

33. The device according to claim 32, characterized in that, The determining module is further configured to: After determining that the infant's state is the sleep detection state, the first detection is performed on the infant's state according to the video stream to obtain a fourth detection result. If the fourth detection result does not meet the first condition, the infant's state is determined to be a motion state, which is used to characterize that the infant is in a motion state.

34. The device according to any one of claims 20-33, characterized in that, When the infant posture includes a lying position, the current posture includes a current lying position, and the safe posture includes at least one of the following: Lying on your back, lying on your left side, lying on your right side.

35. The device according to any one of claims 20-34, characterized in that, The reminder module is also used for: If it is determined that the current posture is a dangerous posture and the duration of the dangerous posture reaches a second preset duration, a second target instruction is output. The second target instruction is used to remind the infant that there is a safety risk.

36. The device according to claim 35, characterized in that, The dangerous postures include: Sleeping on one's stomach, or lying prone.

37. The device according to any one of claims 20-36, characterized in that, The statistics module is also used for: After calculating the duration of the current posture, the duration is displayed on the target device.

38. The device according to any one of claims 20-37, characterized in that, The statistics module is also used for: After calculating the duration of the current posture, a posture adjustment strategy is generated based on the historical posture statistics of the infant in the historical video stream and the duration of the current posture. The posture adjustment strategy is included in the first target instruction.

39. A computer device comprising a processor and a memory, the processor being coupled to the memory, characterized in that, The memory is used to store programs; The processor is configured to execute a program in the memory, causing the computer device to perform the method as described in any one of claims 1-19.

40. A computer storage medium, characterized in that, The device stores computer-readable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-19.

41. A computer program product, characterized in that, The computer program product includes computer-readable instructions that, when executed by a processor, implement the method as described in any one of claims 1-19.

42. A chip, the chip comprising a processor and a data interface, characterized in that, The processor reads instructions stored in the memory through the data interface and executes the method as described in any one of claims 1-19.