A method for intelligently judging human body postures and a nursing device
By obtaining real-life angle video frames in the care equipment and extracting human body data information, and using machine learning models to calculate the human body posture probability, the problems of low accuracy and high misjudgment rate in the prior art are solved, and higher recognition accuracy and reduced misjudgment are achieved.
Patent Information
- Application Number
- CN202011080377.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-10-10
AI Technical Summary
When judging human postures, existing care equipment is difficult to process images at non-horizontal angles, and it is easy to misjudgment non-human objects, resulting in low judgment accuracy.
By obtaining video frames from the real-life angle of the camera device, extracting the human body's key point data information, coordinate data information and enclosing frame data information, using machine learning training model to calculate the human body's posture probability, and eliminating background misjudgment.
It improves the accuracy of identification and judgment of care equipment, reduces misjudgment in background scenes, and enhances the accurate recognition of human posture.
Smart Images

Figure CN112132110B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nursing devices, and in particular, to a method for intelligently judging human postures and a nursing device. Background Art
[0002] Currently, most common nursing devices on the market use machine learning technology to judge human postures for intelligent nursing. However, it is found in practice that since most of the machine learning posture training libraries used in existing nursing devices are based on images taken at horizontal angles, when the nursing device obtains images taken at non-horizontal angles, they do not match the training data of general machine learning posture training libraries, making it difficult to ensure the judgment accuracy rate of the nursing device; moreover, most existing nursing devices will judge many non-human objects in the scene as humans, especially some human-shaped objects such as calendars and mannequins, resulting in misjudgments and thus making it difficult to ensure the judgment accuracy rate of the nursing device. Summary of the Invention
[0003] Embodiments of the present invention disclose a method for intelligently judging human postures and a nursing device, which can effectively improve the recognition and judgment accuracy rate of the nursing device while eliminating most misjudgments in the background scene.
[0004] In a first aspect of the embodiments of the present invention, a method for intelligently judging human postures is disclosed, and the method includes:
[0005] Obtain human data information in the video frame according to the video frame collected at the actual scene angle of the imaging device; wherein, the human data information at least includes human key point data information, human coordinate data information, and human key point bounding box data information;
[0006] Calculate the corresponding human posture probability in the video frame from the machine learning training model according to the human key point data information, the human coordinate data information, and the human key point bounding box data information;
[0007] Judge the maximum probability posture in the human posture probability as the human posture in the video frame.
[0008] As another optional implementation manner, in the first aspect of the embodiments of the present invention, the calculating the corresponding human posture probability in the video frame from the machine learning training model according to the human key point data information, the human coordinate data information, and the human key point bounding box data information includes:
[0009] Obtain N key point information from the human body key point data information, the human body coordinate data information, and the human body key point bounding box data information to obtain N-dimensional data information; wherein, the N key point information at least includes the coordinate positions of N key points.
[0010] Calculate the distances between each of the N key point coordinate positions to obtain N(N - 1) / 2-dimensional data information.
[0011] Calculate the central point coordinates of the set of the N key point coordinate positions to obtain two-dimensional data information.
[0012] According to the N-dimensional data information, the N(N - 1) / 2-dimensional data information, and the two-dimensional data information, calculate K-dimensional judgment data information; wherein, the K-dimensional judgment data information is obtained by adding the N-dimensional data information, the N(N - 1) / 2-dimensional data information, and the two-dimensional data information.
[0013] Perform a dot product operation on the K-dimensional judgment data information and the machine learning training model to obtain the probability of the human body posture corresponding to the video frame.
[0014] As another alternative implementation, in the first aspect of the embodiments of the present invention, before obtaining the human body data information in the video frame according to the actual scene angle of the imaging device, the method further includes:
[0015] Obtain the background information in the video frame according to the background maintenance algorithm.
[0016] Regard the objects in the video frame other than the background information as the photographed human body to perform the step of obtaining the human body data information in the video frame according to the actual scene angle of the imaging device.
[0017] As another alternative implementation, in the first aspect of the embodiments of the present invention, after regarding the objects in the video frame other than the background information as the photographed human body and before obtaining the human body data information in the video frame according to the actual scene angle of the imaging device, the method further includes:
[0018] If the area of the minimum rectangular bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea, determine it as a mis-identification of the human body image to obtain the key point coordinates of the mis-identified human body image.
[0019] Regard the N key point coordinate positions other than the key point coordinates of the mis-identified human body image as the target key point coordinate positions to perform the step of obtaining the human body data information in the video frame according to the actual scene angle of the imaging device.
[0020] As another alternative implementation, in the first aspect of the embodiments of the present invention, after using the objects other than the background information in the video frame as the photographed human body and before determining that the human body image is mis-recognized to obtain the coordinate positions of the key points of the mis-recognized human body image, the method further includes:
[0021] Calculating the horizontal offset dx and the vertical offset dy between the center point coordinates and the horizontal coordinate of the midpoint of the lower edge of the image in the video frame;
[0022] According to the horizontal offset dx and the vertical offset dy, calculating the horizontal direction decay value decayX and the vertical direction decay value decayY of the center point coordinates;
[0023] According to the horizontal direction decay value decayX and the vertical direction decay value decayY, calculating the minimum bounding area judgment threshold minArea;
[0024] Judging whether the area of the minimum rectangle bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea.
[0025] As another alternative implementation, in the first aspect of the embodiments of the present invention, the method further includes:
[0026] Obtaining the machine learning training model established by the terminal device.
[0027] A second aspect of the embodiments of the present invention discloses a nursing device, and the nursing device includes:
[0028] A first acquisition unit, configured to acquire human body data information in the video frame according to a video frame collected at the actual scene angle of the imaging device; wherein, the human body data information at least includes human body key point data information, human body coordinate data information, and human body key point bounding box data information;
[0029] A first calculation unit, configured to calculate the corresponding human body posture probability in the video frame from the machine learning training model according to the human body key point data information, the human body coordinate data information, and the human body key point bounding box data information;
[0030] A first determination unit, configured to determine the maximum probability posture in the human body posture probability as the human body posture in the video frame.
[0031] As an alternative implementation, in the second aspect of the embodiments of the present invention, the first calculation unit further includes:
[0032] An obtaining subunit, configured to obtain N key point information from the human key point data information, the human coordinate data information, and the human key point bounding box data information, so as to obtain N-dimensional data information; wherein, the N key point information at least includes N key point coordinate positions;
[0033] A first calculation subunit, configured to calculate the distances between every two of the N key point coordinate positions, so as to obtain N(N-1) / 2-dimensional data information;
[0034] A second calculation subunit, configured to calculate the center point coordinates of the set of the N key point coordinate positions, so as to obtain two-dimensional data information;
[0035] A third calculation subunit, configured to calculate K-dimensional judgment data information according to the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information; wherein, the K-dimensional judgment data information is obtained by adding the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information;
[0036] A fourth calculation subunit, configured to perform a dot product operation on the K-dimensional judgment data information and the machine learning training model, so as to obtain the human body posture probability corresponding to the video frame.
[0037] A third aspect of the embodiments of the present invention discloses a nursing device, where the nursing device includes:
[0038] A memory storing executable program code;
[0039] A processor coupled to the memory;
[0040] The processor calls the executable program code stored in the memory and executes a method for intelligently judging a human body posture disclosed in the first aspect of the embodiments of the present invention.
[0041] A fourth aspect of the embodiments of the present invention discloses a computer-readable storage medium, which stores a computer program, wherein the computer program enables a computer to execute a method for intelligently judging a human body posture disclosed in the first aspect of the embodiments of the present invention.
[0042] A fifth aspect of the embodiments of the present invention discloses a computer program product, when the computer program product runs on a computer, enabling the computer to execute some or all of the steps of any one of the methods for intelligently judging a human body posture in the first aspect.
[0043] A sixth aspect of the embodiments of the present invention discloses an application publishing platform for publishing computer program products. When the computer program product runs on a computer, the computer is caused to execute some or all of the steps of any one of the intelligent human pose judgment methods in the first aspect.
[0044] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0045] In the embodiments of the present invention, after the nursing device obtains the human body data information in the video frame according to the video frame collected from the actual scene angle of the imaging device, the nursing device can calculate the corresponding human body pose probability in the video frame from the machine learning training model according to the above human body data information, and determine the pose with the maximum probability in the human body pose probability as the human body pose in the video frame; wherein, the human body data information at least includes human key point data information, human coordinate data information, and human key point bounding box data information. It can be seen that the embodiments of the present invention can effectively improve the recognition and judgment accuracy of the nursing device while eliminating most misjudgments in the background scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 is a flowchart of a method for intelligently judging human body poses disclosed in the embodiments of the present invention;
[0048] Figure 2 is a flowchart of another method for intelligently judging human body poses disclosed in the embodiments of the present invention;
[0049] Figure 3 is a structural diagram of a nursing device disclosed in the embodiments of the present invention;
[0050] Figure 4 is a structural diagram of another nursing device disclosed in the embodiments of the present invention;
[0051] Figure 5 is a structural diagram of another nursing device disclosed in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] It should be noted that the terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "having" in the embodiments of the present invention and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0054] The embodiments of the present invention disclose a method for intelligently judging human postures and a nursing device, which can effectively improve the recognition and judgment accuracy rate of the nursing device while excluding most misjudgments in the background scene. The following is a detailed description in conjunction with the accompanying drawings.
[0055] Embodiment 1
[0056] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for intelligently judging human postures disclosed in the embodiments of the present invention. As Figure 1 shown, the method for intelligently judging human postures may include the following steps.
[0057] 101. The nursing device obtains the human data information in the video frame according to the video frame collected at the actual scene angle of the imaging device; wherein, the human data information at least includes human key point data information, human coordinate data information, and human key point bounding box data information.
[0058] In the embodiments of the present invention, the nursing device may be an electronic device such as a wearable watch, a tablet computer, a mobile phone, a home nursing device, or a monitoring and nursing device for the elderly, children, students, the general public, and families, which is not limited in the embodiments of the present invention.
[0059] As an optional implementation manner, the nursing device may set a certain time interval for shooting, and perform human body recognition on the video frames collected at the actual scene angle within this time interval. If the nursing device fails to recognize the human body in the video frames shot within consecutive time intervals, the nursing device may temporarily not execute step 101 until the nursing device recognizes the human body in the video frames shot within consecutive time intervals, and then the nursing device may execute step 101 again;
[0060] In addition, when the probability that the nursing device fails to recognize a human body in video frames captured within a continuous time interval is higher than a specified threshold, the nursing device may temporarily not execute step 101. Until the probability that the nursing device recognizes a human body in video frames captured within a continuous time interval is higher than the specified threshold, the nursing device may execute step 101 again.
[0061] 102. The nursing device calculates the corresponding human body pose probability in the video frame from the machine learning training model according to the human body key point data information, the human body coordinate data information, and the human body key point bounding box data information.
[0062] In this embodiment, the machine learning training model is modeled by a terminal device (such as a PC terminal device), and then the modeled machine learning training model can be imported and stored in the nursing device. When the nursing device obtains the human body data information in the video frame, it can directly obtain the imported and stored machine learning training model;
[0063] In addition, for the machine learning training model, the terminal device (such as a PC terminal device) can first obtain various human body pose sample images of each region taken from the actual scene angle sent by the nursing device, and use these as training sample images; among them, the various human body pose sample images of each region taken from the actual scene angle at least include images of different distance positions between the photographed human body and the camera device;
[0064] In addition, after the terminal device (such as a PC terminal device) obtains the training sample images sent by the nursing device, it can obtain Y key point information from the training sample images through human body model modeling software such as Pose software (Openpose, Posenet) to obtain Y-dimensional data information; among them, the Y key point information at least includes the coordinate positions of Y key points;
[0065] In addition, after the terminal device (such as a PC terminal device) obtains the Y-dimensional data information, it can calculate the distances between each of the Y key point coordinate positions to obtain Y(Y - 1) / 2-dimensional data information, and calculate the center point coordinates of the set of Y key point coordinate positions by the mean center method to obtain the central two-dimensional coordinate data information D;
[0066] In addition, according to the Y-dimensional data information, the Y(Y - 1) / 2-dimensional data information, and the central two-dimensional coordinate data information D, the terminal device (such as a PC terminal device) calculates the training data Z; where the training data Z is the sum of the Y-dimensional data information, the Y(Y - 1) / 2-dimensional data information, and the central two-dimensional coordinate data information D, that is, Z = Y + Y(Y - 1) / 2 + D;
[0067] Also, after the terminal device (such as a PC terminal device) uses Y+Y(Y-1) / 2+D as the training data Z, the human poses in the training sample images can be used as the labeled categories, and a multi-layer neural network model + Softmax classifier can be established for training to obtain a machine learning training model.
[0068] In this embodiment, after the care device obtains the human data information of the video frame to be recognized, the care device can obtain N key point information from the training sample images through human model modeling software such as Pose software (Openpose, Posenet) to obtain N-dimensional data information; among them, the N key point information at least includes the coordinate positions of N key points.
[0069] Also, calculate the distances between each of the N key point coordinate positions to obtain N(N-1) / 2-dimensional data information, and calculate the center point coordinates of the set of N key point coordinate positions by the mean center method to obtain two-dimensional data information R.
[0070] Also, according to the N-dimensional data information, N(N-1) / 2-dimensional data information, and two-dimensional data information R, the care device calculates K-dimensional judgment data information; where the K-dimensional judgment data information is the sum of the N-dimensional data information, N(N-1) / 2-dimensional data information, and two-dimensional data information R, that is, K = N + N(N-1) / 2 + R.
[0071] Also, perform a dot product operation on the K-dimensional judgment data information and the machine learning training model to calculate the corresponding human pose probability in the video frame.
[0072] 103. The care device determines the pose with the maximum probability in the human pose probability as the human pose in the video frame.
[0073] In this embodiment, the care device in the present application can immediately judge the living rules and abnormalities of the human body through the poses of the human body in different scenarios. For example, lying on the ground is regarded as a dangerous state, or lying in bed, and the sleep time will also be counted by timing. When it times out, it is regarded as abnormal, and when it is recognized as a sitting posture and the time is too long, it is regarded as abnormal.
[0074] In this embodiment, for the same portrait picture, when it is in different positions in the picture, there will be different distortions. Correspondingly, different distances will also correspond to different regions of the image. For example, those at close range are generally located in the lower area of the lens. By adding the coordinate data of the human body in the image to the dimension of the machine learning training data, through experiments, the accuracy of the human pose image recognition of the care device can be effectively improved greatly.
[0075] In Figure 1 In the method for intelligently judging human poses shown below, the care device is taken as an example of the execution subject for description. It should be noted thatFigure 1 The execution subject of the method for intelligently judging human body postures shown may also be an independent device associated with the nursing device, which is not limited in the embodiments of the present invention.
[0076] It can be seen that implementing Figure 1 the described method for intelligently judging human body postures can effectively improve the recognition and judgment accuracy rate of the nursing device while eliminating misjudgments in most background scenes.
[0077] In addition, implementing Figure 1 the described method for intelligently judging human body postures can effectively and greatly improve the recognition accuracy rate of human body posture images of the nursing device by adding the coordinate data of the human body in the image to the dimension of the machine learning training data.
[0078] Embodiment 2
[0079] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another method for intelligently judging human body postures disclosed in the embodiments of the present invention. As Figure 2 shown, the method for intelligently judging human body postures may include the following steps:
[0080] 201. The nursing device obtains the machine learning training model established by the terminal device.
[0081] In this embodiment, the machine learning training model is modeled by the terminal device (such as a PC terminal device), and then the modeled machine learning training model can be imported and stored in the nursing device. When the nursing device obtains the human body data information in the video frame, it can directly obtain the imported and stored machine learning training model;
[0082] Moreover, for the machine learning training model, the terminal device (such as a PC terminal device) can first obtain various human body posture sample images of each area taken from the actual scene angle sent by the nursing device, and use them as training sample images; among them, the various human body posture sample images of each area taken from the actual scene angle at least include the images of different distance positions between the photographed human body and the camera device;
[0083] Moreover, after the terminal device (such as a PC terminal device) obtains the training sample images sent by the nursing device, it can obtain Y key point information from the training sample images through human body model modeling software such as Pose software (Openpose, Posenet) to obtain Y-dimensional data information; among them, the Y key point information at least includes the coordinate positions of Y key points;
[0084] Moreover, after the Y - dimensional data information is obtained by the terminal device (such as a PC terminal device), the distances between the coordinate positions of every Y key points can be calculated to obtain Y(Y - 1) / 2 - dimensional data information, and the central point coordinates of the set of Y key point coordinate positions can be calculated by the mean - center method to obtain the central two - dimensional coordinate data information D;
[0085] Moreover, according to the Y - dimensional data information, Y(Y - 1) / 2 - dimensional data information, and the central two - dimensional coordinate data information D, the terminal device (such as a PC terminal device) calculates the training data Z; where the training data Z is the sum of the Y - dimensional data information, Y(Y - 1) / 2 - dimensional data information, and the central two - dimensional coordinate data information D, that is, Z = Y+Y(Y - 1) / 2 + D;
[0086] Moreover, after the terminal device (such as a PC terminal device) uses Y+Y(Y - 1) / 2 + D as the training data Z, the human body posture in the training sample image can be used as the marked category, and a multi - layer neural network model + Softmax classifier can be established for training to obtain a machine - learning training model.
[0087] 202. The care - taking device obtains the background information in the video frame according to the background maintenance algorithm.
[0088] In this embodiment, since there may be mis - identification phenomena in the human body posture recognition and judgment of the care - taking device, in order to reduce the mis - identification probability as much as possible, the present application can use various methods to reduce mis - identification, such as the background maintenance method and the minimum bounding box threshold determination method.
[0089] 203. The care - taking device regards the objects other than the background information in the video frame as the photographed human body.
[0090] In this embodiment, the running process in the care - taking device includes a background maintenance process, that is, the scenes in the video frame that do not move for a long time or move occasionally (such as moving a chair) can be determined as the background, and only the objects moving in front of the background are judged as human bodies and then the posture is judged, excluding most mis - judgments in the background scenes.
[0091] 204. The care - taking device calculates the horizontal offset dx and the vertical offset dy between the central point coordinates and the abscissa of the mid - point of the lower edge of the image in the video frame.
[0092] In this embodiment, the care-taking device can obtain the coordinate positions of N key points from the training sample images through human body model modeling software such as Pose software (Openpose, Posenet), and calculate the center point coordinates (X, Y) of the set of N key point coordinate positions by the mean center method. Then, calculate the absolute value (offset) of the difference between the horizontal and vertical coordinates of the center point coordinates (X, Y) and the midpoint coordinates (X1, Y1) of the lower edge of the image in the video frame, that is, dx = abs(X - X1), dy = abs(Y - Y1).
[0093] 205. The care-taking device calculates the decay value decayX in the horizontal coordinate direction and the decay value decayY in the vertical coordinate direction of the center point coordinates according to the horizontal coordinate offset dx and the vertical coordinate offset dy.
[0094] 206. The care-taking device calculates the minimum bounding area judgment threshold minArea according to the decay value decayX in the horizontal coordinate direction and the decay value decayY in the vertical coordinate direction.
[0095] In this embodiment, the formula for the decay value decayX in the horizontal coordinate direction is decayX = a1 * dx 2 + a2 * dx + a3, and the formula for the decay value decayY in the vertical coordinate direction is decayY = b1 * dy 2 + b2 * dy + b3. The formula for the minimum bounding area judgment threshold minArea is minArea = A + decayY + decayX;
[0096] Among them, a1, a2, a3, b1, b2, b3, and A are empirical constants related to the device and its camera parameters. These empirical constants can be: -0.006, -5, 76, -0.001, -32, 110, 6000 respectively, and this embodiment does not make any limitations.
[0097] 207. The care-taking device determines whether the area of the minimum rectangular bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea. If so, execute steps 208 to 216. If not, end this process.
[0098] 208. If the area of the minimum rectangular bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea, the care-taking device determines that the human body image is mis-identified to obtain the key point coordinate positions of the mis-identified human body image.
[0099] In this embodiment, the care-taking device of the present application is installed indoors at a certain fixed height and angle. Therefore, when the human body is at different positions in the picture, the recognized size will not be lower than a threshold value. Those lower than this value are considered mis-identifications;
[0100] In addition, after the care device determines that the above image is a mis-identified human image, the coordinate positions of the key points recognized on the image can be used as the coordinate positions of the mis-identified key points of the human image, and the coordinate positions of the mis-identified key points of the human image can be removed from the N coordinate positions of the key points used to judge the human posture, thereby further improving the judgment accuracy of the care device for the human posture.
[0101] 209. The care device uses the N coordinate positions of the key points except the coordinate positions of the mis-identified key points of the human image as the target coordinate positions of the key points.
[0102] 210. The care device obtains the human data information in the video frame according to the video frame collected at the actual angle of the camera device; among them, the human data information at least includes human key point data information, human coordinate data information, and human key point bounding box data information.
[0103] 211. The care device obtains N key point information from the human key point data information, human coordinate data information, and human key point bounding box data information to obtain N-dimensional data information; among them, the N key point information at least includes N coordinate positions of the key points.
[0104] 212. The care device calculates the distance between each of the N coordinate positions of the key points to obtain N(N - 1) / 2-dimensional data information.
[0105] 213. The care device calculates the center point coordinates of the set of N coordinate positions of the key points to obtain two-dimensional data information.
[0106] 214. The care device calculates K-dimensional judgment data information according to the N-dimensional data information, N(N - 1) / 2-dimensional data information, and two-dimensional data information; where the K-dimensional judgment data information is the sum of the N-dimensional data information, N(N - 1) / 2-dimensional data information, and two-dimensional data information.
[0107] 215. The care device performs a dot product operation on the K-dimensional judgment data information and the machine learning training model to obtain the corresponding human posture probability in the video frame.
[0108] In this embodiment, after the care device obtains the human data information for the video frame to be recognized, the care device can obtain N key point information from the training sample images through human model modeling software such as Pose software (Openpose, Posenet) to obtain N-dimensional data information; among them, the N key point information at least includes N coordinate positions of the key points;
[0109] Further, calculate the distances between every N key point coordinate positions to obtain N(N - 1) / 2 - dimensional data information, and calculate the central point coordinate of the set of N key point coordinate positions by the mean - center method to obtain two - dimensional data information R;
[0110] Further, according to the N - dimensional data information, N(N - 1) / 2 - dimensional data information, and two - dimensional data information R, the care - taking device calculates K - dimensional judgment data information; where the K - dimensional judgment data information is the sum of the N - dimensional data information, N(N - 1) / 2 - dimensional data information, and two - dimensional data information R, that is, K = N+N(N - 1) / 2 + R;
[0111] Further, perform a dot - product operation on the K - dimensional judgment data information and the machine - learning training model to calculate the corresponding human body pose probability in the video frame.
[0112] 216. The care - taking device determines the pose with the maximum probability in the human body pose probability as the human body pose in the video frame.
[0113] As an alternative implementation, in the embodiments of the present invention, the care - taking device in the present application can immediately judge the living rules and abnormalities of the human body through the poses of the human body in different scenarios. For example, lying on the ground is regarded as a dangerous state, or lying in bed, and the sleep time will also be counted through timing. When it times out, it is regarded as abnormal, and when it is recognized as a sitting posture and the time is too long, it is regarded as abnormal;
[0114] Further, when the care - taking device detects human body pose abnormal information, it can send relevant prompt information. The prompt information sent by the care - taking device includes:
[0115] Send a display information prompt to the mobile device associated with the care - taking device;
[0116] And / or, a voice information prompt sent by the care - taking device.
[0117] It can be seen that implementing Figure 2 the other method for intelligently judging human body poses described can effectively improve the recognition and judgment accuracy of the care - taking device while eliminating most misjudgments in the background scenes.
[0118] In addition, implementing Figure 2 the other method for intelligently judging human body poses described can judge the background in the scene and distinguish the foreground objects in the background at different lights and different times, thereby more accurately reducing misjudgments.
[0119] In addition, implementing Figure 2 the other method for intelligently judging human body poses described can reduce misrecognition by adopting multiple methods, and thus can effectively reduce the misrecognition probability.
[0120] Embodiment III
[0121] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a care device disclosed in an embodiment of the present invention. As Figure 3 shown, the care device 300 may include a first acquisition unit 301, a first calculation unit 302, and a first determination unit 303, where:
[0122] The first acquisition unit 301 is configured to obtain human data information in the video frame according to the video frames collected at the real scene angle of the imaging device; wherein, the human data information at least includes human key point data information, human coordinate data information, and human key point bounding box data information.
[0123] The calculation unit 302 is configured to calculate the corresponding human pose probability in the video frame from the machine learning training model according to the human key point data information, the human coordinate data information, and the human key point bounding box data information.
[0124] The determination unit 303 is configured to determine the pose with the maximum probability in the human pose probability as the human pose in the video frame.
[0125] In an embodiment of the present invention, the care device may be an electronic device such as a wearable watch, a tablet computer, a mobile phone, a home care monitor, or a surveillance care device for use by the elderly, children, students, the general public, and families, which is not limited in the embodiments of the present invention.
[0126] As an optional implementation manner, the care device may be set to take pictures at a certain time interval, and perform human body recognition on the video frames collected at the real scene angle within the time interval. If the care device fails to recognize a human body in the video frames taken within consecutive time intervals, the first acquisition unit 301 may temporarily not execute the step of obtaining human data information in the video frame until the care device recognizes a human body in the video frames taken within consecutive time intervals, and then the first acquisition unit 301 may execute the step of obtaining human data information in the video frame again;
[0127] And, if the probability that the care device fails to recognize a human body in the video frames taken within consecutive time intervals is higher than a specified threshold, the first acquisition unit 301 may temporarily not execute the step of obtaining human data information in the video frame until the probability that the care device recognizes a human body in the video frames taken within consecutive time intervals is higher than the specified threshold, and then the first acquisition unit 301 may execute the step of obtaining human data information in the video frame again.
[0128] In this embodiment, the machine learning training model is modeled by a terminal device (such as a PC terminal device). Subsequently, the modeled machine learning training model can be imported and stored in the care device. When the care device obtains the human data information in the video frame, it can directly obtain the imported and stored machine learning training model;
[0129] Moreover, for the machine learning training model, the terminal device (such as a PC terminal device) can first obtain various human pose sample images of each area taken from the actual scene angle sent by the care device, and use these as training sample images; among them, the various human pose sample images of each area taken from the actual scene angle at least include images of different distance positions between the photographed human body and the camera device;
[0130] Moreover, after the terminal device (such as a PC terminal device) obtains the training sample images sent by the care device, it can obtain Y key point information from the training sample images through human model modeling software such as Pose software (Openpose, Posenet) to obtain Y-dimensional data information; among them, the Y key point information at least includes the coordinate positions of Y key points;
[0131] Moreover, after the terminal device (such as a PC terminal device) obtains the Y-dimensional data information, it can calculate the distance between each of the Y key point coordinate positions to obtain Y(Y - 1) / 2-dimensional data information, and calculate the center point coordinates of the set of Y key point coordinate positions by the mean center method to obtain the central two-dimensional coordinate data information D;
[0132] Moreover, according to the Y-dimensional data information, Y(Y - 1) / 2-dimensional data information, and the central two-dimensional coordinate data information D, the terminal device (such as a PC terminal device) calculates the training data Z; where the training data Z is the sum of the Y-dimensional data information, Y(Y - 1) / 2-dimensional data information, and the central two-dimensional coordinate data information D, that is, Z = Y + Y(Y - 1) / 2 + D;
[0133] Moreover, after the terminal device (such as a PC terminal device) uses Y + Y(Y - 1) / 2 + D as the training data Z, it can use the human pose in the training sample images as the labeled category, and establish a multi-layer neural network model + Softmax classifier for training to obtain the machine learning training model.
[0134] In this embodiment, after the care device obtains the human data information of the video frame to be recognized, the care device can obtain N key point information from the training sample images through human model modeling software such as Pose software (Openpose, Posenet) to obtain N-dimensional data information; among them, the N key point information at least includes the coordinate positions of N key points;
[0135] Moreover, calculate the distances between every N key-point coordinate positions to obtain N(N-1) / 2-dimensional data information, and calculate the central point coordinate of the set of N key-point coordinate positions by the mean center method to obtain two-dimensional data information R;
[0136] Moreover, according to the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information R, the care device calculates K-dimensional judgment data information; where the K-dimensional judgment data information is the sum of the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information R, that is, K = N + N(N-1) / 2 + R;
[0137] Moreover, perform a dot operation on the K-dimensional judgment data information and the machine learning training model to calculate the corresponding human body posture probability in the video frame.
[0138] In this embodiment, the care device in the present application can immediately judge the living rules and abnormalities of the human body through the postures of the human body in different scenarios. For example, lying on the ground is regarded as a dangerous state, or lying in bed, and the sleep time will also be counted by timing. When the time exceeds the limit, it is regarded as abnormal. And when it is recognized as a sitting posture and the time is too long, it is regarded as abnormal.
[0139] In this embodiment, for the same portrait picture, when it is located at different positions in the picture, there will be different distortions. Correspondingly, different distances will also correspond to different regions of the image. For example, in the case of a short distance, it is generally located in the lower region of the lens. By adding the coordinate data of the human body in the image to the dimension of the machine learning training data, through experiments, the recognition accuracy of the human body posture image of the care device can be effectively improved greatly.
[0140] It can be seen that the Figure 3 described care device can effectively improve the recognition and judgment accuracy of the care device while excluding misjudgments in most background scenarios.
[0141] In addition, the Figure 3 described care device can effectively improve the recognition accuracy of the human body posture image of the care device greatly by adding the coordinate data of the human body in the image to the dimension of the machine learning training data.
[0142] Embodiment 4
[0143] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of another care device disclosed in the embodiment of the present invention. Among them, Figure 4 the shown care device is optimized from the Figure 3 shown care device. Compared with the Figure 3 shown care device, Figure 4 the shown first calculation unit 302 may further include:
[0144] Obtain a sub - unit 3021, which is used to obtain N key - point information from human key - point data information, human coordinate data information, and human key - point bounding box data information to obtain N - dimensional data information; where the N key - point information includes at least N key - point coordinate positions.
[0145] A first calculation sub - unit 3022, which is used to calculate the distances between every two of the N key - point coordinate positions to obtain N(N - 1) / 2 - dimensional data information.
[0146] A second calculation sub - unit 3023, which is used to calculate the center - point coordinates of the set of N key - point coordinate positions to obtain two - dimensional data information.
[0147] A third calculation sub - unit 3024, which is used to calculate K - dimensional judgment data information according to the N - dimensional data information, N(N - 1) / 2 - dimensional data information, and two - dimensional data information; where the K - dimensional judgment data information is the sum of the N - dimensional data information, N(N - 1) / 2 - dimensional data information, and two - dimensional data information.
[0148] A fourth calculation sub - unit 3025, which is used to perform a dot - product operation on the K - dimensional judgment data information and a machine - learning training model to obtain the corresponding human - pose probability in the video frame.
[0149] In this embodiment, after the first acquisition unit 301 acquires human - data information for the video frame to be recognized, the acquisition sub - unit 3021 can obtain N key - point information from the training - sample images through human - model modeling software such as Pose software (Openpose, Posenet) to obtain N - dimensional data information; where the N key - point information includes at least N key - point coordinate positions;
[0150] And, the first calculation sub - unit 3022 calculates the distances between every two of the N key - point coordinate positions to obtain N(N - 1) / 2 - dimensional data information, and the second calculation sub - unit 3023 calculates the center - point coordinates of the set of N key - point coordinate positions by the mean - center method to obtain two - dimensional data information R;
[0151] And, the third calculation sub - unit 3024 calculates K - dimensional judgment data information according to the N - dimensional data information, N(N - 1) / 2 - dimensional data information, and two - dimensional data information R. The care - taking device calculates the K - dimensional judgment data information; where the K - dimensional judgment data information is the sum of the N - dimensional data information, N(N - 1) / 2 - dimensional data information, and two - dimensional data information R, that is, K = N+N(N - 1) / 2+R;
[0152] And, the fourth calculation sub - unit 3025 performs a dot - product operation on the K - dimensional judgment data information and the machine - learning training model to calculate the corresponding human - pose probability in the video frame.
[0153] As an alternative implementation, Figure 4 the shown nursing device may further include:
[0154] A second acquisition unit 304, configured to, before the first acquisition unit 301 acquires human body data information in a video frame according to the actual scene angle of the imaging device, acquire background information in the video frame according to a background maintenance algorithm.
[0155] In this embodiment, since there may be misrecognition phenomena in the human body posture recognition and judgment of the nursing device, in order to reduce the misrecognition probability as much as possible, this application may adopt various methods to reduce misrecognition, such as the background maintenance method and the minimum bounding box threshold determination method.
[0156] A second determination unit 305, configured to use an object other than the background information in the video frame as the photographed human body to execute the step of the first acquisition unit 301 acquiring human body data information in the video frame according to the actual scene angle of the imaging device.
[0157] In this embodiment, the running process in the nursing device includes a background maintenance process, that is, a scene in the video frame that does not move for a long time or moves occasionally (such as moving a chair) can be determined as the background, and only an object moving in front of the background is used for human body judgment and then posture judgment, excluding most misjudgments in the background scene.
[0158] As an alternative implementation, Figure 4 the shown nursing device may further include:
[0159] A third determination unit 306, configured to, after the second determination unit 305 uses an object other than the background information in the video frame as the photographed human body and before the first acquisition unit 301 acquires human body data information in the video frame according to the actual scene angle of the imaging device, if the area of the minimum rectangular bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea, determine it as a misrecognition of the human body image to obtain the key point coordinate positions of the misrecognized human body image.
[0160] In this embodiment, the nursing device of this application is installed indoors at a certain fixed height and angle, so when the human body is at different positions in the picture, the recognized size will not be lower than a threshold, and those lower than this value are considered misrecognized.
[0161] Correspondingly, the second determination unit 305 is further configured to use the N key point coordinate positions except the key point coordinate positions of the misrecognized human body image as the target key point coordinate positions to execute the step of the first acquisition unit 301 acquiring human body data information in the video frame according to the actual scene angle of the imaging device.
[0162] As an alternative embodiment, Figure 4 the care-taking device shown may further include:
[0163] A second calculation unit 307, configured to calculate a horizontal offset dx and a vertical offset dy between the center point coordinates and the abscissa of the midpoint coordinate of the lower edge of the image in the video frame after the second determination unit 305 takes the object other than the background information in the video frame as the photographed human body and before the third determination unit 306 determines that the human body image is mis-identified.
[0164] In this embodiment, the care-taking device can obtain the coordinate positions of N key points from the training sample images through human body model modeling software such as Pose software (Openpose, Posenet), calculate the center point coordinates (X, Y) of the set of N key point coordinate positions by the mean center method, and then calculate the absolute value (offset) of the difference between the horizontal and vertical coordinates of the center point coordinates (X, Y) and the midpoint coordinates (X1, Y1) of the lower edge of the image in the video frame, that is, dx = abs(X - X1), dy = abs(Y - Y1).
[0165] Correspondingly, the second calculation unit 307 is further configured to calculate a horizontal decay value decayX and a vertical decay value decayY of the center point coordinates according to the horizontal offset dx and the vertical offset dy.
[0166] Correspondingly, the second calculation unit 307 is further configured to calculate a minimum bounding area judgment threshold minArea according to the horizontal decay value decayX and the vertical decay value decayY.
[0167] In this embodiment, the formula for the horizontal decay value decayX is decayX = a1 * dx 2 + a2 * dx + a3, and the formula for the vertical decay value decayY is decayY = b1 * dy 2 + b2 * dy + b3, and the formula for the minimum bounding area judgment threshold minArea is minArea = A + decayY + decayX;
[0168] where a1, a2, a3, b1, b2, b3, and A are empirical constants related to the device and its camera parameters. These empirical constants can be: -0.006, -5, 76, -0.001, -32, 110, 6000 respectively, and this embodiment does not make any limitations.
[0169] A judgment unit 308, configured to judge whether the area of the minimum rectangular bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea.
[0170] As an alternative implementation, Figure 4 the illustrated care device may further include:
[0171] a third acquisition unit 309, configured to acquire a machine learning training model established by a terminal device.
[0172] In this embodiment, the machine learning training model is modeled by a terminal device (such as a PC terminal device), and then the modeled machine learning training model can be imported and stored in the care device. When the care device acquires the human data information in the video frame, it can directly acquire the imported and stored machine learning training model;
[0173] Moreover, for the machine learning training model, the terminal device (such as a PC terminal device) can first acquire various human pose sample images of each area taken from the actual scene angle sent by the care device, and use these as training sample images; among them, the various human pose sample images of each area taken from the actual scene angle at least include images of different distance positions between the photographed human body and the camera device;
[0174] Moreover, after the terminal device (such as a PC terminal device) acquires the training sample images sent by the care device, it can obtain Y key point information from the training sample images through human model modeling software such as Pose software (Openpose, Posenet) to obtain Y-dimensional data information; among them, the Y key point information at least includes the coordinate positions of Y key points;
[0175] Moreover, after the terminal device (such as a PC terminal device) obtains the Y-dimensional data information, it can calculate the distances between each of the Y key point coordinate positions to obtain Y(Y-1) / 2-dimensional data information, and calculate the center point coordinates of the set of Y key point coordinate positions by the mean center method to obtain the central two-dimensional coordinate data information D;
[0176] Moreover, according to the Y-dimensional data information, Y(Y-1) / 2-dimensional data information, and the central two-dimensional coordinate data information D, the terminal device (such as a PC terminal device) calculates the training data Z; where the training data Z is the sum of the Y-dimensional data information, Y(Y-1) / 2-dimensional data information, and the central two-dimensional coordinate data information D, that is, Z = Y + Y(Y-1) / 2 + D;
[0177] Moreover, after the terminal device (such as a PC terminal device) uses Y + Y(Y-1) / 2 + D as the training data Z, it can use the human poses in the training sample images as the marked categories, and establish a multi-layer neural network model + Softmax classifier for training to obtain the machine learning training model.
[0178] As an alternative embodiment, in the embodiments of the present invention, the care device in the present application can immediately judge the living rules and abnormalities of the human body through the postures of the human body in different scenarios. For example, lying on the ground is regarded as a dangerous state, or lying in bed, the sleep time will also be counted through timing. When the time exceeds the limit, it is regarded as abnormal. And when it is recognized as a sitting posture, if the time is too long, it is regarded as abnormal;
[0179] Moreover, when the care device detects abnormal information about the human body posture, it can issue relevant prompt information. The prompt information issued by the care device includes:
[0180] Sending a display information prompt to a mobile device associated with the care device;
[0181] And / or, a voice information prompt issued by the care device.
[0182] It can be seen that the described care device can effectively improve the recognition and judgment accuracy of the care device while eliminating misjudgments in most background scenarios. Figure 4 Moreover, the described care device can judge the background in the scene and distinguish the foreground objects in the background at different times and under different lights, thereby reducing misjudgments more accurately.
[0183] In addition, the described care device can reduce misrecognition by adopting various methods, and thus can effectively reduce the probability of misrecognition. Figure 4 Moreover, the described care device can judge the background in the scene and distinguish the foreground objects in the background at different times and under different lights, thereby reducing misjudgments more accurately.
[0184] In addition, the described care device can reduce misrecognition by adopting various methods, and thus can effectively reduce the probability of misrecognition. Figure 4 Moreover, the described care device can reduce misrecognition by adopting various methods, and thus can effectively reduce the probability of misrecognition.
[0185] Embodiment 5
[0186] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of another care device disclosed in the embodiments of the present invention. As Figure 5 shown, the care device may include:
[0187] A memory 501 storing executable program code;
[0188] A processor 502 coupled to the memory 501;
[0189] Wherein, the processor 502 calls the executable program code stored in the memory 501 and executes Figures 1 to 4 Any method for intelligently judging the human body posture.
[0190] The embodiments of the present invention disclose a computer-readable storage medium that stores a computer program, wherein the computer program enables a computer to execute Figures 1 to 2 Any method for intelligently judging the human body posture.
[0191] An embodiment of the present invention also discloses a computer program product. When the computer program product runs on a computer, it causes the computer to execute some or all of the steps of the methods in the above method embodiments.
[0192] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0193] The above has introduced in detail a method for intelligently judging human postures and a care device disclosed in the embodiments of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for intelligently judging human body postures, characterized in that, it includes: Obtain the human body data information in the video frame according to the video frame collected at the actual scene angle of the imaging device; wherein, the human body data information at least includes human body key point data information, human body coordinate data information, and human body key point bounding box data information; According to the human body key point data information, the human body coordinate data information, and the human body key point bounding box data information, calculate the corresponding human body posture probability in the video frame from the machine learning training model, including: Obtain N key point information from the human body key point data information, the human body coordinate data information, and the human body key point bounding box data information to obtain N-dimensional data information; wherein, the N key point information at least includes the coordinate positions of N key points; Calculate the distances between every two of the N key point coordinate positions to obtain N(N-1) / 2-dimensional data information; Calculate the central point coordinate of the set of the N key point coordinate positions to obtain two-dimensional data information; According to the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information, calculate K-dimensional judgment data information; wherein, the K-dimensional judgment data information is obtained by adding the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information; Perform a dot product operation on the K-dimensional judgment data information and the machine learning training model to obtain the corresponding human body posture probability in the video frame; Judge the posture with the maximum probability in the human body posture probability as the human body posture in the video frame.
2. The method according to claim 1, characterized in that, Before obtaining the human body data information in the video frame according to the video frame collected at the actual scene angle of the imaging device, the method further includes: Obtain the background information in the video frame according to the background maintenance algorithm; Regard the objects in the video frame other than the background information as the photographed human body to execute the step of obtaining the human body data information in the video frame according to the video frame collected at the actual scene angle of the imaging device.
3. The method according to claim 2, characterized in that, After regarding the objects in the video frame other than the background information as the photographed human body and before obtaining the human body data information in the video frame according to the video frame collected at the actual scene angle of the imaging device, the method further includes: If the area of the minimum rectangular bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea, determine it as a misrecognition of the human body image to obtain the key point coordinate positions of the misrecognized human body image; Regard the N key point coordinate positions other than the key point coordinate positions of the misrecognized human body image as the target key point coordinate positions to execute the step of obtaining the human body data information in the video frame according to the video frame collected at the actual scene angle of the imaging device.
4. The method according to claim 3, characterized in that, After using the objects other than the background information in the video frame as the photographed human body and before determining that the human body image is misrecognized to obtain the coordinate positions of the key points of the misrecognized human body image, the method further includes: Calculating a horizontal offset dx and a vertical offset dy between the center point coordinates and the horizontal coordinate of the midpoint of the lower edge of the image in the video frame; Calculating a horizontal direction attenuation value decayX and a vertical direction attenuation value decayY of the center point coordinates according to the horizontal offset dx and the vertical offset dy; Calculating a minimum bounding area judgment threshold minArea according to the horizontal direction attenuation value decayX and the vertical direction attenuation value decayY; Judging whether the area of the minimum rectangular bounding box corresponding to the N key point coordinate positions is lower than the minimum bounding area judgment threshold minArea.
5. The method according to any one of claims 1 to 4, characterized in that the method further includes: obtaining the machine learning training model established by the terminal device.
6. A care-taking device, characterized in that the care-taking device includes: A first obtaining unit, configured to obtain human body data information in the video frame according to a video frame collected at the actual scene angle of the imaging device; wherein, the human body data information at least includes human body key point data information, human body coordinate data information, and human body key point bounding box data information; A first calculating unit, configured to calculate the corresponding human body posture probability in the video frame from the machine learning training model according to the human body key point data information, the human body coordinate data information, and the human body key point bounding box data information; A first determining unit, configured to determine the posture with the maximum probability in the human body posture probability as the human body posture in the video frame; Wherein, the first calculating unit includes: An obtaining subunit, configured to obtain N key point information from the human body key point data information, the human body coordinate data information, and the human body key point bounding box data information to obtain N-dimensional data information; wherein, the N key point information at least includes N key point coordinate positions; A first calculating subunit, configured to calculate the distance between each of the N key point coordinate positions to obtain N(N-1) / 2-dimensional data information; A second calculating subunit, configured to calculate the center point coordinates of the set of the N key point coordinate positions to obtain two-dimensional data information; A third calculating subunit, configured to calculate K-dimensional judgment data information according to the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information; wherein, the K-dimensional judgment data information is obtained by adding the N-dimensional data information, the N(N-1) / 2-dimensional data information, and the two-dimensional data information; A fourth calculating subunit, configured to perform a dot product operation on the K-dimensional judgment data information and the machine learning training model to obtain the corresponding human body posture probability in the video frame.
7. A care-taking device, characterized in that the care-taking device includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the method for intelligently judging human body postures according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program causes a computer to execute the method for intelligently judging human body postures according to any one of claims 1-6.
Citation Information
Patent Citations
Target posture recognition method and device, and camera
CN111104816A
Attitude acquisition method and training method and device of key point coordinate positioning model
CN111126272A