An Infrared Image Fall Detection Method Based on Improved Alphapose
By improving the Alphapose algorithm, combining infrared images and human object detection network, the problem of high accuracy and error detection rate of existing fall detection methods in dim environments is solved, and high accuracy fall detection in dim environments is achieved.
Patent Information
- Application Number
- CN202111668347.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The existing fall detection methods cannot effectively consider the posture changes of the human body during the fall, and it is easy to cause false detection of fall-like actions.
The infrared image fall detection method based on improved Alphapose is adopted, video frames are collected through infrared cameras, and the human target is detected using the yolov5s human target detection network, and the posture category is generated with the posture classification library. The human target is input to the key point detection network to obtain the position information of the skeleton key point, and the fall judgment is made based on key point analysis and posture classification.
It can work normally in a dim environment, reduce false detection, improve the accuracy of fall detection, and prevent misidentification of fall actions through Mosaic data enhancement and key point analysis.
Smart Images

Figure CN114299050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human behavior recognition, and mainly relates to an infrared image fall detection method based on improved Alphapose. Background Art
[0002] Human fall detection can effectively detect fall behaviors in videos and reduce the risk that people cannot call for help in time after falling. Most traditional computer vision-based human fall detection methods use visible light images, and such methods have poor effects in dim environments affected by light. Existing fall detection methods can generally be divided into three categories: (1) detection methods based on Freeman chain codes; (2) detection methods based on key points; (3) detection methods based on the aspect ratio and the change rate of the centroid. These methods do not fully consider the posture change law of the human body during the fall process and are prone to misdetecting similar fall actions. Balancing the timing of the fall action and the relevance of the front and back actions is a difficult point in improving the accuracy.
[0003] In view of the above problems, it is urgent to improve the original fall detection technology. Summary of the Invention
[0004] 1. Object of the Invention
[0005] The present invention provides an infrared image fall detection method based on improved Alphapose to solve the technical problem in the above background art that the current existing fall detection methods cannot consider the posture change law of the human body during the fall process and are prone to misdetecting similar fall actions.
[0006] 2. Technical Solution
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] An infrared image fall detection method based on improved Alphapose includes:
[0009] Signal acquisition, collecting infrared video frames as input and inputting human target infrared images;
[0010] Generating key point position information and posture categories, using the yolov5s human target detection network to detect human target infrared images, comparing with the posture classification library to complete the classification of human postures to generate posture categories, and then inputting the obtained human target infrared images into the key point detection network to obtain the key point position information of the skeleton;
[0011] Key point analysis, using the height of the human target box in the previous frame as a reference object, calculating the relative movement speed of the hip joint key point and the change of the human aspect ratio, comparing with the key point position information and the posture category, and outputting a fall judgment signal.
[0012] Falling determination: After receiving the falling judgment signal, combined with pose classification, continue to detect the human pose in subsequent frames and output the falling judgment result.
[0013] Preferably, the method for generating the pose classification library includes:
[0014] Obtain images, and collect human body images through an infrared camera;
[0015] Label training samples, pre-define pose categories, use a labeling tool to frame the human body position in the infrared image, label the pose category of each human body, and generate the corresponding xml file;
[0016] Data preprocessing, adopting the method of Mosaic data augmentation;
[0017] Generate a deep training model, use Yolov5s as the human body detection deep network of Alphapose, train the human body detection model, generate a pose classification library, and complete the classification of the human body pose while extracting the human target region box and inputting the key point detection network to predict the key point positions.
[0018] Preferably, the Mosaic data augmentation method is to randomly extract four pictures from the dataset each time, and generate a new training picture by means of random cropping, random scaling, and random combination.
[0019] Preferably, the Yolov5s network directly classifies the human body pose, and at the same time, from the human body region box of the input human target infrared image, inputs the key point detection network to detect the key points of the human body skeleton, and outputs the key point position information and the predicted pose category together.
[0020] Preferably, the method for obtaining the human body aspect ratio is:
[0021] Assume that both the i-th frame and the (i - 1)-th frame are single-person infrared images, 2 ≤ i ≤ N, where N represents the total number of frames of the infrared video, and a hip joint key points are detected;
[0022] Set the upper left corner of the image as the origin, the horizontal right direction as the positive direction of the X-axis, and the vertical downward direction as the positive direction of the Y-axis to establish a rectangular coordinate system;
[0023] Denote the ordinate of the key point numbered n in the human body skeleton of the i-th frame as The center point M of point b and point c represents the hip joint, then the ordinate of point M in the i-th frame
[0024] Denote the height of the human body target box in the i-th frame as H i , and the width as W i , then the human body aspect ratio P in the i-th framei = W i / H i ;
[0025] The i, a, b, and c are numbers.
[0026] Preferably, the method for obtaining the relative movement speed of the hip joint key points is:
[0027] The relative movement speed of the midpoint M in the vertical direction in the i-th frame is
[0028] Preferably, the method for obtaining the output fall judgment signal is: set a threshold Th greater than 0, and the fall judgment reference value
[0029] When F i equals 1, output a fall judgment signal.
[0030] Preferably, the method for fall determination is:
[0031] In the subsequent human postures to be counted, if the human posture category in the i-th frame image is l i , then the corresponding human image score is s i , and the relationship between the two can be expressed by the formula:
[0032]
[0033] When the human posture category l i in the image is any one of the three postures of "sit_flat", "lie", and "push_up", the score of the human image is recorded as 1, otherwise the score is recorded as 0;
[0034] When F i = 1, that is, when the fall judgment signal is received in the i-th frame, continue to count the human postures of the subsequent d frames of images. If the sum of the cumulative image scores is greater than e, record it as FF i = 1, and the calculation formula is
[0035] When it is recorded as FF i = 1, output a judgment of fall;
[0036] The "sit_flat" means sitting on the ground, the "lie" means lying down, the "push_up" means propping on the ground, the FF i is the fall judgment result, the s i is the human image score, the l i is the human posture category, and the else means that F i is not equal to 1 or the total score of s i is less than or equal to e.
[0037] Preferably, d is equal to 20 and e is equal to 10.
[0038] 3. Beneficial effects
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] By using an infrared camera for data collection, the present invention enables normal use in a dim environment under the influence of light. At the same time, the infrared image can protect personal privacy and is applicable to all-weather human fall detection.
[0041] By using the Mosaic data augmentation method to preprocess the training data, through Mosaic data augmentation, the scene content can be enriched, the sample diversity can be enhanced, and the anti-interference ability of the target detection method can be improved.
[0042] By combining the key point analysis results and posture classification for further determination, after it is judged through key point analysis that a fall may occur, the human postures in subsequent multiple frames are continuously counted to prevent misidentifying actions similar to falling, such as quickly squatting down or bending over to pick up something, as a fall. Description of the drawings
[0043] Figure 1 It is a structural diagram of the improved Alphapose algorithm;
[0044] Figure 2 It is a schematic diagram of the Alphapose key point detection result;
[0045] Figure 3 It is a general method flow chart;
[0046] Figure 4 It is a schematic diagram of Mosaic data augmentation. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] Embodiment
[0049] Refer to Figures 1-4 , an infrared image fall detection method based on an improved Alphapose, which is characterized by including:
[0050] Signal acquisition: Collect infrared video frames as input and input the infrared image of the human target;
[0051] Generate key point position information and pose categories. Use the yolov5s human target detection network to detect the infrared image of the human target, compare it with the pose classification library to complete the classification of the human pose and generate pose categories, and then input the obtained infrared image of the human target into the key point detection network to obtain the key point position information of the skeleton;
[0052] Key point analysis: Using the height of the human target box in the previous frame as the reference object, calculate the relative movement speed of the hip joint key points and the transformation of the human body width-to-height ratio, and output a fall judgment signal.
[0053] Fall determination: After receiving the fall judgment signal, combined with pose classification, continue to detect the human pose in subsequent frames, compare it with the key point position information and pose categories, and output the fall judgment result.
[0054] In specific implementation, the input infrared image of the human target is the original infrared image, and it is very likely that there is no human body. Only after human target detection can the human target image be obtained. That is, the detected human body is cropped from the original infrared image and sent to the key point detection network.
[0055] Reference Figures 1-4 , the method for generating the pose classification library includes:
[0056] Obtain images: Collect human body images through an infrared camera;
[0057] Label training samples: Pre-define pose categories, use the labeling tool to frame the human body position in the infrared image, label the pose category of each human body, and generate the corresponding xml file;
[0058] Data preprocessing: Adopt the method of Mosaic data augmentation;
[0059] Generate a deep training model: Use Yolov5s as the human detection deep network of Alphapose to train the human detection model, generate the pose classification library, and while extracting the human target region box and inputting it into the key point detection network to predict the key point position, complete the classification of the human pose. The improved Alphapose algorithm of the present invention uses the Yolov5s network to directly classify the human pose, and at the same time extracts the human body region box from the input infrared image, inputs it into the human pose estimation network to detect the key points of the human skeleton, and outputs the key point information and the predicted pose category together.
[0060] In specific implementation, a training sample library is first constructed, which includes a large number of pictures and a large number of annotation files. Each picture corresponds to an annotation file, and the annotation file records the position of the human body in its corresponding picture and the corresponding category. The training sample library is trained through a training network to obtain the target detection model we need. Then, the trained model is used to infer the input infrared pictures. As long as a human target is detected, the probabilities of the corresponding single person in eight categories will be obtained. Finally, the category with the highest probability is selected as its final category.
[0061] Reference Figures 1-4 , the way of Mosaic data augmentation is to randomly extract four pictures from the dataset each time, and generate a new training picture by means of random cropping, random scaling and random combination. Through Mosaic data augmentation, the scene content can be enriched, the sample diversity can be enhanced, and the anti-interference ability of the target detection algorithm can be improved.
[0062] Reference Figures 1-4 , the Yolov5s network completes the direct classification of human postures. At the same time, from the human body region box of the input infrared image of the human target, the key point detection network for the human body is input to detect the key points of the human skeleton, and the position information of the key points and the predicted posture category are output together.
[0063] Reference Figures 1-4 , in specific implementation, during the falling process, when the height of the human target box in the previous frame is significantly smaller than its width, even if the falling speed remains unchanged, due to the decrease in the height of the target box in the previous frame, the calculated speed will be too large. Therefore, this method introduces the aspect ratio of the human target box in the previous frame as a limiting condition to prevent the slight fluctuation of point M from being misidentified as a fall due to an excessive aspect ratio of the target box in the previous frame. The method for obtaining the aspect ratio of the human body is as follows:
[0064] Let the i-th frame and the (i - 1)-th frame both be single-person infrared images, 2 ≤ i ≤ N, where N represents the total number of frames of the infrared video, and a key points of the hip joint are detected. The present invention does not make specific limitations on the number of key points of the hip joint. In this embodiment, 18 key points of the human skeleton are preferably adopted;
[0065] Set the upper left corner of the image as the origin, the horizontal right direction as the positive direction of the X axis, and the vertical downward direction as the positive direction of the Y axis to establish a rectangular coordinate system;
[0066] Denote the ordinate of the key point numbered n in the human skeleton of the i-th frame as The center point M of point b and point c represents the hip joint, then the ordinate of point M in the i-th frame
[0067] Denote the height of the human target box in the i-th frame as H i , and the width as W i , then the aspect ratio P of the human body in the i-th framei = W i / H i ;
[0068] i, a, b, c are numbers. The present invention does not limit the specific values of i, a, b, and c. In this embodiment, b is preferably 11 and c is 12. The corresponding ordinate in this embodiment The formula is
[0069] Reference Figures 1-4 , when a human body target standing or walking falls, the most intuitive manifestation is the rapid decline of the hip joint position in the vertical direction. Therefore, by monitoring the moving speed of the hip joint in the sequence of frames, the generated falling action can be detected in a timely manner. However, as the distance between the moving human body target and the camera increases, the displacement of its hip joint in the image becomes smaller and smaller, which is likely to cause missed detection. To address this problem, the present invention uses the height of the human body target box in the previous frame as a reference object to calculate the relative moving speed of the hip joint key point. The method for obtaining the relative moving speed of the hip joint key point is as follows:
[0070] The relative moving speed of point M in the vertical direction in the i-th frame is
[0071] Reference Figures 1-4 , the method for obtaining the falling judgment signal is: set a threshold Th greater than 0, and the falling judgment reference value
[0072] When F i equals 1, output the falling judgment signal.
[0073] Reference Figures 1-4 , the method for falling determination is:
[0074] In the subsequent human body postures, if the human body posture category in the i-th frame image is l i , then the corresponding human body image score is s i , and the relationship between the two can be expressed by the formula:
[0075]
[0076] When the human body posture category l in the image i is any one of the three postures of "sit_flat", "lie", and "push_up", the score of this human body image is recorded as 1, otherwise the score is recorded as 0;
[0077] When F i = 1, that is, when the falling judgment signal is received in the i-th frame, continue to count the human body postures of the subsequent d frames of images. If the sum of the cumulative image scores is greater than e, it is recorded as FF i= 1, and the calculation formula is
[0078] When recorded as FF i = 1, the output judgment is a fall;
[0079] sit_flat is sitting on the ground, lie is lying down, push_up is propping on the ground, FF i is the fall judgment result, s i is the human body image score, l i is the human body pose category, else is F i is not equal to 1 or the total score of s i is less than or equal to e.
[0080] Reference Figures 1-4 , d is equal to 20, e is equal to 10. The present invention does not specifically limit the specific values of d and e. In this embodiment, d is preferably equal to 20 and e is equal to 10. The corresponding fall result judgment formula is
[0081] In a specific implementation, the process of generating the key point position information and the pose category is the running process of the improved Alphapose algorithm. The structure diagram of the improved Alphapose algorithm is for reference Figure 1 , The Alphapose algorithm is a skeleton key point detection algorithm that can detect the human skeleton in an image. It proposes a regional multi-person pose estimation framework RMPE, which is mainly composed of SSTN, PP-NMS, PGPG, and Parallel SPPE. Among them, PGPG is used to generate a large number of training samples, and Parallel SPPE acts as an additional regularization term to avoid local minima, while SSTN is composed of STN, SPPE, and SDTN. Alphapose first uses an object detection algorithm to detect a single image to obtain a single-person human object box, which is used as the input of RMPE and enters the SSTN module. The detected human object box may have the target deviated from the center or the proportion of the human body in the object box is too small, affecting the subsequent pose estimation effect. Therefore, it is necessary to first use STN to extract a high-quality single-person area, then input SPPE to obtain the estimated pose, and then SDTN inverse-transforms the estimated pose into the original human area box. Redundant human area boxes input into the pose estimation network may detect redundant poses.
[0082] Alphapose proposes PP-NMS to eliminate redundant poses. It defines a pose distance to measure the similarity between poses and establishes a criterion for eliminating redundant poses based on this. PP-NMS first selects the pose with the highest confidence as a reference and eliminates the bounding boxes in the area close to this reference according to the elimination criterion. This process is repeated multiple times until all redundant detection boxes are eliminated. The present invention uses Yolov5s as the human body detector of Alphapose. The improved algorithm uses the Yolov5s network to directly classify human poses, extracts the human body bounding box from the input infrared image at the same time, inputs it into the human pose estimation network to detect the key points of the human skeleton, and outputs the key point information and the predicted pose category together.
[0083] Among them, SSTN is the Symmetric Spatial Transformer Network, PP-NMS is the Parametric Pose Non-Maximum Suppression, PGPG is the Pose-Guided Proposal Generator, Parallel SPPE is the Parallel Single-Person Pose Estimator, RMPE is the Regional Multi-Person Pose Estimation Framework, STN is the Spatial Transformer Network, SPPE is the Single-Person Pose Estimator, and SDTN is the Spatial Inverse Transformer Network.
[0084] In the real-time detection process of the present invention, the real-time video stream is used as the input to detect the human poses and the key points of the skeleton in each frame of the current picture. Starting from the second frame, if the hip joint positions of the same person are detected in both the previous frame and the current frame, that is, the "unoccluded" state, then key point analysis is performed to calculate the relative movement rate and direction of the hip joint key points between these two adjacent frames. When the conditions for possible falling are met, a fall determination is made, and the pose categories in the subsequent 20 frames of images are counted. If the final fall condition is met, it is determined as a fall. The schematic diagram of a complete fall detection process in the case of real-time detection is as Figure 3 shown.
[0085] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An infrared image fall detection method based on improved Alphapose, characterized in that, Including: Signal acquisition, which acquires infrared video frames as input and inputs the infrared image of the human target; Generating key point position information and pose categories, using the yolov5s human target detection network to detect the infrared image of the human target, comparing it with the pose classification library to complete the classification of the human pose to generate the pose categories, and then inputting the obtained infrared image of the human target into the key point detection network to obtain the key point position information of the skeleton; Key point analysis, using the height of the human target box in the previous frame as a reference object, calculating the relative movement speed of the hip joint key points and the transformation of the human body width-to-height ratio, comparing with the key point position information and the pose categories, and outputting a fall judgment signal; Fall determination, after receiving the fall judgment signal, combining the pose classification, continuing to detect the human pose in subsequent frames, and outputting a fall judgment result; The method for obtaining the human body width-to-height ratio is: Suppose the th frame and the th frame are both single-person infrared images, , represents the total number of frames of the infrared video, and a hip joint key points are detected; Taking the upper left corner of the image as the origin, the horizontal right direction as the positive direction of the X axis, and the vertical downward direction as the positive direction of the Y axis to establish a rectangular coordinate system; Record the vertical coordinate of the key point numbered in the skeleton of the human body in the frame. The center point of point b and point c represents the hip joint. Then the vertical coordinate of point in the i-th frame is Record the height of the human target box in the th frame as , and the width as . Then the aspect ratio of the human body in the th frame is The i, a, b, and c are numbers; The method for obtaining the relative movement speed of the hip joint key points is: Frame midpoint has a relative moving speed in the vertical direction of ; The method for obtaining the output fall judgment signal is: set a threshold Th greater than 0, and the fall judgment reference value ; When equals 1, output the fall determination signal.
2. The infrared image fall detection method based on improved Alphapose according to claim 1, characterized in that: The method for generating the pose classification library includes: Obtaining images, collecting human body images through an infrared camera; Labeling training samples, predefining pose categories, using a labeling tool to frame the human body position in the infrared image, labeling the pose categories of each human body, and generating corresponding xml files; Data preprocessing, adopting the Mosaic data augmentation method; Generating a deep training model, using Yolov5s as the human detection deep network of Alphapose, training the human detection model to generate the pose classification library, and while extracting the human target region box and inputting it into the key point detection network to predict the key point position, completing the classification of the human pose.
3. The infrared image fall detection method based on improved Alphapose according to claim 2, characterized in that: The Mosaic data augmentation method is to randomly extract four pictures from the dataset each time and generate a new training picture by means of random cropping, random scaling, and random combination.
4. The infrared image fall detection method based on improved Alphapose according to claim 2, characterized in that: The Yolov5s network completes the direct classification of the human pose, and at the same time, from the human region box of the input infrared image of the human target, inputs it into the key point detection network to detect the key points of the human skeleton, and outputs the key point position information and the predicted pose categories together.
5. The infrared image fall detection method based on improved Alphapose according to claim 1, characterized in that: The method for fall determination is: In the subsequent human pose statistics, for the th frame image, if the human pose category is , then the corresponding human image score is , and the relationship between the two can be expressed by the formula: ; When the human body pose category in the image is any one of the three poses "sit_flat", "lie", and "push_up", the score of the human body image is recorded as 1, otherwise the score is recorded as 0; When , that is, when the th frame receives the fall judgment signal, continue to count the human postures of the subsequent d frames of images. If the sum of the cumulative image scores is greater than e, it is recorded as , and the calculation formula is ; When recorded as , the output is judged as a fall; The "sit_flat" means sitting on the ground, the "lie" means lying down, and the "push_up" means propping on the ground. The is the fall judgment result, and the is the human body image score, and the is the human body posture category. The "else" means not equal to 1 or the total score is less than or equal to e.
6. The infrared image fall detection method based on improved Alphapose according to claim 5, characterized in that: The d is equal to 20 and the e is equal to 10.
Citation Information
Patent Citations
Pedestrian falling detection method based on Gaussian mixture model and neural network
CN110991274A
Behavior detection method and device and computer readable storage medium
CN112395978A