Image-based fall behavior recognition method and device
By acquiring the human skeleton points, outlines, and clarity in surveillance images, and combining the width, height, and skeleton point coordinates, the system can identify fall states, solving the accuracy problem of fall behavior in surveillance scenarios, improving the recognition rate, and reducing the false recognition rate.
Patent Information
- Application Number
- CN202211666081.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-12-23
AI Technical Summary
Existing technologies struggle to accurately identify falls in surveillance scenarios, especially in complex and ever-changing environments where factors such as occlusion between people, objects obstructing people, people occluding themselves, and lighting conditions reduce the accuracy of identification.
By acquiring human skeleton points, human contours, and image clarity from the image to be identified, it is determined whether there is a complete human image in the image with a clarity greater than a preset value. The fall state is identified based on the width, height, and coordinates of the skeleton points. Deep learning models and pre-trained state detection models are used to improve the recognition accuracy.
It effectively avoids misidentification in cases of image occlusion and low clarity, improving the accuracy of fall recognition, especially in cases of falls from multiple angles and complex scenarios, reducing the misidentification rate.
Smart Images

Figure CN116229502B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to an image-based method and device for recognizing fall behavior. Background Technology
[0002] The identification of personnel falling behavior is an important technology used in the field of safety supervision. This technology can analyze video footage in real time in a monitoring scenario to see if there is any abnormal behavior such as falling. If a fall behavior is detected, it can be promptly reported to the relevant management personnel for handling.
[0003] However, in real-world surveillance scenarios, the direction in which a person falls is often uncertain; they may fall at any angle. In such cases, existing technologies that determine the fall by calculating the positional relationships between body parts struggle to guarantee accuracy. Furthermore, surveillance scenarios are often complex and dynamic, typically involving occlusion between people, objects obstructing people, and the effects of lighting on video image quality. All of these factors contribute to a decrease in the final recognition accuracy. Summary of the Invention
[0004] In view of the above, this application provides an image-based method and device for recognizing fall behavior, the purpose of which is to improve the accuracy of fall behavior recognition.
[0005] In a first aspect, this application provides an image-based method for recognizing fall behavior, the method comprising:
[0006] Obtain human skeleton points, human contours, and / or the clarity of the image to be identified;
[0007] Based on the human skeleton points, the human outline, and / or the clarity of the image to be identified, determine whether there is a complete human image in the image to be identified with a clarity greater than a preset value.
[0008] If so, based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points, identify whether the human body in the image to be identified is in a fallen state.
[0009] Secondly, this application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0010] Memory, used to store computer programs;
[0011] When a processor executes a program stored in memory, it implements the image-based fall behavior recognition method described in any embodiment of the first aspect.
[0012] The technical solutions provided in this application have the following advantages compared with the prior art:
[0013] This application obtains human skeleton points, human contours, and / or the clarity of the image to be identified. Based on the clarity of the human skeleton points, human contours, and / or the image to be identified, it detects whether there is a complete human image in the image to be identified with a clarity greater than a preset value. Only when there is a complete human image in the image to be identified with a clarity greater than the preset value is the image to be identified whether the human body in the image to be identified is in a fallen state. This can avoid misidentification caused by still recognizing the image when the human body in the image is occluded or the image clarity is low. In real-world scenarios, there are often factors such as occlusion between people, occlusion by objects, occlusion by the person themselves, and low image clarity. Furthermore, since the human skeleton points differ when the width of the human image is greater than its height compared to when the height is greater than its width, this application, when detecting a complete human image with a clarity greater than a preset value in the image to be identified, identifies whether the human in the image to be identified is in a fallen state based on the width of the human image, the height of the human image, and / or the position coordinates of the human skeleton points. By considering the width and height of the human image as factors in determining whether the human is in a fallen state, the accuracy of fall behavior recognition can be improved. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a preferred embodiment of the image-based fall behavior recognition method of this application;
[0017] Figure 2 This is a schematic diagram of human skeleton points in an embodiment of this application;
[0018] Figure 3 This is a schematic diagram of the human body outline in an embodiment of this application;
[0019] Figure 4 This is a schematic diagram illustrating a human body image in an embodiment of this application where the width is greater than the height.
[0020] Figure 5 This is a schematic diagram illustrating a human body image in an embodiment of this application where the width is less than or equal to the height.
[0021] Figure 6 This is a schematic diagram of a preferred embodiment of the image-based fall behavior recognition device of this application;
[0022] Figure 7 This is a schematic diagram of a preferred embodiment of the electronic device of this application;
[0023] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0025] It should be noted that the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of those features. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.
[0026] This application provides an image-based method for recognizing fall behavior. (Refer to...) Figure 1 The diagram shown is a flowchart illustrating an embodiment of the image-based fall behavior recognition method of this application. This method can be executed by an electronic device, which can be implemented in software and / or hardware. The image-based fall behavior recognition method includes:
[0027] Step S10: Obtain the human skeleton points, human contour and / or the clarity of the image to be identified in the image to be identified;
[0028] Step S20: Based on the human skeleton points, the human outline, and / or the clarity of the image to be identified, determine whether there is a complete human image in the image to be identified with a clarity greater than a preset value;
[0029] Step S30: If yes, identify whether the human body in the image to be identified is in a fallen state based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points.
[0030] In this embodiment, the image to be recognized may be an image captured in real time by a surveillance camera in a surveillance scenario. For example, a video stream is obtained from the surveillance camera, and the decoded video stream is used as a frame image to be recognized. Taking the image captured in real time by the camera as an example to illustrate this solution, it can be understood that the actual application scenario of this solution is not limited to this. The image to be recognized may also be a static image captured non-real-time, or an image pre-stored in a database, etc. It should be noted that if there is no human body image in the image to be recognized, that is, no human body skeleton points and human body contours can be obtained from the image to be recognized, then the next frame image corresponding to the image to be recognized can be used as the image to be recognized, and it is determined whether there is a human body image in the image to be recognized. If so, the human body skeleton points, human body contours and / or the clarity of the image to be recognized are obtained.
[0031] The human body skeleton points in the image to be recognized can be obtained by using a human body pose estimation algorithm (such as the OpenPose algorithm). After the image to be recognized is grayscale processed, the human body contour in the image to be recognized can be obtained. The clarity of the image to be recognized can be obtained by using the Laplacian gradient method.
[0032] Due to the complexity of the surveillance scenario, there are usually factors such as occlusion between people, occlusion of people by objects, self-occlusion of people, changes in lighting, and poor imaging quality of the video picture in the surveillance scenario, which will affect the final recognition accuracy. Therefore, it is necessary to detect whether there is a complete human body image with a clarity greater than a preset value in the image to be recognized according to the human body skeleton points, human body contours and / or the clarity of the image to be recognized. If the human body skeleton points are incomplete, or the human body contour is incomplete, or the clarity of the image to be recognized is less than or equal to the preset value (for example, 0.5), then there is no complete human body image with a clarity greater than the preset value in the image to be recognized. If the human body skeleton points are complete and the human body contour is complete, and the clarity of the image to be recognized is greater than the preset value, then there is a complete human body image with a clarity greater than the preset value in the image to be recognized. The preset value can be modified through a JSON configuration file without modifying the code, which is convenient for real-time adjustment in the actual application scenario.
[0033] As shown Figure 2 in the figure, it is a schematic diagram of human body skeleton points in an embodiment of the present application. The human body skeleton points include the position information of 17 key points of the human body skeleton, that is, the nose (point 0), the left eye (point 1), the right eye (point 2), the left ear (point 3), the right ear (point 4), the left shoulder (point 5), the right shoulder (point 6), the left elbow (point 7), the right elbow (point 8), the left wrist (point 9), the right wrist (point 10), the left hip (point 11), the right hip (point 12), the left knee (point 13), the right knee (point 14), the left ankle (point 15) and the right ankle (point 16). As shown Figure 3 The diagram shown is a schematic representation of the human body contour in an embodiment of this application. The human body contour is the semantic segmentation information of three independent parts: the head, upper body, and lower body. The sharpness of the image to be recognized includes the imaging quality of the human body in the image (e.g., whether it is blurry, dark, or clearly distinguishable to the naked eye). When no complete human body image with a sharpness greater than a preset value is detected in the image to be recognized, the image to be recognized can be filtered out, and the next frame of the image to be recognized can be used as the image to be recognized for continued detection. Filtering out images that do not contain complete images and images with a sharpness less than or equal to the preset value, and retaining only clear images in which the complete human body can be seen, can effectively improve the recognition rate of human fall behavior recognition.
[0034] For example, when the human skeletal points have the above-mentioned Figure 2 The image is considered complete if it contains a human body image with a resolution greater than the preset value, provided that the image contains the 17 key points shown and the semantic segmentation information of the head, upper body, and lower body of the human body. If the image only contains the upper body, then the human body image in the image is incomplete.
[0035] In actual monitoring scenarios, falls often occur from multiple angles. Figure 4 The image shown is a schematic diagram illustrating a human body image in an embodiment of this application where the width is greater than the height. (See reference...) Figure 5 The diagram illustrates a human image where the width is less than or equal to the height, as per an embodiment of this application. When a complete human image with a clarity greater than a preset value is detected in the image to be identified, the system identifies whether the human body in the image to be identified is in a fallen state based on the width of the human image, the height of the human image, and / or the position coordinates of the human skeleton points. For example, in the case where the width of the human image is greater than the height of the human image (… Figure 4 As shown in the example, the system identifies whether a person in the image is in a fallen state based on the coordinates of the human skeleton points. This applies to cases where the width of the human image is less than or equal to the height of the human image. Figure 5 As shown in the example, a deep learning model is used to identify whether a person has fallen. By considering the relationship between the width and height of the human image, different recognition methods can be selected to improve accuracy and reduce false positives.
[0036] Specifically, identifying whether the human body in the image to be identified is in a fallen state based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points includes:
[0037] Determine whether the width of the human body image is greater than the height of the human body image shown;
[0038] If the width of the human body image is greater than the height of the human body image, the system identifies whether the human body in the image to be identified is in a fallen state based on the ratio of the width to the height of the human body image and the position coordinates of the human skeleton points.
[0039] If the width of the human body image is less than or equal to the height of the human body image, the system identifies whether the human body in the image to be identified is in a fallen state based on a pre-trained state detection model.
[0040] When the width of a human image is greater than its height, the coordinates of the human skeleton points in a standing and fallen state will differ. Therefore, the ratio of the width to the height of the human image, along with the coordinates of the skeleton points, can be used to identify whether a person in the image is in a fallen state. However, when the width of a human image is less than or equal to its height, the coordinates of the skeleton points in a standing and fallen state are very close, making it impossible to determine a fallen state using only the skeleton coordinates. In this case, a pre-trained state detection model can be used to identify whether a person in the image is in a fallen state. This state detection model is a binary classification model, trained using a ResNet18 network. During training, images of a person in a fallen state are used as positive samples and labeled, containing data on fallen states from various angles and directions. Images of a person not in a fallen state are used as negative samples and labeled, containing data on non-fallen states such as squatting, sitting, and standing normally. After the human body image is detected by a binary classification state detection model, the score probabilities of the two state categories are obtained. The score probability of the falling state is used to determine whether the human body in the human body image is in a falling state. For example, if the score of the falling state is greater than 0.6, it is determined whether the human body in the human body image is in a falling state.
[0041] Since a person's fall usually lasts for a period of time in real-world scenarios, after identifying whether the person in the image to be identified is in a fall state, it is possible to further identify each frame of the image within a certain time period. Based on the identification results of multiple frames within a time period, it can be determined whether the person is in a fall state. For example, 12 frames of the image to be identified within the next second can be obtained as the images to be identified, and the 12 frames can be identified sequentially. Only when 10 out of the 12 frames are identified as the person in a fall state can the person be determined to be in a fall state, thereby improving the accuracy of identifying the person's fall state.
[0042] This application obtains human skeleton points, human contours, and / or the clarity of the image to be identified. Based on the clarity of the human skeleton points, human contours, and / or the image to be identified, it detects whether there is a complete human image in the image to be identified with a clarity greater than a preset value. Only when there is a complete human image in the image to be identified with a clarity greater than the preset value is the image to be identified whether the human body in the image to be identified is in a fallen state. This can avoid misidentification caused by still recognizing the image when the human body in the image is occluded or the image clarity is low. In real-world scenarios, there are often factors such as occlusion between people, occlusion by objects, occlusion by the person themselves, and low image clarity. Furthermore, since the human skeleton points differ when the width of the human image is greater than its height compared to when the height is greater than its width, this application, when detecting a complete human image with a clarity greater than a preset value in the image to be identified, identifies whether the human in the image to be identified is in a fallen state based on the width of the human image, the height of the human image, and / or the position coordinates of the human skeleton points. Using the width and height of the human image as factors to consider when determining whether the human is in a fallen state can improve the accuracy of the identification.
[0043] In one embodiment, obtaining human skeleton points, human contours, and / or the sharpness of the image to be identified includes:
[0044] Using a pre-built human detection model, the presence of human information in the image to be identified is detected;
[0045] If human information is present in the image to be identified, the human skeleton points, human contours, and / or the clarity of the image to be identified are obtained according to the pre-trained fusion model.
[0046] The human detection model can be trained based on a lightweight YOLOv5-s model. Using a lightweight model reduces the system's computational load, ensuring fast operation when deployed on edge devices and meeting practical recognition needs. The human detection model detects the presence of a human body in an image to be identified. If no human body is found, the next frame is used as the next image for detection. If a human body is present, the model uses a pre-trained fusion model to obtain the human skeleton points, human contours, and / or the image's clarity.
[0047] Further, obtaining the human skeleton points, human contours, and / or the sharpness of the image to be identified based on the pre-trained fusion model includes:
[0048] The human skeleton points in the image to be identified are obtained using the human skeleton point detection network in the fusion model.
[0049] The human contour in the image to be identified is obtained by using the human contour detection network in the fusion model.
[0050] The image sharpness detection network in the fusion model is used to obtain the sharpness of the image to be identified.
[0051] The pre-trained fusion model integrates three branch networks: a human skeleton point detection network for outputting human skeleton point information, a human contour detection network for outputting human contour information, and an image sharpness detection network for outputting the sharpness of the image to be recognized. A High-Resolution Net (HRNet) is used to construct the branch networks for outputting human skeleton point and human contour information; HRNet has wide applications in image segmentation and human skeleton point detection tasks. A 10-layer convolutional neural network structure is then constructed as the branch network for outputting human image quality information. The model parameters of the three branch networks are fused to obtain the fusion model that outputs the three types of information. Because the original HRNet network structure has too many layers, it causes a decrease in the overall recognition speed of the model during training. Therefore, model compression techniques (network channel pruning) are used to prune the original HRNet network structure, reducing the number of parameters and the model size, thereby reducing the computational load and ensuring the running speed of the model when deployed on edge devices, thus meeting actual recognition requirements. During training, since the three branch networks cannot be trained simultaneously, they are trained sequentially. First, the human skeleton detection network is trained separately. After its training is complete, its parameters are frozen. Next, the human contour detection network is trained separately. After its training is complete, the parameters of the two trained branch models are frozen. Finally, the image sharpness detection network is trained separately. The final result is a fusion model of the three branch networks. After training the fusion model, the image to be recognized is input into it to obtain the human skeleton points, human contours, and / or the sharpness of the image.
[0052] Further, determining whether a complete human image with a resolution greater than a preset value exists in the image to be identified based on the human skeleton points, the human outline, and / or the clarity of the image to be identified includes:
[0053] Determine whether the human skeleton points in the image to be identified are complete;
[0054] If so, determine whether the human body outline in the image to be identified is complete;
[0055] If so, determine whether the clarity of the image to be identified is greater than a preset value;
[0056] If so, then the image to be identified contains a complete human image with a clarity greater than a preset value.
[0057] To determine whether the human skeleton points in the image to be identified are complete, the confidence scores of multiple key points in the human skeleton points can be used to determine whether the human body is occluded. Six points in the human skeleton points—left shoulder, right shoulder, left hip, right hip, left knee, and right knee—can be used as the basis for judgment. When the confidence score of one of the key points is lower than a preset value (e.g., 0.5), it indicates that the human body may be occluded, that is, the human skeleton points in the image to be identified are incomplete.
[0058] To determine whether the human body contour in the image to be recognized is complete, we can determine whether the contour information of the head, upper body and lower body can be obtained by completely segmenting the image to be recognized, and determine whether the human body contour in the image is complete based on the area covered by the contour. This branch network performs further detection on the basis of human skeleton point detection, achieving the effect of dual detection.
[0059] The purpose of determining the sharpness of an image to be recognized is to detect whether there are any blurry or indistinguishable human figures in the image. During training, this branch network categorizes the sharpness of human figures in the image into three categories: normal, blurry, and overly dark. Therefore, the judgment is made by checking whether the score probability of the normal category is greater than a preset value (e.g., 0.5). If the score probability of the normal category is greater than the preset value, the human figure in the image is determined to be sharp; otherwise, the human figure in the current image is determined to be blurry or overly dark. Only when the human skeleton points in the image to be recognized are complete, the human figure outline is complete, and the sharpness of the image to be recognized is greater than the preset value, is the image to be recognized considered to contain a complete human figure with a sharpness greater than the preset value.
[0060] Further, the step of identifying whether the human body in the image to be identified is in a fallen state based on the ratio of the width to the height of the human body image and the position coordinates of the human skeleton points includes:
[0061] Using the position coordinates of the human skeleton points, calculate the first ordinate value of the center point between the left and right hip points of the human body, and calculate the second ordinate value of the center point between the left and right knee points of the human body.
[0062] Based on the coordinate information of the left and right shoulder points of the human body, and the coordinate information of the left and right hip points of the human body, calculate the height and width of the upper body of the human body.
[0063] Using the position coordinates of the human skeleton points, calculate the first distance in the horizontal direction and the second distance in the vertical direction from the head to the ankle of the human body.
[0064] Based on the relationship between the upper body height and the upper body width, the relationship between the ratio of the width to the height of the human body image and a preset threshold, the relationship between the first ordinate value and the second ordinate value, and the relationship between the first distance and the second distance, it is determined whether the human body in the image to be identified is in a fallen state.
[0065] Assume the human skeleton points are represented by the following symbols: left shoulder point p[lshoulder], right shoulder point p[rshoulder], left hip point p[lhip], right hip point p[rhip], left ankle point p[lankle], right ankle point p[rankle], left knee point p[lknee], right knee point p[rknee], nose point p[nose]. "p[].x" and "p[].y" represent the x-coordinate and y-coordinate of the point in a coordinate system with the top-left corner of the image to be identified as the origin, respectively. For example, p[lshoulder].x represents the x-coordinate of the left shoulder point, p[lshoulder].y represents the y-coordinate of the left shoulder point, p[rshoulder].x represents the x-coordinate of the right shoulder point, and p[rshoulder].y represents the y-coordinate of the right shoulder point.
[0066] Using the position coordinates of the human skeleton points, calculate the first ordinate value of the center point between the left and right hip points. The first ordinate value of the center point between the left and right hip points is denoted as hip_c_y, where hip_c_y = (p[lhip].y + p[rhip].y) / 2;
[0067] The second ordinate of the center point between the left and right knee points of a human body is denoted as knee_c_y, where Knee_c_y = (p[lknee].y + p[rknee].y) / 2.
[0068] Based on the coordinates of the left and right shoulder points and the left and right hip points, the height and width of the upper body can be calculated. For example, the distance between the left and right shoulder points can be calculated as the width of the upper body, and the distance between the shoulder and hip can be calculated as the height of the upper body.
[0069] By calculating the coordinates of the center point between the left and right ankles and obtaining the coordinates of the nose, the first distance in the horizontal direction and the second distance in the vertical direction from the head to the ankle can be calculated.
[0070] Based on the relationship between the upper body height and the upper body width, the relationship between the ratio of the width to the height of the human body image and a preset threshold, the relationship between the first ordinate value and the second ordinate value, and the relationship between the first distance and the second distance, it is determined whether the human body in the image to be identified is in a fallen state.
[0071] The step of identifying whether the human body in the image to be identified is in a fallen state based on the relationship between the upper body height and the upper body width, the relationship between the ratio of the width to the height of the human body image and a preset threshold, the relationship between the first ordinate value and the second ordinate value, and the relationship between the first distance and the second distance includes:
[0072] If the height of the upper body is greater than the width of the upper body, and the first ordinate value is greater than or equal to the second ordinate value, and the first distance is less than the second distance, and the ratio is greater than the preset threshold, then the human body in the human image is in a fallen state.
[0073] Otherwise, the human body in the human body image is not in a fallen state.
[0074] The preset threshold is a value greater than 1. When the preset threshold is greater than 1, it means that the width of the human body image is greater than the height of the human body image. If the upper body height is less than or equal to the upper body width, or the first ordinate value is less than the second ordinate value, or the first distance is greater than or equal to the second distance, then the human body in the image is not in a fallen state. By combining the upper body height and width, the horizontal and vertical distances from the head to the ankles, the first ordinate value of the center point between the left and right hip points and the second ordinate value of the center point between the left and right knee points, and the ratio of the width to the height of the human body image, a more accurate determination of whether the human body in the image is in a fallen state can be made.
[0075] Further, the step of calculating the upper body height and upper body width of the human body based on the coordinate information of the left and right shoulder points and the coordinate information of the left and right hip points includes:
[0076] Calculate the first difference between the x-coordinate of the left shoulder point and the x-coordinate of the right shoulder point, and calculate the second difference between the x-coordinate of the left hip point and the x-coordinate of the right hip point;
[0077] Calculate the third difference between the ordinate value of the left hip point and the ordinate value of the left shoulder point, and calculate the fourth difference between the ordinate value of the right hip point and the ordinate value of the right shoulder point;
[0078] Calculate the first average of the absolute values of the first difference and the second difference, and use the first average as the upper body width;
[0079] Calculate the second average of the absolute values of the third difference and the fourth difference, and use the second average as the upper body height.
[0080] p[lshoulder].x represents the x-coordinate of the left shoulder point, p[lshoulder].y represents the y-coordinate of the left shoulder point, p[rshoulder].x represents the x-coordinate of the right shoulder point, p[rshoulder].y represents the y-coordinate of the right shoulder point, p[lhip].x represents the x-coordinate of the left hip point, p[lhip].y represents the y-coordinate of the left hip point, p[rhip].x represents the x-coordinate of the right hip point, p[rhip].y represents the y-coordinate of the right hip point.
[0081] First difference = p[lshoulder].x p[rshoulder].x;
[0082] The second difference = p[lhip].x p[rhip].x;
[0083] The third difference = p[lhip].yp[lshoulder].y;
[0084] The fourth difference = p[rhip].yp[rshoulder].y;
[0085] Width of the upper body:
[0086] u_w=(abs(p[lshoulder].xp[rshoulder].x)+abs(p[lhip].xp[rhip].x)) / 2;
[0087] Height of the human upper body:
[0088] u_h = (abs(p[lhip].yp[lshoulder].y) + abs(p[rhip].yp[rshoulder].y)) / 2. Here, abs represents the absolute value.
[0089] Further, the step of using the position coordinates of the human skeleton points to calculate the first distance in the horizontal direction and the second distance in the vertical direction from the head to the ankle of the human body includes:
[0090] Using the position coordinates of the human skeleton points, calculate the third abscissa and third ordinate of the center point between the left and right ankle points of the human body;
[0091] Obtain the fourth abscissa and fourth ordinate values of the nose point of the human body from the position coordinates of the human skeleton points;
[0092] The absolute value of the subtraction of the third horizontal coordinate value from the fourth horizontal coordinate value is taken as the first distance;
[0093] The absolute value of the difference between the fourth and third ordinate values is taken as the second distance.
[0094] The x-coordinate of the center point between the left and right ankle points is denoted as the third x-coordinate value, ankle_c_x, where ankle_c_x = (p[lnakle].x + p[rankle].x) / 2. The y-coordinate is denoted as the third y-coordinate value, ankle_c_y, where ankle_c_y = (p[lankle].y + p[rankle].y) / 2. The x-coordinate of the nose point can be directly obtained from the coordinates of the human skeleton points and is denoted as the fourth x-coordinate value, p[nose].x, and the y-coordinate of the nose point is denoted as the fourth y-coordinate value, p[nose].y. The absolute value of the result obtained by subtracting the third x-coordinate value from the fourth x-coordinate value is taken as the distance from the head to the ankle in the x-coordinate direction of the human body and is denoted as the first distance, head_ankle_dis_x, i.e., head_ankle_dis_x = abs(p[nose].x - ankle_c_x).
[0095] The absolute value of the difference between the fourth and third ordinate values is taken as the distance from the head to the ankle of the human body, denoted as the second distance head_ankle_dis_y. That is, head_ankle_dis_y = abs(p[nose].y-ankle_c_y), where abs represents the symbol for taking the absolute value.
[0096] Reference Figure 6 The diagram shown is a functional module schematic of the image-based fall behavior recognition device 100 of this application.
[0097] The image-based fall recognition device 100 described in this application can be installed in an electronic device. Depending on the functions implemented, the image-based fall recognition device 100 may include an acquisition module 110, a detection module 120, and a recognition module 130. The module described in this application can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and which are stored in the memory of the electronic device.
[0098] In this embodiment, the functions of each module / unit are as follows:
[0099] Acquisition module 110: used to acquire human skeleton points, human contours and / or the clarity of the image to be identified in the image to be identified;
[0100] Detection module 120: used to determine whether there is a complete human image with a resolution greater than a preset value in the image to be identified based on the human skeleton points, the human outline and / or the resolution of the image to be identified;
[0101] The recognition module 130 is used to identify whether the human body in the image to be recognized is in a fallen state if there is a complete human body image in the image to be recognized with a clarity greater than a preset value, based on the width of the human body image, the height of the human body image and / or the position coordinates of the human skeleton points.
[0102] In one embodiment, obtaining human skeleton points, human contours, and / or the sharpness of the image to be identified includes:
[0103] Using a pre-built human detection model, the presence of human information in the image to be identified is detected;
[0104] If human information is present in the image to be identified, the human skeleton points, human contours, and / or the clarity of the image to be identified are obtained according to the pre-trained fusion model.
[0105] In one embodiment, obtaining human skeleton points, human contours, and / or the sharpness of the image to be identified based on a pre-trained fusion model includes:
[0106] The human skeleton points in the image to be identified are obtained using the human skeleton point detection network in the fusion model.
[0107] The human contour in the image to be identified is obtained by using the human contour detection network in the fusion model.
[0108] The image sharpness detection network in the fusion model is used to obtain the sharpness of the image to be identified.
[0109] In one embodiment, determining whether a complete human image with a resolution greater than a preset value exists in the image to be identified based on the human skeleton points, the human outline, and / or the clarity of the image to be identified includes:
[0110] Determine whether the human skeleton points in the image to be identified are complete;
[0111] If so, determine whether the human body outline in the image to be identified is complete;
[0112] If so, determine whether the clarity of the image to be identified is greater than a preset value;
[0113] If so, then the image to be identified contains a complete human image with a clarity greater than a preset value.
[0114] In one embodiment, identifying whether the human body in the image to be identified is in a fallen state based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points includes:
[0115] Determine whether the width of the human body image is greater than the height of the human body image shown;
[0116] If the width of the human body image is greater than the height of the human body image, the system identifies whether the human body in the image to be identified is in a fallen state based on the ratio of the width to the height of the human body image and the position coordinates of the human skeleton points.
[0117] If the width of the human body image is less than or equal to the height of the human body image, the system identifies whether the human body in the image to be identified is in a fallen state based on a pre-trained state detection model.
[0118] In one embodiment, identifying whether the human body in the image to be identified is in a fallen state based on the ratio of the width to the height of the human body image and the position coordinates of the human skeleton points includes:
[0119] Using the position coordinates of the human skeleton points, calculate the first ordinate value of the center point between the left and right hip points of the human body, and calculate the second ordinate value of the center point between the left and right knee points of the human body.
[0120] Based on the coordinate information of the left and right shoulder points of the human body, and the coordinate information of the left and right hip points of the human body, calculate the height and width of the upper body of the human body.
[0121] Using the position coordinates of the human skeleton points, calculate the first distance in the horizontal direction and the second distance in the vertical direction from the head to the ankle of the human body.
[0122] Based on the relationship between the upper body height and the upper body width, the relationship between the ratio of the width to the height of the human body image and a preset threshold, the relationship between the first ordinate value and the second ordinate value, and the relationship between the first distance and the second distance, it is determined whether the human body in the image to be identified is in a fallen state.
[0123] In one embodiment, calculating the upper body height and upper body width of the human body based on the coordinate information of the left and right shoulder points and the coordinate information of the left and right hip points includes:
[0124] Calculate the first difference between the x-coordinate of the left shoulder point and the x-coordinate of the right shoulder point, and calculate the second difference between the x-coordinate of the left hip point and the x-coordinate of the right hip point;
[0125] Calculate the third difference between the ordinate value of the left hip point and the ordinate value of the left shoulder point, and calculate the fourth difference between the ordinate value of the right hip point and the ordinate value of the right shoulder point;
[0126] Calculate the first average of the absolute values of the first difference and the second difference, and use the first average as the upper body width;
[0127] Calculate the second average of the absolute values of the third difference and the fourth difference, and use the second average as the upper body height.
[0128] In one embodiment, calculating the first distance in the horizontal direction and the second distance in the vertical direction from the head to the ankle of the human body using the position coordinates of the human skeleton points includes:
[0129] Using the position coordinates of the human skeleton points, calculate the third abscissa and third ordinate of the center point between the left and right ankle points of the human body;
[0130] Obtain the fourth abscissa and fourth ordinate values of the nose point of the human body from the position coordinates of the human skeleton points;
[0131] The absolute value of the subtraction of the third horizontal coordinate value from the fourth horizontal coordinate value is taken as the first distance;
[0132] The absolute value of the difference between the fourth and third ordinate values is taken as the second distance.
[0133] In one embodiment, identifying whether the human body in the image to be identified is in a fallen state based on the relationship between the upper body height and the upper body width, the relationship between the ratio of the width to the height of the human body image and a preset threshold, the relationship between the first ordinate value and the second ordinate value, and the relationship between the first distance and the second distance includes:
[0134] If the height of the upper body is greater than the width of the upper body, and the first ordinate value is greater than or equal to the second ordinate value, and the first distance is less than the second distance, and the ratio is greater than the preset threshold, then the human body in the human image is in a fallen state.
[0135] Otherwise, the human body in the human body image is not in a fallen state.
[0136] Reference Figure 7 The diagram shown is a schematic diagram of a preferred embodiment of the electronic device 1 of this application.
[0137] The electronic device 1 includes, but is not limited to, a memory 11, a processor 12, a display 13, and a communication interface 14. The electronic device 1 can connect to a network via the communication interface 14. The network can be an intranet, the Internet, a Global System for Mobile communication (GSM), a Wideband Code Division Multiple Access (WCDMA) network, a 4G network, a 5G network, Bluetooth, Wi-Fi, a voice communication network, or other wireless or wired networks.
[0138] The memory 11 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the hard disk or memory of the electronic device 1. In other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped with the electronic device 1. Of course, the memory 11 may include both the internal storage unit and its external storage device of the electronic device 1. In this embodiment, the memory 11 is typically used to store the operating system and various computer programs installed on the electronic device 1, such as the program code of the image-based fall behavior recognition program 10. In addition, the memory 11 can also be used to temporarily store various types of data that have been output or will be output.
[0139] In some embodiments, processor 12 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. Processor 12 is typically used to control the overall operation of the electronic device 1, such as performing data interaction or communication-related control and processing. In this embodiment, processor 12 is used to run program code stored in memory 11 or process data, such as running program code for an image-based fall behavior recognition program 10.
[0140] The display 13 may be referred to as a display screen or display unit. In some embodiments, the display 13 may be an LED display, a liquid crystal display, a touch liquid crystal display, or an organic light-emitting diode (OLED) touch screen, etc. The display 13 is used to display information processed in the electronic device 1 and to display a visual working interface.
[0141] The communication interface 14 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface), which is typically used to establish a communication connection between the electronic device 1 and other electronic devices.
[0142] Figure 7Only an electronic device 1 with components 11-14 and an image-based fall behavior recognition program 10 is shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0143] In the above embodiments, when the processor 12 executes the image-based fall behavior recognition program 10 stored in the memory 11, it can perform the following steps:
[0144] Obtain human skeleton points, human contours, and / or the clarity of the image to be identified;
[0145] Based on the human skeleton points, the human outline, and / or the clarity of the image to be identified, determine whether there is a complete human image in the image to be identified with a clarity greater than a preset value.
[0146] If so, based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points, identify whether the human body in the image to be identified is in a fallen state.
[0147] The storage device can be the memory 11 of the electronic device 1, or it can be other storage devices that are communicatively connected to the electronic device 1.
[0148] For a detailed explanation of the above steps, please refer to the above. Figure 6 Functional block diagram of an embodiment of an image-based fall behavior recognition device 100 and... Figure 1 Description of a flowchart of an embodiment of an image-based fall behavior recognition method.
[0149] Furthermore, this application embodiment also proposes a computer-readable storage medium, which can be non-volatile or volatile. The computer-readable storage medium can be any one or any combination of several of the following: hard disk, multimedia card, SD card, flash memory card, SMC, read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, etc. The computer-readable storage medium includes a data storage area and a program storage area. The program storage area stores an image-based fall behavior recognition program 10. When the image-based fall behavior recognition program 10 is executed by a processor, it performs the following operations:
[0150] Obtain human skeleton points, human contours, and / or the clarity of the image to be identified;
[0151] Based on the human skeleton points, the human outline, and / or the clarity of the image to be identified, determine whether there is a complete human image in the image to be identified with a clarity greater than a preset value.
[0152] If so, based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points, identify whether the human body in the image to be identified is in a fallen state.
[0153] The specific implementation of the computer-readable storage medium in this application is largely the same as the specific implementation of the image-based fall behavior recognition method described above, and will not be repeated here.
[0154] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware simulation platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, electronic device, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0156] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for recognizing fall behavior based on images, characterized in that, The method includes: Obtain human skeleton points, human contours, and / or the clarity of the image to be identified; Based on the human skeleton points, the human outline, and / or the clarity of the image to be identified, determine whether there is a complete human image in the image to be identified with a clarity greater than a preset value. If so, based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points, identify whether the human body in the image to be identified is in a fallen state; The step of identifying whether the human body in the image to be identified is in a fallen state based on the width of the human body image, the height of the human body image, and / or the position coordinates of the human skeleton points includes: Determine whether the width of the human body image is greater than the height of the human body image shown; If the width of the human body image is greater than the height of the human body image, the first ordinate value of the center point between the left and right hip points of the human body is calculated using the position coordinates of the human body skeleton points, and the second ordinate value of the center point between the left and right knee points of the human body is calculated. Based on the coordinate information of the left and right shoulder points of the human body, and the coordinate information of the left and right hip points of the human body, calculate the height and width of the upper body of the human body. Using the position coordinates of the human skeleton points, calculate the first distance in the horizontal direction and the second distance in the vertical direction from the head to the ankle of the human body. If the height of the upper body is greater than the width of the upper body, and the first ordinate value is greater than or equal to the second ordinate value, and the first distance is less than the second distance, and the ratio of the width of the human body image to the height of the human body image is greater than a preset threshold, then the human body in the human body image is in a fallen state. Otherwise, the human body in the image is not in a fallen state; If the width of the human body image is less than or equal to the height of the human body image, the system identifies whether the human body in the image to be identified is in a fallen state based on a pre-trained state detection model.
2. The image-based fall behavior recognition method as described in claim 1, characterized in that, The acquisition of human skeleton points, human contours, and / or the sharpness of the image to be identified includes: Using a pre-built human detection model, the presence of human information in the image to be identified is detected; If human information is present in the image to be identified, the human skeleton points, human contours, and / or the clarity of the image to be identified are obtained according to the pre-trained fusion model.
3. The image-based fall behavior recognition method as described in claim 2, characterized in that, The step of obtaining human skeleton points, human contours, and / or the sharpness of the image to be identified based on a pre-trained fusion model includes: The human skeleton points in the image to be identified are obtained using the human skeleton point detection network in the fusion model. The human contour in the image to be identified is obtained by using the human contour detection network in the fusion model. The image sharpness detection network in the fusion model is used to obtain the sharpness of the image to be identified.
4. The image-based fall behavior recognition method as described in claim 1 or 3, characterized in that, The step of determining whether a complete human image with a resolution greater than a preset value exists in the image to be identified based on the human skeleton points, the human outline, and / or the clarity of the image to be identified includes: Determine whether the human skeleton points in the image to be identified are complete; If so, determine whether the human body outline in the image to be identified is complete; If so, determine whether the clarity of the image to be identified is greater than a preset value; If so, then the image to be identified contains a complete human image with a clarity greater than a preset value.
5. The image-based fall behavior recognition method as described in claim 1, characterized in that, The step of calculating the upper body height and upper body width of the human body based on the coordinate information of the left and right shoulder points and the coordinate information of the left and right hip points includes: Calculate the first difference between the x-coordinate of the left shoulder point and the x-coordinate of the right shoulder point, and calculate the second difference between the x-coordinate of the left hip point and the x-coordinate of the right hip point; Calculate the third difference between the ordinate value of the left hip point and the ordinate value of the left shoulder point, and calculate the fourth difference between the ordinate value of the right hip point and the ordinate value of the right shoulder point; Calculate the first average of the absolute values of the first difference and the second difference, and use the first average as the upper body width; Calculate the second average of the absolute values of the third difference and the fourth difference, and use the second average as the upper body height.
6. The image-based fall behavior recognition method as described in claim 1, characterized in that, The step of calculating the first distance in the horizontal direction and the second distance in the vertical direction from the head to the ankle of the human body using the position coordinates of the human skeleton points includes: Using the position coordinates of the human skeleton points, calculate the third abscissa and third ordinate of the center point between the left and right ankle points of the human body; Obtain the fourth abscissa and fourth ordinate values of the nose point of the human body from the position coordinates of the human skeleton points; The absolute value of the result obtained by subtracting the third horizontal coordinate value from the fourth horizontal coordinate value is taken as the first distance; The absolute value of the difference between the fourth and third ordinate values is taken as the second distance.
7. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the image-based fall behavior recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Body fall detection method
CN106503643A
Pedestrian falling detection method based on Gaussian mixture model and neural network
CN110991274A
Human body fall recognition method and device, terminal equipment, and storage medium
CN113673316A
Video-based face image optimization method and device, equipment and storage medium
CN114037942A