Human body key point detection method and related device

Through a monocular camera, the 3D key points of the human body are detected, and the posture matching and angle correction technology are used to solve the problems of high cost and complexity of multi-camera detection, achieving low-cost and high-accuracy 3D key point detection, improving the accuracy and user experience of posture detection.

CN115147339BActive Publication Date: 2025-08-08HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110351266.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2025-08-08
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

In the prior art, the method of detecting 3D key points in a human body through multiple cameras is expensive and has high computational complexity, resulting in inaccurate detection of 3D key points position information.

Method used

A monocular camera is used to detect the 3D key points of the human body. By detecting whether the user's posture matches the preset posture, the angle between the human body model and the image plane is calculated for correction, reducing the error of the 3D key point position information of the perspective deformation of the image.

Benefits of technology

It reduces cost and calculation complexity, improves the detection accuracy of 3D key point position information, improves the accuracy of posture detection, and improves the user's experience in fitness or somatosensory games.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147339B_ABST
    Figure CN115147339B_ABST
Patent Text Reader

Abstract

The present application provides a method for detecting key points of a human body and a related device. The method can be applied to electronic devices equipped with one or more cameras. The electronic device can identify the 3D key points of a user in an image captured by the camera. The electronic device can detect whether the user's posture matches a preset posture based on the above 3D key points. Using a set of 3D key points determined when the user's posture matches the preset posture, the electronic device can calculate the angle between the human body model determined by this set of 3D key points and the image plane. The electronic device can use the angle to correct the position information of the 3D key points. This method can save costs, reduce the error caused by image perspective deformation to the position information of 3D key points, and improve the accuracy of the detection of the position information of 3D key points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a method for detecting key points of a human body and related devices. Background Art

[0002] Human keypoint detection is the foundation of many computer vision tasks. By detecting the three-dimensional (3D) keypoints of the human body, posture detection, action classification, smart fitness, and somatosensory gaming can be achieved.

[0003] Electronic devices can capture images through cameras and identify two-dimensional (2D) key points of the human body in the images. Based on the 2D key points, electronic devices can use technologies such as deep learning to estimate the 3D key points of the human body. In the above method of estimating 3D key points using 2D key points, due to factors such as differences in user height and distance between the user and the camera, the images captured by the camera will have varying degrees of perspective distortion. Image perspective distortion can cause errors in the detected 2D key points. Therefore, the 3D key points estimated using the above 2D key points will also have errors.

[0004] Currently, electronic devices use multiple cameras to detect the user's distance from the cameras to determine the degree of perspective distortion of a person in an image. Furthermore, electronic devices can correct 3D key points based on the degree of perspective distortion of the person in the image. However, this method requires multiple cameras and the design of their placement to improve detection accuracy, which is not only costly but also computationally complex in determining 3D key points. Summary of the Invention

[0005] The present application provides a method for detecting key points of a human body and a related device. The method can be applied to electronic devices equipped with one or more cameras. The electronic device can identify the 3D key points of a user in an image captured by the camera. The electronic device can detect whether the user's posture matches a preset posture based on the above 3D key points. Using a set of 3D key points determined when the user's posture matches the preset posture, the electronic device can calculate the angle between the human body model determined by this set of 3D key points and the image plane. The electronic device can use the angle to correct the position information of the 3D key points. This method can save costs, reduce the error caused by image perspective deformation to the position information of 3D key points, and improve the accuracy of the detection of the position information of 3D key points.

[0006] In a first aspect, the present application provides a method for detecting key points of a human body, which can be applied to an electronic device comprising one or more cameras. In this method, the electronic device can obtain a first image of a first user through a camera. The electronic device can determine a first set of 3D key points of the first user based on the first image. The electronic device can determine whether multiple 3D key points in the first set of 3D key points meet a first condition. If the first condition is met, the electronic device can determine a first compensation angle based on the multiple 3D key points. The first set of 3D key points is rotationally corrected using the first compensation angle.

[0007] In conjunction with the first aspect, in some embodiments, if the first condition is not met, the electronic device may perform rotation correction on the first set of 3D key points using a second compensation angle. The second compensation angle is determined based on the second set of 3D key points. The second set of 3D key points is the most recent set of 3D key points that meet the first condition before acquiring the first image.

[0008] In conjunction with the first aspect, in some embodiments, a method for an electronic device to obtain a first image of a first user via a camera may include: the electronic device may determine a first moment based on first multimedia information, where the first moment is the moment when the first multimedia information indicates that the user performed an action that satisfies a first condition. The electronic device may obtain the first image of the first user via the camera within a first time period starting at the first moment.

[0009] The first time period may be the time period during which the first multimedia message instructs the user to perform an action that satisfies the first condition. For example, if the time required to complete the action that satisfies the first condition is 1 second, then the first time period is 1 second starting from the first moment. Optionally, the first time period may be a fixed time period.

[0010] In conjunction with the first aspect, in some embodiments, if the 3D key points corresponding to the action indicated by the multimedia information at the second moment do not meet the first condition, the electronic device may use a third compensation angle to perform rotation correction on the 3D key points determined based on images captured within a second time period starting from the second moment. The third compensation angle is determined based on a third group of 3D key points. The third group of 3D key points is the most recent group of 3D key points that meet the first condition before the second moment.

[0011] The second time period may be a time period during which the first multimedia message instructs the user to perform an action that does not satisfy the first condition. For example, if the time required to complete the action that satisfies the first condition is 1 second, then the second time period is 1 second starting from the second moment. Optionally, the second time period may be a fixed time period.

[0012] In combination with the first aspect, in some embodiments, the method for the electronic device to determine whether multiple 3D key points in the first group of 3D key points meet the first condition can be that the electronic device can determine whether multiple 3D key points in the first group of 3D key points match the 3D key points corresponding to the first action, and the first action is at least one upright action of the upper body and legs.

[0013] If the first action is an upper body upright action, the first compensation angle may be the angle between the straight line between the neck point and the chest and abdomen points in the first group of 3D key points and the image plane.

[0014] If the first action is a leg-standing action, the first compensation angle is the angle between the straight line where any two 3D key points among the hip point, knee point, and ankle point in the first group of 3D key points are located and the image plane.

[0015] In conjunction with the first aspect, in some embodiments, the multiple 3D key points in the first group of 3D key points include hip points, knee points, and ankle points. The method for the electronic device to determine whether the multiple 3D key points in the first group of 3D key points meet the first condition can be that the electronic device can calculate a first angle between the straight line containing the left hip point and the left knee point, and the straight line containing the left knee point and the left foot point in the first group of 3D key points, and a second angle between the straight line containing the right hip point and the right knee point, and the straight line containing the right knee point and the right foot point in the first group of 3D key points. The electronic device can determine whether the multiple 3D key points in the first group of 3D key points meet the first condition by detecting whether the difference between the first angle and 180° is less than the first difference, and whether the difference between the second angle and 180° is less than the first difference.

[0016] The situation where multiple 3D key points in the first group of 3D key points meet the first condition includes: the difference between the first angle and 180° is less than the first difference and / or the difference between the second angle and 180° is less than the first difference.

[0017] It can be seen from the above method that when it is detected that the user's 3D key points meet the first condition, the electronic device can use the 3D key points determined based on the image to determine the degree of perspective deformation of the character in the image. When the above-mentioned user's 3D key points meet the first condition, the electronic device can detect that the user is standing with his upper body upright and / or his legs upright. In other words, when the user's upper body is upright and / or his legs are upright, the electronic device can use the 3D key points determined based on the image to determine the degree of perspective deformation of the character in the image. Furthermore, the electronic device can correct the position information of the 3D key points determined based on the image, reduce the error caused by the image perspective deformation to the position information of the 3D key points, and improve the accuracy of the detection of the position information of the 3D key points. This method only requires one camera, which not only saves costs, but also has a low computational complexity for key point detection.

[0018] In addition, each time a user's 3D key points are detected to meet the first condition, if the angle between the straight line containing the 3D key points of the upper body and / or the 3D key points of the legs and the image plane can be used to correct the position information of the 3D key points, the electronic device can update the compensation angle. The updated compensation angle can more accurately reflect the degree of perspective deformation of the person in the image captured by the camera at the user's current position. This can reduce the impact of changes in the position between the user and the camera on the correction of the position information of the 3D key points, thereby improving the accuracy of the position information of the 3D key points after correction. In other words, the electronic device can use the updated compensation angle to correct the position information of the 3D key points. The human body model determined by the corrected 3D key points can more accurately reflect the user's posture, improving the accuracy of posture detection. In this way, in fitness or somatosensory games, the electronic device can more accurately determine whether the user's posture is correct and whether the amplitude of the user's movement meets the requirements, etc., so that the user has a better experience when playing fitness or somatosensory games.

[0019] In conjunction with the first aspect, in some embodiments, the second compensation angle is the angle between the line containing the upper body 3D key points and / or the leg 3D key points in the second group of 3D key points and the image plane. The third compensation angle is the angle between the line containing the upper body 3D key points and / or the leg 3D key points in the third group of 3D key points and the image plane.

[0020] In conjunction with the first aspect, in some embodiments, the second compensation angle is the sum of the change in the first compensation angle and a third angle; the third angle is the angle between the line containing the upper body 3D key points and / or the leg 3D key points in the second set of 3D key points and the image plane; the change in the first compensation angle is determined based on a first height H1, a first distance Y1, a second height H2, and a second distance Y2, where H1 and Y1 are the height of the first user and the distance between the first user and the camera when the first image was captured, respectively; and H2 and Y2 are the height of the first user and the distance between the first user and the camera when the second image was captured, respectively; the key points of the first user in the first image are the first set of 3D key points, and the key points of the first user in the second image are the second set of 3D key points. In the time period between the capture of the second image and the capture of the first image, the greater the decrease in the height of the first user, the greater the increase in the second compensation angle relative to the third angle; the greater the decrease in the distance between the first user and the camera, the greater the increase in the second compensation angle relative to the third angle; and the smaller H1, the greater the increase in the second compensation angle relative to the third angle.

[0021] The third compensation angle is the sum of the change in the second compensation angle and the fourth angle. The fourth angle is the angle between the line containing the upper body 3D key points and / or the leg 3D key points in the third set of 3D key points and the image plane. The change in the second compensation angle is determined based on the third height H3, the third distance Y3, the fourth height H4, and the fourth distance Y4, where H3 and Y3 are the height of the first user and the distance between the first user and the camera during the second time period, respectively. H4 and Y4 are the height of the first user and the distance between the first user and the camera when the third image was captured, respectively. The key points of the first user in the third image are the third set of 3D key points. From the time the first user's height decreases between the capture of the second image and the second time period, the third compensation angle increases more compared to the fourth angle. The greater the decrease in the distance between the first user and the camera, the greater the increase in the third compensation angle compared to the fourth angle. The smaller H3 is, the greater the increase in the third compensation angle compared to the fourth angle.

[0022] The above-mentioned first compensation angle transformation amount and second compensation angle change amount are both determined by a first model, and the first model is obtained by training multiple groups of training samples. A group of training samples includes: the change in the lower edge of the human body's position in the image between the image with an earlier acquisition time and the image with a later acquisition time, the change in the height of the human body in the image, the height of the human body in the image with a later acquisition time, and the change in the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the two groups of 3D key points determined based on the image with an earlier acquisition time and the image with a later acquisition time and the image plane; multiple 3D key points in the two groups of 3D key points all meet the first condition.

[0023] It can be seen from the above method that even if the user's 3D key points do not meet the first condition, the electronic device can determine the degree of perspective deformation of the character in the image in real time and correct the position information of the 3D key points, thereby improving the accuracy of 3D key point detection.

[0024] In conjunction with the first aspect, in some embodiments, if it is determined that the angle between the line containing the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the image plane is less than a fifth angle, the electronic device may determine the first compensation angle based on the multiple 3D key points. If it is determined that the angle between the line containing the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the plane image is greater than the fifth angle, the electronic device may perform rotation correction on the first group of 3D key points using the second compensation angle.

[0025] The fifth angle can be set based on empirical values, which is not limited in the present embodiment.

[0026] It can be seen from the above method that the electronic device can determine whether the angle obtained by the above calculation can be used as a compensation angle for correcting the position information of the 3D key point by setting the fifth angle, thereby avoiding the compensation angle being an impossible value due to errors in the calculation of the above angle, and improving the accuracy of detecting 3D key points.

[0027] The third compensation angle is the sum of the second compensation angle change and the fourth angle; the fourth angle is the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in the third group of 3D key points and the image plane; the second compensation angle change is determined based on the third height H3, the third distance Y3, the fourth height H4, and the fourth distance Y4, where H3 and Y3 are the height of the first user and the distance between the first user and the camera during the second time period, respectively, and H4 and Y4 are the height of the first user and the distance between the first user and the camera when the third image is captured, respectively; the key points of the first user in the third image are the third group of 3D key points;

[0028] Among them, in the time period from the capture of the second image to the second time period, the more the height of the first user decreases, the more the third compensation angle increases compared with the fourth angle; the more the distance between the first user and the camera decreases, the more the third compensation angle increases compared with the fourth angle; the smaller H3 is, the more the third compensation angle increases compared with the fourth angle.

[0029] Determining that an angle between a straight line including the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the image plane is smaller than a fifth angle;

[0030] If the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in the first set of 3D key points and the plane image is greater than the fifth angle, the first set of 3D key points is rotated and corrected using the second compensation angle.

[0031] In a second aspect, the present application provides an electronic device, which includes a camera, a display screen, a memory, and a processor, wherein: the camera can be used to capture images, the memory can be used to store computer programs, and the processor can be used to call the computer program, so that the electronic device executes any possible implementation method of the above-mentioned first aspect.

[0032] In a third aspect, the present application provides a computer storage medium comprising instructions, which, when executed on an electronic device, enables the electronic device to execute any possible implementation of the first aspect.

[0033] In a fourth aspect, an embodiment of the present application provides a chip, which is applied to an electronic device. The chip includes one or more processors, which are used to call computer instructions to enable the electronic device to execute any possible implementation method of the above-mentioned first aspect.

[0034] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when the computer program product is run on a device, enables the electronic device to execute any possible implementation method of the first aspect.

[0035] It is understandable that the electronic device provided in the second aspect, the computer storage medium provided in the third aspect, the chip provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to execute the methods provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a position distribution map of key points of a human body provided in an embodiment of the present application;

[0037] Figure 2A 2D coordinate system schematic diagram for determining position information of 2D key points provided in an embodiment of the present application;

[0038] Figure 2B This is a schematic diagram of a 3D coordinate system for determining position information of 3D key points provided in an embodiment of the present application;

[0039] Figure 3 This is a schematic diagram of a scene for detecting key points of a human body provided in an embodiment of the present application;

[0040] Figure 4A and Figure 4B is a schematic diagram of 3D key points detected by the electronic device 100 provided in an embodiment of the present application;

[0041] Figure 5 1 is a schematic structural diagram of an electronic device 100 provided in an embodiment of the present application;

[0042] Figures 6A to 6D Schematic diagram of some scenes of human key point detection provided by the embodiments of the present application;

[0043] Figure 7A and Figure 7B is a schematic diagram of 3D key points detected by the electronic device 100 provided in an embodiment of the present application;

[0044] Figure 8 This is a flow chart of a method for detecting key points of a human body provided in an embodiment of the present application;

[0045] Figures 9A to 9C Schematic diagram of some scenes of human key point detection provided by the embodiments of the present application;

[0046] Figure 10This is a flow chart of another method for detecting key points of a human body provided in an embodiment of the present application;

[0047] Figure 11 This is a flow chart of another method for detecting key points of a human body provided in an embodiment of the present application;

[0048] Figure 12 This is another position distribution map of key points on the human body provided in an embodiment of the present application;

[0049] Figure 13 This is a flowchart of another method for detecting key points of the human body provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] The following is a clear and detailed description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0051] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0052] Figure 1 The position distribution diagram of the key points of the human body is shown as an example. Figure 1 As shown, the key points of the human body may include: head point, neck point, left shoulder point, right shoulder point, right elbow point, left elbow point, right hand point, left hand point, right hip point, left hip point, middle point between left and right hips, right knee point, left knee point, right ankle point, and left ankle point. The key points are not limited to the above key points. Other key points may also be included in the embodiments of the present application, which are not specifically limited here.

[0053] The concepts of 2D key points and 3D key points involved in the embodiments of the present application are described in detail below.

[0054] 1. 2D key points

[0055] The 2D key points in the embodiments of the present application may represent key points distributed on a 2D plane. The electronic device may capture an image of the user through a camera and identify the 2D key points of the user in the image. The 2D plane may be an image plane on which the image captured by the camera is located. The electronic device may specifically identify the 2D key points of the user in the image by determining the position information of the key points of the user on the 2D plane. The position information of the 2D key points may be represented by two-dimensional coordinates in the 2D plane.

[0056] In one possible implementation, the position information of each 2D key point may be the position information of one of the aforementioned 2D key points with one 2D key point as a reference point. For example, with the midpoint between the left and right hips as the reference point, the position information of the midpoint between the left and right hips on the 2D plane may be the coordinates (0, 0). The electronic device may then determine the position information of the other 2D key points based on the relative positions of the other 2D key points and the midpoint between the left and right hips.

[0057] In another possible implementation, the electronic device may establish a Figure 2A The 2D coordinate system x_i-y_i is shown. The 2D coordinate system can have a vertex in the image captured by the camera as its origin, the horizontal direction of the object in the image as the direction of the x_i axis of the 2D coordinate system, and the vertical direction of the object in the image as the direction of the y_i axis of the 2D coordinate system. The position information of each 2D key point can be the two-dimensional coordinate of the user's key point in the 2D coordinate system.

[0058] The embodiment of the present application does not limit the method for determining the position information of the above-mentioned 2D key points.

[0059] The electronic device can identify a set of 2D key points from a frame of image captured by the camera. This set of 2D key points may include Figure 1 A set of 2D key points can be used to define a human body model on a 2D plane.

[0060] 2. 3D key points

[0061] The 3D key points in the embodiments of the present application may represent key points distributed in a 3D space. Based on 2D key points, the electronic device may estimate the user's 3D key points using technologies such as deep learning. The above-mentioned 3D space may be the 3D space where the camera of the electronic device is located. The electronic device determines the user's 3D key points by specifically determining the position information of the user's key points in the 3D space. The position information of the 3D key points may be represented by three-dimensional coordinates in the 3D space. Compared to 2D key points, the position information of the 3D key points includes the depth information of the user's key points. That is, the position information of the 3D key points may reflect the distance of the user's key points relative to the camera.

[0062] In one possible implementation, the position information of each 3D key point may be the position information of one of the aforementioned 3D key points with one 3D key point as a reference point. For example, with the midpoint between the left and right hips as the reference point, the position information of the midpoint between the left and right hips in the 3D control may be the coordinates (0, 0, 0). The electronic device may then determine the position information of other 3D key points based on their relative positions to the midpoint between the left and right hips.

[0063] In another possible implementation, the electronic device can establish a 3D space where the camera is located. Figure 2B The 3D coordinate system xyz is shown. The 3D coordinate system can be based on the optical center of the camera as the origin, the direction of the camera optical axis (i.e., the direction perpendicular to the image plane) as the z-axis, and the direction of the image plane as the z-axis. Figure 2A The directions of the x_i axis and the y_i axis in the 2D coordinate system are respectively the directions of the x axis and the y axis of the 3D coordinate system. The position information of each 3D key point can be the three-dimensional coordinates of the key point of the user in the 3D coordinate system.

[0064] Figure 2B The 3D coordinate system shown is a right-handed coordinate system. Wherein -x can represent the negative direction of the x-axis. The embodiment of the present application does not limit the method for establishing the above-mentioned 3D coordinate system. For example, the above-mentioned 3D coordinate system can also be a left-handed coordinate system. The electronic device can determine the three-dimensional coordinates of the user's key points in the left-handed coordinate system.

[0065] The method of estimating the position information of 3D key points using deep learning is not limited to that of deep learning. The electronic device can determine the position information of 3D key points using other methods. The specific implementation process of the electronic device estimating the position information of 3D key points based on 2D key points can refer to the implementation of the existing technology, and this embodiment of the application will not be described in detail.

[0066] The electronic device may determine a set of 3D key points using a set of 2D key points. Alternatively, the electronic device may determine a set of 3D key points using multiple sets of 2D key points determined from multiple consecutive frames of images. The above set of 3D key points may include Figure 1 A set of 3D key points can be used to define a human body model in 3D space.

[0067] In some embodiments, the electronic device can perform 3D key point detection on a single frame or multiple consecutive frames of images, and determine a set of 3D key points from the single frame or multiple consecutive frames of images. This means that the electronic device does not need to first determine the 2D key points of the image and then estimate the user's 3D key points based on the 2D key points. The embodiments of this application do not limit the specific method for the electronic device to perform 3D key point detection.

[0068] In the subsequent embodiments of this application, the human body key point detection method provided by this application is specifically introduced by taking the method in which an electronic device first determines the 2D key points of an image and then estimates the 3D key points of the user based on the 2D key points as an example.

[0069] Perspective distortion of people in images captured by the camera can affect the accuracy of 3D keypoint detection. The following details the impact of perspective distortion on 3D keypoint detection.

[0070] Figure 3 The schematic diagram exemplarily shows a scenario in which the electronic device 100 realizes intelligent fitness by detecting 3D key points of the human body.

[0071] like Figure 3 As shown, the electronic device 100 may include a camera 193. The electronic device 100 may capture images through the camera 193. The images captured by the camera 193 may include images of the user during exercise. The electronic device 100 may display the images captured by the camera 193 on the user interface 210. The present embodiment of the application does not limit the content displayed on the user interface 210.

[0072] The electronic device 100 can identify 2D key points of a user in an image captured by the camera 193. Based on these 2D key points, the electronic device 100 can estimate a set of 3D key points of the user. A set of 3D key points can be used to determine a human body model of the user in 3D space. The human body model determined from the 3D key points can reflect the user's posture. The more accurate the position information of the 3D key points, the more accurately the human body model determined from the 3D key points can reflect the user's posture.

[0073] Due to differences in the pitch angle, field of view, and height of the camera on different products. During the process of the camera capturing images, the distance between the user and the camera may change. Affected by the above factors, the image captured by the camera will undergo perspective deformation. If the electronic device 100 uses the image with perspective deformation to determine the user's 2D key points, then the aspect ratio of the human body model determined by the above 2D key points will be different from the aspect ratio of the user's actual body. The electronic device 100 uses the above 2D key points to determine the user's 3D key points. There will be errors in the position information of the above 3D key points. The human body model determined by the above 3D key points is difficult to accurately reflect the user's posture.

[0074] Figure 4A and Figure 4B A group of 3D key points determined by the electronic device 100 when the user is in a standing posture is exemplarily shown from different directions.

[0075] Figure 4A The 3D coordinate system shown can be the aforementioned Figure 2B The 3D coordinate system xyz is shown. Figure 4A A set of 3D key points of the user are shown from the direction of the z-axis. This orientation is equivalent to observing the user's posture from behind the user. Figure 4A It can be seen that the posture of the user reflected by the human body model determined by this set of 3D key points is a standing posture.

[0076] Figure 4B The 3D coordinate system shown can be the aforementioned Figure 2B The 3D coordinate system xyz is shown. z can represent the positive direction of the z-axis. The positive direction of the z-axis is the direction in which the camera points to the object being photographed. Figure 4B A set of 3D key points of the user is shown from the negative direction of the x-axis. This orientation is equivalent to observing the user's posture from the side of the user. Figure 4B It can be seen that the human body model determined by this set of 3D key points leans forward in the direction of the camera 193. When the user is in a standing position, the human body model determined by a set of 3D key points should be perpendicular to the plane x-0-z. Due to the perspective deformation of the image, the human body model determined by a set of 3D key points will have the following Figure 4B The greater the degree to which a group of 3D human models lean forward toward the direction of the camera 193 , the greater the degree of perspective deformation of the human figures in the image captured by the camera 193 .

[0077] This application provides a method for detecting key points on a human body, applicable to electronic devices equipped with a monocular camera. The electronic device can detect 3D key points on a human body using images captured by the monocular camera. The monocular camera can be camera 193 in the aforementioned electronic device 100. This method can reduce costs and improve the accuracy of detecting 3D key point location information.

[0078] Specifically, the electronic device can identify the 2D key points of the user in the image captured by the camera and estimate the 3D key points of the user based on the 2D key points. The electronic device can determine whether the user's legs are upright based on the above 3D key points. Using a set of 3D key points determined when the user's legs are upright, the electronic device can calculate the compensation angle between the human body model determined by this set of 3D key points and the image plane. Furthermore, the electronic device can use the compensation angle to correct the position information of the 3D key points, reduce the error caused by the perspective deformation of the image to the position information of the 3D key points, and improve the accuracy of the detection of the position information of the 3D key points.

[0079] The above-mentioned image plane is the plane where the image captured by the camera is located (that is, the plane perpendicular to the optical axis of the camera).

[0080] When the user's legs are upright, the angle between the line containing the user's legs and the image plane should be zero or close to zero. Upon detecting that the user's legs are upright, the electronic device can determine the degree of perspective distortion of the person in the image based on the angle between the line containing the user's leg 3D key points and the image plane, and correct the position information of the 3D key points. Compared to detecting 3D key points using images captured by multiple cameras, this method only requires a single camera, resulting in lower costs and computational complexity.

[0081] When a user is exercising or playing a somatosensory game, the area of activity generally does not change significantly. In other words, the degree of perspective distortion of the character in the image captured by the camera does not change significantly during the exercise or somatosensory game. During the exercise or somatosensory game, the electronic device can detect the angle between the line containing the 3D key points of the user's legs and the image plane when the user's legs are upright at a certain time during the exercise or somatosensory game, and use this angle as a compensation angle to correct the position information of the 3D key points determined during the exercise or somatosensory game.

[0082] Optionally, during a fitness or somatosensory game, if the electronic device detects the user's legs standing upright multiple times, the electronic device can update the compensation angle. That is, the electronic device can use the angle between the line containing the 3D key points of the user's legs when the user's legs were most recently detected standing upright and the image plane as the compensation angle to correct the position information of the 3D key points determined in subsequent stages of the fitness or somatosensory game. The above method can reduce the impact of changes in the position between the user and the camera on the correction of the 3D key point position information, thereby improving the accuracy of the corrected 3D key point position information.

[0083] For example, in a somatosensory game scenario, the electronic device can determine the user's 3D key points based on the image captured by the camera, and determine the virtual character in the somatosensory game based on the above 3D key points. The electronic device can present the above virtual character on the user interface. The above virtual character can reflect the user's posture. For example, the user jumps forward, and the above virtual character also jumps forward. Before the position information of the 3D key points is corrected, there is an error in the 3D key point position information due to the perspective deformation of the image. Then there is a difference between the posture of the above virtual character and the actual posture of the user. For example, the user is actually in a standing posture, and the above virtual character may be in a posture leaning forward. Then the action actually completed by the user may be standard, but the electronic device 100 will determine that the user's action is not standard and indicate that the user has failed to pass the level. This will affect the user's gaming experience.

[0084] After the position information of the 3D key points is corrected, the posture of the virtual character is more closely aligned with the user's actual posture. In this way, during fitness or somatosensory games, the electronic device 100 can more accurately determine whether the user's posture is correct and whether the amplitude of the user's movements meets the requirements, thereby providing the user with a better experience when playing fitness or somatosensory games.

[0085] Figure 5 A structural diagram of an electronic device 100 provided in an embodiment of the present application is exemplarily shown.

[0086] like Figure 5 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, etc.

[0087] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0088] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0089] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0090] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0091] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality devices.

[0092] The charging management module 140 is used to receive charging input from a charger. While the charging management module 140 is charging the battery 142 , it can also provide power to the electronic device 100 through the power management module 141 .

[0093] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.

[0094] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0095] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0096] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0097] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0098] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0099] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0100] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.

[0101] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0102] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0103] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0104] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0105] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0106] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.

[0107] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0108] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0109] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0110] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0111] The speaker 170A, also called a "horn", is used to convert audio electrical signals into sound signals.

[0112] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals.

[0113] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals.

[0114] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0115] The sensor module 180 may include a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0116] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.

[0117] Motor 191 can generate vibration prompts. Motor 191 can be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. Motor 191 can also correspond to different vibration feedback effects for touch operations acting on different areas of the display screen 194. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0118] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0119] Not limited to Figure 5 The electronic device 100 may include more or fewer components than those shown. The electronic device 100 in the embodiment of the present application may be a television, a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a portable multimedia player (PMP), a dedicated media player, an AR (augmented reality) / VR (virtual reality) device, or other types of electronic devices. The embodiment of the present application does not limit the specific type of the electronic device 100.

[0120] The following describes a method for detecting key points of a human body provided in an embodiment of the present application, based on a scenario in which a user follows a fitness course in the electronic device 100 and the electronic device 100 detects the user's 3D key points.

[0121] The electronic device 100 may store multiple fitness courses. Optionally, the electronic device 100 may obtain multiple fitness courses from a cloud server. A fitness course typically includes multiple movements, wherein a preset rest period may be provided between two consecutive movements, and any two movements may be the same or different. Fitness courses may be recommended by the electronic device based on the user's historical fitness data, or may be selected by the user based on actual needs. Fitness courses may be played locally or online. No specific limitations are imposed here.

[0122] A fitness course may include multiple sub-courses, each of which may include one or more consecutive exercises. These sub-courses may be divided based on exercise type, exercise purpose, exercised body parts, etc. This is not specifically limited here.

[0123] For example, a fitness course includes three sub-courses, wherein the first sub-course is warm-up exercise, the second sub-course is main exercise, and the third sub-course is stretching exercise, and any of the three sub-courses includes one or more consecutive movements.

[0124] In the embodiment of the present application, the fitness course may include one or more types of content in the form of video, animation, voice, text, etc., which are not specifically limited here.

[0125] Phase 1: Start a fitness class.

[0126] like Figure 6A As shown, Figure 6A The user interface 61 on the electronic device 100 for displaying the application programs installed on the electronic device 100 is exemplarily shown.

[0127] The user interface 61 may include an icon 611 for the fitness application, as well as icons for other applications (such as email, gallery, and music). Any application icon can be used to respond to a user operation, such as a touch operation, to enable the electronic device 100 to launch the application corresponding to the icon. The user interface 61 may also include more or less content, which is not limited in this embodiment of the application.

[0128] In response to a user operation on the fitness icon 611, the electronic device 100 may display the following information: Figure 6B The fitness course interface 62 shown. The fitness course interface 62 may include an application title bar 621, a function bar 622, and a display area 623. Among them:

[0129] The application title bar 621 may be used to indicate that the current page is used to display the setting interface of the electronic device 100. The application title bar 621 may be in the form of text information "Smart Fitness", an icon, or other forms.

[0130] The function bar 622 may include: user center control, course recommendation control, fat burning area control, body shaping area control, and body shaping area control. The function bar 622 is not limited to the above controls, and may include more or fewer controls.

[0131] In response to a user operation on any control in the function bar 622 , the electronic device 100 may display the content indicated by the control in the display area 623 .

[0132] For example, in response to a user operation on a user center control, the electronic device 100 may display the interface content of the user's personal center in the display area 623. In response to a user operation on a course recommendation control, the electronic device 100 may display one or more recommended fitness courses in the display area 623. Figure 6B As shown, display area 623 displays course covers for multiple recommended courses. The course covers may include the course category, duration, and name of the corresponding fitness course. In response to a user operation on any course cover, electronic device 100 may open the fitness course corresponding to the course cover and display the exercise content of the fitness course.

[0133] The embodiments of the present application do not limit the above-mentioned user operations. For example, the user can also control the electronic device 100 to execute corresponding instructions (such as starting a fitness application, starting a fitness course, etc.) through a remote control.

[0134] The fitness course interface 62 may also include more or less content, which is not limited in this embodiment of the present application.

[0135] In response to a user action on the course cover of any fitness course (e.g., a fitness course titled "Full Body Fat Burning Beginner"), electronic device 100 can start the fitness course. During the fitness course playback, electronic device 100 needs to capture images using its camera. Before the fitness course playback begins, electronic device 100 can prompt the user that the camera is about to be turned on.

[0136] Phase 2: Determine the target user and the initial compensation angle, and use the initial compensation angle to correct the position information of the 3D key points.

[0137] The target user may represent a user who needs the electronic device 100 to detect 3D key points and record motion data while the electronic device 100 is playing a fitness course.

[0138] In some embodiments, as Figure 6C As shown, the electronic device 100 can collect facial information of the user to identify the target user who is exercising. Identifying the target user facilitates the electronic device 100 to track the user for whom key point detection is required and accurately obtain the target user's motion data. This prevents other users other than the target user from interfering with the electronic device 100's detection of the target user's key points and acquisition of the target user's motion data within the camera's shooting range.

[0139] During the process of determining the target user, the electronic device 100 may also determine an initial compensation angle. Before the initial compensation angle is updated, the electronic device 100 may use the initial compensation angle to correct the position information of the 3D key points determined based on the 2D key points after the fitness course starts playing.

[0140] For example, the electronic device 100 may display a target user determination interface 63. The target user determination interface 63 may include a prompt 631 and a user image 632. The prompt may be used to prompt the user to perform operations related to determining the target user and the initial compensation angle. The prompt may be a text message such as "Please stand in the area where you are exercising and face the camera." This embodiment of the application does not limit the form and specific content of the prompt. The user image 632 is an image of the target user captured by the camera.

[0141] The user images captured by the camera can be used by the electronic device 100 to identify target users for key point detection and motion data recording during a fitness course. The electronic device 100 can utilize a target tracking algorithm to track the target user. The implementation of the target tracking algorithm can be referenced to the specific implementation of target tracking algorithms in the prior art and will not be elaborated upon here.

[0142] The above prompts the user to maintain a standing posture in the area where exercise is performed, which can facilitate the electronic device 100 to determine the initial compensation angle.

[0143] Specifically, following prompt 631, the user maintains a standing posture within the area where the user will be exercising while the electronic device 100 identifies the target user. The electronic device 100 may capture one or more frames of images containing the user's motion posture and identify one or more sets of 2D key points of the user from the one or more frames of images. Furthermore, based on the one or more sets of 2D key points, the electronic device 100 may estimate a set of 3D key points of the user. Based on the set of 3D key points of the user, the electronic device 100 may determine whether the user's legs are upright. If the user's legs are upright, the electronic device 100 may calculate the angle between the line containing the 3D key points of the user's legs and the image plane. If the angle is less than a preset angle, the electronic device 100 may determine the angle as the initial compensation angle. Otherwise, the electronic device 100 may use a default angle as the initial compensation angle. The default angle may be pre-stored in the electronic device 100. The above default angle can be used as a universal compensation angle to correct the position information of 3D key points, reducing the error caused by the perspective deformation of the characters in the image. The above default angle can be set based on empirical values. The embodiment of this application does not limit the value of the above default angle.

[0144] The electronic device 100 determines whether the angle between the line on which the 3D key point is located and the image plane is less than a preset angle, and if the angle is less than the preset angle, determines the angle as the initial compensation angle. This can avoid an initial compensation angle that is impossible due to an error in calculating the angle between the line on which the 3D key point is located and the image plane. The preset angle can be, for example, 45°. The embodiment of the present application does not limit the value of the preset angle.

[0145] Once the target user and initial compensation angle are determined, electronic device 100 can play the fitness course. During the fitness course, electronic device 100 can capture images using a camera and identify the user's 2D key points in the images. Based on these 2D key points, electronic device 100 can determine the user's 3D key points. Electronic device 100 can use the initial compensation angle to correct the position information of these 3D key points.

[0146] The electronic device 100 can perform 2D keypoint detection on each frame of image. For each frame of image, the electronic device 100 can identify a set of 2D keypoints corresponding to the frame of image. The electronic device 100 can use this set of 2D keypoints to determine a set of 3D keypoints. This set of 3D keypoints can be used to determine a human body model. This human body model can reflect the user's posture in the frame of image.

[0147] Optionally, the electronic device 100 can perform 2D key point detection on each frame of image. For multiple consecutive frames of image, the electronic device 100 can identify multiple groups of 2D key points corresponding to the multiple consecutive frames of image. The electronic device 100 can use these multiple groups of 2D key points to determine a set of 3D key points. This set of 3D key points can be used to determine a human body model. This human body model can reflect the posture of the user in the multiple consecutive frames of image. Compared with the 3D key points determined using the 2D key points identified from a single frame of image, the 3D key points determined using the 2D key points identified from multiple consecutive frames of image can more accurately reflect the user's posture.

[0148] The embodiment of the present application does not limit the method of determining the target user. For example, the electronic device 100 can determine the target user by detecting whether the user's action matches a preset action. The above preset action can be bending the arms upward. Before starting to play the fitness course, the electronic device 100 can Figure 6C The target user determination interface 63 shown displays an example of an action of bending arms upward, as well as a prompt for prompting the user to do so. In addition, the electronic device 100 can perform human posture detection based on images captured by the camera. The electronic device 100 can determine a user whose posture is the same as that of bending arms upward as the target user. When the user performs the above-mentioned target user determination process and instructs the user to complete the action (such as bending arms upward), the legs are upright. Then, during the above-mentioned target user determination process, the electronic device 100 can also determine the initial compensation angle according to the method of the above-mentioned embodiment.

[0149] The embodiment of the present application does not limit the type of the above-mentioned preset actions.

[0150] In some embodiments, when playing a fitness course, the electronic device 100 can obtain the movements included in the fitness course. When playing a fitness course, the electronic device 100 can detect users whose movements match the movements in the fitness course to determine the target user. In addition, the electronic device 100 can use the above-mentioned default angle as the initial compensation angle and use the default angle to correct the position information of the 3D key points determined by the electronic device 100. That is, the electronic device 100 can not use Figure 6C The method shown is used to determine the target user and the initial compensation angle.

[0151] Alternatively, if the electronic device 100 detects the user's legs standing upright at the start of a fitness class, or within a period of time after the start of a fitness class (e.g., within 3 seconds), the electronic device 100 may determine the initial compensation angle based on the 3D key points used when detecting the user's legs standing upright. The initial compensation angle is the angle between the line containing the leg 3D key points and the image plane.

[0152] Phase 3: Update the compensation angle and use the updated compensation angle to correct the position information of the 3D key points.

[0153] When playing a fitness course, the electronic device 100 can obtain the movements included in the fitness course. The electronic device 100 can determine whether the movements in the fitness course include leg upright movements (such as standing movements, front standing arms raising movements, etc.).

[0154] like Figure 6D As shown, during the playing of the fitness course, the electronic device 100 can display Figure 6D The user interface 64 shown. The user interface 64 may include a fitness course window 641 and a user fitness window 642.

[0155] The fitness course window 641 may be used to display the specific content of the fitness course, such as an image of a coach performing movements in the fitness course.

[0156] The user fitness window 642 may be used to display the image of the target user captured by the camera in real time.

[0157] The embodiment of the present application does not limit the distribution of the fitness course window 641 and the user fitness window 642 on the user interface 64. The fitness course window 641 and the user fitness window 642 may also contain more content, which is not limited in the embodiment of the present application.

[0158] When the fitness course is played to time t1, the action that the fitness course instructs the user to complete is the action of straightening the legs, such as standing. Figure 6D As shown, the action of the coach in the fitness course window 641 is a standing action. The user can perform the standing action according to the instructions of the coach in the fitness course window 641. The electronic device 100 can use the image captured by the camera around time t1 to obtain the 3D key points. The electronic device 100 can determine whether the user's legs are upright based on the 3D key points. If the user's legs are upright, the electronic device 100 can calculate the angle between the straight line where the 3D key points of the user's legs are located and the image plane. If the angle is smaller than the preset angle in the aforementioned embodiment, the electronic device 100 can determine the angle as the compensation angle. The electronic device 100 can use the compensation angle to update the previous compensation angle, and use the updated compensation angle to correct the position information of the 3D key points.

[0159] When the user is instructed to stand up straight for the first time during a fitness class, the compensation angle calculated by the electronic device 100 is updated to the initial compensation angle in the aforementioned embodiment. When the user is instructed to stand up straight for the second or subsequent times during a fitness class, the compensation angle calculated by the electronic device 100 is updated to the compensation angle calculated the last time the electronic device 100 detected the user standing up straight.

[0160] As can be seen from the above embodiment, the electronic device 100 can detect whether the user's legs are straight when instructed to do so during a fitness class. If the user's legs are straight, the electronic device 100 can update the compensation angle used to correct the position information of the 3D key points. The updated compensation angle is the angle between the line containing the 3D key points of the user's legs and the image plane, calculated by the electronic device 100 when the user's legs are straight. The updated compensation angle can more accurately reflect the degree of perspective distortion of the person in the image captured by the camera at the user's current position. In other words, the electronic device 100 can use the updated compensation angle to correct the position information of the 3D key points. The human body model determined by the corrected 3D key points can more accurately reflect the user's posture, improving the accuracy of posture detection. In this way, during fitness or somatosensory gaming, the electronic device 100 can more accurately determine whether the user's posture is correct and whether the amplitude of the user's movements meets the requirements, thereby providing the user with a better experience while playing fitness or somatosensory gaming.

[0161] In some embodiments, the electronic device 100 can determine a set of 3D key points during the playback of a fitness course through a set of 2D key points identified from a frame of image or multiple sets of 2D key points identified from multiple consecutive frames of image. The set of 3D key points can determine a human body model. The electronic device 100 can determine whether the user's legs are upright through the set of 3D key points. If the user's legs are upright, the electronic device 100 can calculate the angle between the straight line where the 3D key points of the user's legs are located and the image plane. If the angle is smaller than the preset angle in the aforementioned embodiment, the electronic device 100 can update the compensation angle. The updated compensation angle is the above-mentioned angle. The electronic device 100 can use the updated compensation angle to correct the position information of the 3D key points.

[0162] That is, each time the electronic device 100 determines a set of 3D key points, it may use the set of 3D key points to determine whether the compensation angle for correcting the position information of the 3D key points can be updated.

[0163] Not limited to the above-mentioned smart fitness scenario, the human body key point detection method provided in the embodiment of the present application can also be applied to other scenarios where posture detection is achieved by detecting 3D key points.

[0164] The following describes a method provided by an embodiment of the present application for determining whether a user's legs are upright and determining a compensation angle when the user's legs are upright.

[0165] In some embodiments, the user's legs being upright may mean both of the user's legs are upright.

[0166] When the user's legs are both upright, the angle between the thigh and calf of each leg should be 180° or a value close to 180°. The electronic device 100 can use the 3D key points determined based on the 2D key points to determine whether the angle between the user's thigh and calf is close to 180°, thereby determining whether the user's legs are upright.

[0167] Figure 7A Schematic diagram of a set of 3D key points is shown as an example. Figure 7A As shown, the angle between the straight line where the user's left thigh is located (i.e., the straight line where the left hip point and the left knee point are located) and the straight line where the left calf is located (i.e., the straight line where the left knee point and the left ankle point are located) is β1. The angle between the straight line where the user's right thigh is located (i.e., the straight line where the right hip point and the right knee point are located) and the straight line where the right calf is located (i.e., the straight line where the right knee point and the right ankle point are located) is β2. The electronic device 100 can calculate the difference between β1 and 180° and the difference between β2 and 180°. If the difference between β1 and 180° and the difference between β2 and 180° are both less than the preset difference, the electronic device 100 can determine that the user's legs are upright. The above-mentioned preset difference is a value close to 0. The above-mentioned preset difference can be set based on experience. The embodiment of the present application does not limit the size of the above-mentioned preset difference.

[0168] Furthermore, when it is determined that the user's legs are upright, the electronic device 100 can calculate the angle α1 between the line connecting the left hip point and the left ankle point and the image plane, and the angle α2 between the line connecting the right hip point and the right ankle point and the image plane. The electronic device 100 can calculate the average of the angles α1 and α2 to obtain the following: Figure 7B If the angle α is smaller than the preset angle in the above embodiment, the electronic device 100 may determine the angle α as a compensation angle and use the compensation angle to correct the position information of the 3D key point.

[0169] Optionally, when it is determined that the user's legs are upright, the electronic device 100 may determine that the straight line connecting the left hip point and the left ankle point is Figure 7A The first projection line on the y-0-z plane of the 3D coordinate system shown, and the line connecting the right hip point and the right ankle point are located in Figure 7A The electronic device 100 can calculate the first projection line and the second projection line on the y-0-z plane of the 3D coordinate system. Figure 7AThe angle α3 of the positive direction of the y-axis of the 3D coordinate system shown, and the angle between the second projection line and Figure 7A The electronic device 100 can calculate the average of the angle α3 and the angle α4 to obtain the following: Figure 7B If the angle α is smaller than the preset angle in the above embodiment, the electronic device 100 may determine the angle α as a compensation angle and use the compensation angle to correct the position information of the 3D key point.

[0170] In some embodiments, the user's legs standing upright may refer to the user standing upright on one leg. The electronic device 100 may determine whether any one of the user's legs is standing upright according to the method of the aforementioned embodiment. If it is determined that one of the user's legs is standing upright, the electronic device 100 may determine that the user's legs are standing upright. Furthermore, the electronic device 100 may calculate the angle between the straight line connecting the hip point and the ankle point on the upright leg and the image plane. If the angle is smaller than the preset angle in the aforementioned embodiment, the electronic device 100 may determine the angle as a compensation angle and use the compensation angle to correct the position information of the 3D key points.

[0171] The following describes a method for correcting the position information of 3D key points using a compensation angle α, provided in an embodiment of the present application.

[0172] The electronic device 100 can correct the position information of the 3D key points according to the following formula (1):

[0173]

[0174] in, It is the position information of a 3D key point before correction. It is estimated by the electronic device 100 based on the 2D key points. It is the corrected position information of the above 3D key point.

[0175] Figure 8 The flowchart of a method for detecting key points of a human body provided in an embodiment of the present application is exemplified.

[0176] like Figure 8 As shown, the method may include steps S101 to S108.

[0177] S101. The electronic device 100 may capture images through a camera.

[0178] In some embodiments, the camera and the electronic device 100 may be integrated. Figure 3 As shown, the electronic device 100 may include a camera 193 .

[0179] In other embodiments, the camera and the electronic device 100 may be two separate devices. A communication connection is established between the camera and the electronic device 100. The camera may send the captured image to the electronic device 100.

[0180] In some embodiments, in response to a user operation of starting a fitness course, the electronic device 100 may turn on the camera or send an instruction to the camera to turn on the camera. Figure 6C As shown, in response to a user operation acting on the determination control 624A, the electronic device 100 can turn on the camera.

[0181] The embodiment of the present application does not limit the time when the camera is turned on. For example, the camera can also be turned on before the electronic device 100 receives a user operation to start a fitness course.

[0182] S102 : The electronic device 100 may determine, based on the m frames of images captured by the camera, m groups of 2D key points corresponding to the m frames of images.

[0183] In some embodiments, the electronic device 100 may perform 2D key point detection on images captured by a camera during a fitness course, that is, the m frames of images are captured by the camera during the fitness course.

[0184] The value of m is a positive integer. The electronic device 100 can identify the 2D key points of a group of users in a frame of image. Among them, the electronic device 100 can use a deep learning method (such as the openpose algorithm, etc.) to perform 2D key point detection on the image. The embodiment of the present application does not limit the method for the electronic device 100 to identify 2D key points. The implementation process of the electronic device 100 to identify 2D key points can refer to the implementation process in the prior art and will not be repeated here.

[0185] S103 : The electronic device 100 may estimate a set of 3D key points corresponding to the user and the m frames of images based on the m sets of 2D key points.

[0186] If m is 1, then for each frame of image, the electronic device 100 can identify a set of 2D key points from the frame of image. Furthermore, the electronic device 100 can use the set of 2D key points corresponding to the frame of image to estimate a set of 3D key points corresponding to the frame of image.

[0187] If m is an integer greater than 1, the m frames represent a continuous multi-frame image captured by the camera. The electronic device 100 can estimate a set of 3D key points based on the continuous multi-frame image. Specifically, the electronic device 100 can identify multiple sets of 2D key points from the continuous multi-frame image. Using the multiple sets of 2D key points corresponding to the multi-frame image, the electronic device 100 can estimate a set of 3D key points corresponding to the multi-frame image.

[0188] The embodiment of the present application does not limit the method of estimating 3D key points using 2D key points. The implementation process of the electronic device 100 estimating 3D key points can refer to the implementation process in the prior art and will not be described in detail here.

[0189] S104 : The electronic device 100 may use a set of 3D key points corresponding to the m frames of images to detect whether the user's legs are upright.

[0190] When a set of 3D key points is determined, the electronic device 100 can detect whether the user's legs are upright based on the position information of the set of 3D key points. The method for the electronic device 100 to detect whether the user's legs are upright can refer to the aforementioned Figure 7A The embodiment shown.

[0191] If it is detected that the user's legs are standing upright, the electronic device 100 may execute step S105.

[0192] If it is detected that the user's legs are not upright, the electronic device 100 may execute step S108.

[0193] S105. If it is detected that the user's legs are standing upright, the electronic device 100 may determine the angle between the straight line where the 3D key points of the legs in the above set of 3D key points are located and the image plane, where the image plane is a plane perpendicular to the optical axis of the camera.

[0194] When the user's legs are upright, the angle between the line containing the user's legs and the image plane should be close to 0. The electronic device 100 can determine the degree of perspective distortion of the person in the image captured by the camera by detecting the angle between the line containing the 3D key points of the user's legs and the image plane, thereby correcting the position information of the 3D key points.

[0195] When a set of 3D key points corresponding to the m frames of image are obtained and the user's legs are detected to be upright, the electronic device 100 can determine the angle between the straight line where the 3D key points of the legs in the set of 3D key points are located and the image plane. The method for calculating the angle by the electronic device 100 can refer to the above method. Figure 7A The description of the illustrated embodiment will not be repeated here.

[0196] S106: The electronic device 100 may determine whether the angle is smaller than a preset angle.

[0197] If it is determined that the angle is smaller than the preset angle, the electronic device 100 may execute step S107 to update the compensation angle stored in the electronic device 100 .

[0198] If it is determined that the angle is greater than or equal to the preset angle, the electronic device 100 may execute step S108 to correct the position information of a group of 3D key points corresponding to the m frames of images using the compensation angle stored therein.

[0199] There may be errors in the calculation of the above-mentioned angle by the electronic device 100, resulting in the calculated angle being too large. For example, the above-mentioned angle calculated by the electronic device 100 is 70°. However, when the user's legs are upright, the effect of the perspective deformation of the character in the image on the 3D key points obviously cannot cause the angle between the straight line where the user's leg 3D key point is located and the image plane to reach 70°. The electronic device 100 sets a preset angle to determine whether the angle calculated above can be used as a compensation angle to correct the position information of the 3D key point, thereby avoiding the compensation angle being an impossible value due to the error in calculating the above-mentioned angle, and improving the accuracy of detecting 3D key points.

[0200] The value of the preset angle can be set based on experience. For example, the value of the preset angle can be 45°. The embodiment of the present application does not limit the value of the preset angle.

[0201] S107: If the angle is smaller than the preset angle, the electronic device 100 may use the angle to update the compensation angle stored in the electronic device 100. After the update, the compensation angle stored in the electronic device 100 is the angle.

[0202] If the angle is smaller than the preset angle, the preset angle can be used as a compensation angle for correcting the position information of the 3D key point. The electronic device 100 can use the angle to update its stored compensation angle.

[0203] S108 : The electronic device 100 may use the compensation angle stored in the electronic device 100 to correct the position information of the group of 3D key points.

[0204] The method for the electronic device 100 to correct the position information of the 3D key points using the compensation angle may refer to the aforementioned embodiment.

[0205] In some embodiments, the compensation angle stored in the electronic device 100 may be the compensation angle updated in step S107. The updated compensation angle is the angle between the straight line containing the 3D key points of the leg in the set of 3D key points corresponding to the m frames of image and the image plane.

[0206] In other embodiments, the compensation angle stored in the electronic device 100 may be the initial compensation angle in the aforementioned embodiment. As can be seen from the aforementioned embodiment, the initial compensation angle may be a default angle pre-stored in the electronic device 100. Alternatively, the initial compensation angle may be Figure 6C The figure shows the compensation angle calculated when the user's legs are detected standing upright before a fitness class starts.

[0207] That is, during the time period T1 between the electronic device 100 starting to play the fitness class and the camera capturing the m frames of imagery, the electronic device 100 does not update the initial compensation angle. The electronic device 100 can use the initial compensation angle to correct the position information of each group of 3D key points determined during the time period T1.

[0208] In other embodiments, the compensation angle stored in the electronic device 100 may be the compensation angle calculated when the user's legs are detected to be standing upright for the last time within the above-mentioned time period T1.

[0209] That is to say, within the above-mentioned time period T1, the electronic device 100 detects the user's legs standing upright at least once. When the user's legs are detected to be standing upright, the electronic device 100 can calculate the angle between the 3D key points of the legs and the image plane based on a set of 3D key points determined when the user's legs are standing upright this time. If the angle can be used to correct the position information of the 3D key points, the electronic device 100 can update the compensation angle and use the angle as a new compensation angle to replace the compensation angle calculated the last time the user's legs were standing upright. The electronic device 100 can use the stored compensation angle to correct the position information of the 3D key points until the stored compensation angle is updated. The electronic device 100 can use the updated compensation angle to correct the position information of the 3D key points determined after the compensation angle is updated.

[0210] In some embodiments, when the electronic device 100 plays a fitness course, it determines whether the action in the fitness course includes the action of standing upright. When the fitness course is played to the point where the user is instructed to complete the action of standing upright, the electronic device 100 may execute Figure 8 Step S104 is shown. That is, the electronic device 100 can use the 3D key points determined based on the 2D key points to detect whether the user's legs are upright. If the user's legs are detected to be upright, the electronic device 100 can determine whether the angle between the straight line containing the 3D key points of the legs and the image plane can be used to correct the position information of the 3D key points. If the angle can be used to correct the position information of the 3D key points, the electronic device 100 can use the angle to update the compensation angle.

[0211] When a fitness class is played and the user is instructed to perform an unsteady leg movement, the electronic device 100 can use the currently captured image to determine the user's 2D key points and, based on these 2D key points, determine the user's 3D key points. The electronic device 100 can use the most recently updated compensation angle to correct the 3D key points.

[0212] In other words, when the fitness class instructs the user to straighten their legs, the electronic device 100 can use the 3D key points determined based on the 2D key points to determine whether the user's legs are straight. This eliminates the need for the electronic device 100 to use each set of 3D key points determined based on the 2D key points to determine whether the user's legs are straight and whether to update the compensation angle. This saves computing resources on the electronic device 100.

[0213] The embodiment of the present application does not limit the time when the electronic device 100 updates the compensation angle. In addition to the method in the above embodiment, the electronic device 100 can also regularly or irregularly detect whether the user's legs are upright, and update the compensation angle when the user's legs are detected to be upright.

[0214] Depend on Figure 8 As can be seen from the human body key point detection method shown, when the user's legs are detected to be upright, the electronic device 100 can use the angle between the straight line containing the 3D key points of the legs among the 3D key points determined based on the 2D key points and the image plane to determine the degree of perspective deformation of the person in the image captured by the camera. Furthermore, the electronic device 100 can correct the position information of the 3D key points, reduce the error caused by the image perspective deformation to the position information of the 3D key points, and improve the accuracy of the 3D key point position information detection. This method only requires one camera, which not only saves costs but also reduces the computational complexity of key point detection.

[0215] Additionally, each time the user's legs are detected standing upright, if the angle between the straight line containing the leg's 3D key points and the image plane can be used to correct the position information of the 3D key points, the electronic device 100 can update the compensation angle. The updated compensation angle can more accurately reflect the degree of perspective distortion of the person in the image captured by the camera at the user's current position. This can reduce the impact of changes in the position between the user and the camera on the correction of the 3D key point position information, thereby improving the accuracy of the corrected 3D key point position information. In other words, the electronic device 100 can use the updated compensation angle to correct the position information of the 3D key points. The human body model determined by the corrected 3D key points can more accurately reflect the user's posture, improving the accuracy of posture detection. In this way, during fitness or somatosensory gaming, the electronic device 100 can more accurately determine whether the user's posture is correct and whether the amplitude of the user's movements meets the requirements, thereby providing the user with a better experience when playing fitness or somatosensory gaming.

[0216] The human body key point detection method in the aforementioned embodiment is mainly to determine the degree of perspective deformation of the character in the image captured by the camera when the user's legs are upright. In actual applications, such as when doing fitness or somatosensory games, the user will not keep the legs upright. The user may perform actions such as squatting, lying on the side, etc. When the user's legs are not upright, the electronic device 100 can use the image captured by the camera to determine the 2D key points, and determine one or more groups of 3D key points based on the above 2D key points. According to the aforementioned Figure 8 In the method shown, the electronic device 100 can only use the compensation angle calculated when the user's legs are most recently detected to be upright to correct the position information of one or more groups of 3D key points determined when the user's legs are not upright.

[0217] Although the area where the user exercises is generally not too large, the degree of perspective deformation of the person in the image captured by the camera will not change much, but the change in the relative position between the user and the camera and the change in the user's height in the image captured by the camera will affect the degree of perspective deformation of the person in the image captured by the camera. The compensation angle calculated by the electronic device 100 when the user's legs are upright at the previous moment is difficult to accurately reflect the degree of perspective deformation of the person in the image captured by the camera when the user's legs are not upright at the next moment. The electronic device 100 uses the compensation angle calculated when the user's legs were most recently detected to be upright to correct the position information of one or more groups of 3D key points determined when the user's legs were not upright. After the correction, the position information of the 3D key points is still not accurate enough.

[0218] According to the principle of perspective deformation, when the camera is placed at the same position and angle, the closer the distance between the user and the camera is, the greater the degree of perspective deformation of the person in the image captured by the camera (i.e., the aforementioned Figure 4B The smaller the user height in the image captured by the camera, the greater the degree of perspective deformation of the person in the image captured by the camera (i.e., the aforementioned Figure 4B The larger the compensation angle α shown, the greater the effect. For example, if the user changes from standing to squatting, the user's height in the image captured by the camera decreases, which results in a greater degree of perspective distortion of the person in the image captured by the camera. As the distance between the user and the camera increases, the user's height in the image captured by the camera decreases, which also results in a greater degree of perspective distortion of the person in the image captured by the camera.

[0219] The present application provides a method for detecting key points of a human body, which can determine the degree of perspective deformation of a person in an image in real time without requiring the user to keep his legs upright, and correct the position information of the 3D key points determined based on the 2D key points. Specifically, the electronic device 100 can determine an initial compensation angle. The method for determining the initial compensation angle can refer to the introduction of the aforementioned embodiment and will not be described in detail here. According to the change in the position of the user in two adjacent frames of images, the electronic device 100 can determine the amount of change in the degree of perspective deformation of the person between the two adjacent frames of images. The sum of the amount of change in the degree of perspective deformation of the person in the previous frame of image and the amount of change in the degree of perspective deformation of the person between the two adjacent frames of image is the degree of perspective deformation of the person in the next frame of image. The degree of perspective deformation of the person in the previous frame of image can be based on the initial compensation angle, and the amount of change in the degree of perspective deformation of the person between the two adjacent frames of image before the previous frame of image is accumulated.

[0220] As can be seen from the above method, even if the user's legs are not upright, electronic device 100 can still determine the degree of perspective distortion of the person in the image captured by the camera. Electronic device 100 corrects the position information of 3D key points based on the degree of perspective distortion of the person in the image determined in real time, thereby improving the accuracy of 3D key point detection.

[0221] The following specifically describes a method for determining the degree of change in the perspective deformation of a character between two adjacent frames of image provided by an embodiment of the present application.

[0222] Figure 9A The following is an example of a scene diagram of a user in motion. Figure 9AAs shown, electronic device 100 can capture images via camera 193 and display the images on user interface 910. Between the time camera 193 captures the n-1th frame and the time camera 193 captures the nth frame, the user moves in the direction of camera 193. That is, the distance between the user and camera 193 decreases. n is an integer greater than 1. The user's position in the n-1th frame changes to that in the nth frame.

[0223] Figure 9B and Figure 9C The n-1th frame image and the nth frame image captured by the camera 193 are exemplarily shown.

[0224] The electronic device 100 can perform human body detection on the image and determine the human body rectangular frame of the user in the image. The human body rectangular frame can be used to determine the user's position in the image. The height and width of the human body rectangular frame are adapted to the height and width of the user in the image, respectively. The electronic device 100 can determine the height of the human body rectangular frame in the image coordinate system (i.e., the 2D coordinate system x_i-y_i in the aforementioned embodiment) and the distance of the lower edge from the x_i axis in each frame of the image.

[0225] like Figure 9B As shown, in the n-1th frame image, the distance between the lower edge of the human body rectangle and the x_i axis is y n-1 The height of the human body rectangle is h n-1 .

[0226] like Figure 9C As shown, in the nth frame image, the distance between the lower edge of the human body rectangle and the x_i axis is y n The height of the human body rectangle is h n .

[0227] Between two adjacent frames, the change in the distance Δy between the lower edge of the human body rectangular frame and the x_i axis can reflect the change in the relative position between the user and the camera. The above Δy is the difference between the distance between the lower edge of the human body rectangular frame and the x_i axis in the latter frame and the distance between the lower edge of the human body rectangular frame and the x_i axis in the previous frame. If the user moves closer to the camera, the distance between the lower edge of the human body rectangular frame and the x_i axis becomes smaller, and the degree of perspective deformation of the character in the image becomes larger. Figure 4BThe compensation angle α shown can represent the degree of perspective deformation of the character in the image. In other words, the smaller the distance between the lower edge of the human body rectangular frame and the x_i axis in a frame of image, the larger the compensation angle α used to correct the position information of a set of 3D key points determined based on this frame of image. The larger the absolute value of Δy, the larger the absolute value of Δα. The above Δα is the difference between the latter compensation angle and the previous compensation angle. The above previous compensation angle is used to correct the position information of a set of 3D key points determined based on the previous frame of image in two adjacent frames. The above latter compensation angle is used to correct the position information of a set of 3D key points determined based on the latter frame of image in two adjacent frames.

[0228] If Δy < 0, then Δα > 0. That is, as the user approaches the camera, the degree of perspective distortion of the person in the image increases, and the latter compensation angle is larger than the former. If Δy > 0, then Δα < 0. That is, as the user moves away from the camera, the degree of perspective distortion of the person in the image decreases, and the latter compensation angle is smaller than the former.

[0229] The change in the height of the human frame, Δh, between two adjacent image frames can reflect the change in the user's height in the image. Δh is the difference between the height of the human frame in the subsequent image and the height of the human frame in the previous image. As the height of the human frame decreases, the degree of perspective distortion of the person in the image increases. In other words, the smaller the height of the human frame in a frame, the larger the compensation angle α used to correct the position information of the set of 3D key points determined based on that frame. The larger the absolute value of Δh, the larger the absolute value of Δα.

[0230] If Δh < 0, then Δα > 0. That is, as the height of the user in the image decreases, the degree of perspective distortion increases, and the latter compensation angle is larger than the former. If Δh > 0, then Δα < 0. That is, as the height of the user in the image increases, the degree of perspective distortion decreases, and the latter compensation angle is smaller than the former.

[0231] Between two adjacent image frames, the height h of the human figure's rectangular frame in the latter image can influence the change in the degree of perspective distortion between the two frames. It's understandable that the smaller the user's height in the image, the greater the degree of perspective distortion. Given the same Δh between two adjacent image frames, the smaller the user's height, the greater the change in perspective distortion between the two frames. That is, compared to varying the user's height within a larger range, varying the user's height within a smaller range results in a greater change in perspective distortion.

[0232] The smaller h is, the larger the absolute value of Δα is.

[0233] The electronic device 100 is not limited to determining the change in the relative position between the user and the camera, and the change in the user's height in the image captured by the camera through the above-mentioned human body rectangular frame. The electronic device 100 can also determine the change in the relative position between the user and the camera, the user's height in the image captured by the camera, and the change in the user's height through other methods.

[0234] The electronic device 100 may store a compensation angle determination model. The input of the compensation angle determination model may include Δy and Δh between two adjacent frames of images and the height h of the human body rectangle in the latter frame of the image. The output may be Δα between the two adjacent frames of images.

[0235] The compensation angle determination model may be, for example, a linear model, a nonlinear model, a neural network model, etc. The embodiment of the present application does not limit the type of the compensation angle determination model.

[0236] The compensation angle determination model can be obtained through multiple sets of training samples. These multiple sets of training samples can be determined by using images captured by the camera when the user's legs are upright. A set of training samples may include Δy, Δh between two adjacent frames of images, and the height h and Δα of the human body rectangular frame in the latter frame of the image. Δα in a set of training samples may be the difference between the latter compensation angle and the former compensation angle. The former compensation angle is used to correct the position information of a set of 3D key points determined based on the former frame of the two adjacent frames of images. The latter compensation angle is used to correct the position information of a set of 3D key points determined based on the latter frame of the two adjacent frames of images. The former compensation angle and the latter compensation angle are obtained by the aforementioned Figure 7A The calculation is performed using the method in the embodiment shown.

[0237] The embodiment of the present application does not limit the method for training the above-mentioned compensation angle determination model. For details, reference may be made to the training methods of linear models, neural network models, and the like in the prior art.

[0238] In some embodiments, the electronic device 100 can determine the Δy′, Δh′ between a frame image acquired earlier and a frame image acquired later, as well as the height of the rectangular frame of the human body in the frame image acquired later. By inputting the Δy′, Δh′, and the height of the rectangular frame of the human body in the frame image acquired later into the compensation angle determination model, the electronic device 100 can obtain the compensation angle change between the frame image acquired earlier and the frame image acquired later. This is not limited to determining the compensation angle change between two adjacent frames; the electronic device 100 can also determine the compensation angle change between two frames separated by multiple frames.

[0239] In some embodiments, the electronic device 100 can identify a set of 2D key points from a frame of image and estimate a set of 3D key points corresponding to the frame of image based on the set of 2D key points. That is, each frame of image can correspond to a set of 3D key points. The electronic device 100 can determine the change in the degree of perspective deformation of a character between two adjacent frames of image using the methods described in the aforementioned embodiments. That is, for each frame of image, the electronic device 100 can determine a compensation angle. The electronic device 100 can use the compensation angle corresponding to a frame of image to correct the position information of the set of 3D key points corresponding to the frame of image.

[0240] In some embodiments, the electronic device 100 can identify multiple groups of 2D key points from multiple consecutive frames of images, and estimate a group of 3D key points corresponding to the multiple frames of images based on the multiple groups of 2D key points. For example, the electronic device 100 can use consecutive k frames of images to determine a group of 3D key points. The above k is an integer greater than 1. Based on the change in the degree of perspective deformation of the character between the first k frames of images and the last k frames of images, the electronic device 100 can determine the compensation angle corresponding to the last k frames of images. That is, for each consecutive k frames of images, the electronic device 100 can determine a compensation angle. The electronic device 100 can use the compensation angle corresponding to the consecutive k frames of images to correct the position information of a group of 3D key points corresponding to the consecutive k frames of images.

[0241] When determining the change in the degree of perspective deformation of the person between the first k frames and the last k frames, the electronic device 100 may select one frame from each of the first k frames and the last k frames. The electronic device 100 may calculate the change in the distance between the lower edge of the person's rectangular frame and the x_i axis and the change in the height of the person's rectangular frame between the two selected frames. Based on the compensation angle determination model, the electronic device 100 may determine the change in the degree of perspective deformation of the person between the first k frames and the last k frames.

[0242] In some embodiments, after the electronic device 100 determines a compensation angle corresponding to a group of 3D key points by accumulating the changes in the compensation angles according to the method in the above embodiment, it can smooth the compensation angle and use the smoothed compensation angle to correct the position information of the group of 3D key points. Specifically, the electronic device 100 can obtain multiple compensation angles corresponding to multiple groups of 3D key points before the above group of 3D key points. The electronic device 100 can calculate the weighted average of the compensation angle corresponding to the above group of 3D key points and the multiple compensation angles corresponding to multiple groups of 3D key points before the above group of 3D key points. Among them, the compensation angle corresponding to the 3D key point that is closer to the time of determination of the above group of 3D key points can have a greater weight. The embodiment of the present application does not specifically limit the size of the weight of each compensation angle when calculating the above weighted average. The above weighted average is the smoothed compensation angle.

[0243] Between consecutive images captured by the camera, the distance between the lower edge of the human body rectangular frame and the x_i axis, as well as the height of the human body rectangular frame, in each frame continuously changes. The electronic device 100 performs the aforementioned smoothing on the compensation angle to reduce sudden changes in the calculated compensation angle and improve the accuracy of 3D key point detection.

[0244] Figure 10 The flowchart of another method for detecting key points of a human body provided in an embodiment of the present application is exemplified.

[0245] Here, the electronic device 100 uses each frame image to determine a set of 3D key points as an example. Figure 10 As shown, the method may include steps S201 to S207.

[0246] S201. The electronic device 100 may capture images through a camera.

[0247] Step S201 can refer to the above Figure 8 Step S101 in the method shown.

[0248] S202: The electronic device 100 may determine an initial compensation angle.

[0249] The method for determining the initial compensation angle can refer to the above Figure 6C The description of the illustrated embodiment will not be repeated here.

[0250] S203, the electronic device 100 determines the tth image corresponding to the nth image frame according to the nth image frame captured by the camera. n Group 2D key points and according to the t n A set of 2D key points, estimating the tth corresponding user to the nth frame image n Group 3D key points.

[0251] S204: The electronic device 100 determines the displacement Δy of the lower edge of the human body rectangular frame from the n-1th frame image to the nth frame image based on the n-1th frame image and the nth frame image captured by the camera. n , the change in the height of the human body rectangle Δh n and the height h of the human body rectangle in the nth frame image n .

[0252] The above n is an integer greater than 1. The (n-1)th frame image and the (n)th frame image are any two adjacent frame images in the images captured by the camera.

[0253] The electronic device 100 may perform 2D key point detection on the first frame of image captured by the camera to determine the t1th group of 2D key points corresponding to the first frame of image. Based on the t1th group of 2D key points, the electronic device 100 may estimate the t1th group of 3D key points of the user corresponding to the t1th frame of image. The electronic device 100 may use the initial compensation angle to correct the position information of the t1th group of 3D key points.

[0254] Displacement Δy of the lower edge of the human body rectangle n That is, the change in the distance between the lower edge of the human body rectangle and the x_i axis between the n-1th frame image and the nth frame image.

[0255] The electronic device 100 determines the above Δy n and the above Δh n The method can refer to the above Figure 9B and Figure 9C The description of the illustrated embodiment will not be repeated here.

[0256] S205, the electronic device 100 can be based on the above Δy n , the above Δh n and the above h n , determine from the t n-1 Group 3D key points to tth n The compensation angle change of the group 3D key points, the tth n-1 The set of 3D key points is obtained based on the n-1 frame image captured by the camera.

[0257] The electronic device 100 may use the compensation angle determination model to determine the angle from the tth n-1 Group 3D key points to tth n Group 3D key point compensation angle change Δα n Δα n It can reflect the change in the degree of perspective deformation of the character between the n-1th frame image and the nth frame image.

[0258] The method for the electronic device 100 to obtain the compensation angle determination model may refer to the aforementioned embodiment.

[0259] The electronic device 100 can perform 2D key point detection on the n-1th frame image captured by the camera, and determine the tth key point corresponding to the n-1th frame image. n-1 Group 2D key points. According to the t n-1 The electronic device 100 can estimate the t-th image corresponding to the n-1 frame of the user. n-1 Group 3D key points.

[0260] S206, the electronic device 100 can be based on the t n-1 The compensation angle and the above compensation angle change are used to determine the t n compensation angle, tth n-1 The compensation angle is used to correct the t n-1 The position information of the 3D key points of the group, the tth n-1 The compensation angle is calculated based on the initial compensation angle and the angle from the t1th group of 3D key points to the tth group of 3D key points. n-1 The sum of the compensation angle changes of all two adjacent groups of 3D key points between the groups of 3D key points.

[0261] The electronic device 100 can determine the tth n Compensation angles:

[0262] α n =α n-1 +Δα n (2)

[0263] Among them, α n For the tth n Compensation angle. n-1 For the tth n-1 Compensation angle. α1 is the initial compensation angle mentioned above. From the t1th group of 3D key points to the tth n-1 The compensation angle change between all two adjacent groups of 3D key points. It can reflect the change in the degree of perspective deformation of the character in all two adjacent frames from the 1st frame image to the n-1th frame image.

[0264] In some embodiments, the electronic device 100 can determine the amount of change in the compensation angle of the previous group of 3D key points and the next group of 3D key points. There may be one or more groups of 3D key points between the previous group of 3D key points and the next group of 3D key points. The above-mentioned amount of change in the compensation angle can be determined based on the height of the user, the distance between the user and the camera when the previous frame image is collected, and the height of the user, the distance between the user and the camera when the next frame image is collected. The key points of the user in the previous frame image are the above-mentioned 3D key points in the previous group. The key points of the user in the next frame image are the above-mentioned 3D key points in the next group. The electronic device 100 can determine the compensation angle of the next group of 3D key points based on the compensation angle of the previous group of 3D key points and the above-mentioned compensation angle change. That is, α n =α n-c +Δα n ′. Among them, α n-c It can be the t1th group of 3D key points to the tth group n-c The sum of the compensation angle change of the group key points and the compensation angle of the t1th group 3D key points. Δα n ′ is the compensation angle change between the previous set of 3D key points and the next set of 3D key points.

[0265] S207, the electronic device 100 can use the t n The compensation angle corrects the tth n The position information of the group 3D key points.

[0266] Depend on Figure 10 As can be seen from the method shown, even if the user does not keep his legs upright continuously, the electronic device 100 can still determine the degree of perspective deformation of the character in the image in real time and correct the position information of the 3D key points determined based on the 2D key points. Among them, the electronic device 100 can determine the compensation angle corresponding to each group of 3D key points. The compensation angle corresponding to each group of 3D key points can change with the change of the relative position between the user and the camera and the change of the user's action. A compensation angle corresponding to a group of 3D key points can reflect the degree of perspective deformation of the character in the image used to determine this group of 3D key points. Using a compensation angle corresponding to a group of 3D key points to correct the position information of this group of 3D key points can improve the accuracy of 3D key point detection.

[0267] In the above Figure 10 In the method shown, the compensation angle variation determined by the electronic device 100 using the compensation angle determination model may have a certain error. n When the compensation angle is 1, the electronic device 100 accumulates the 3D key points from the t1th group to the tth group based on the initial compensation angle. nThe compensation angle variation of all adjacent groups of 3D key points between the two groups of 3D key points will be accumulated. n The larger the value of n The error of each compensation angle will also be larger, which will reduce the accuracy of 3D key point detection.

[0268] In some embodiments, based on the above Figure 10 According to the method shown, the electronic device 100 can periodically or irregularly detect whether the user's legs are upright. When the user's legs are detected to be upright, the electronic device 100 can calculate the angle between the straight line where the 3D key points of the legs are located in a group of 3D key points determined based on the 2D key points and the image plane. If the angle is less than the preset angle, the electronic device 100 can determine the angle as the compensation angle corresponding to this group of 3D key points, and use the compensation angle to correct the position information of this group of 3D key points. Furthermore, the electronic device 100 can use the compensation angle corresponding to this group of 3D key points as a basis, according to Figure 10 The method shown determines the amount of change in the compensation angle to determine the subsequent compensation angle corresponding to each group of 3D key points.

[0269] If the user's legs are detected to be upright again, the electronic device 100 can use the angle between the straight line where the leg 3D key points are located in a group of 3D key points determined based on the 2D key points when the user's legs are upright this time and the image plane as a basis to determine the subsequent compensation angles corresponding to each group of 3D key points.

[0270] In this way, the electronic device 100 can reduce the accumulated error when determining the compensation angle by accumulating the compensation angle change. Figure 10 The method shown in the above embodiment can further reduce the error in calculating the degree of perspective deformation of the character in the image and improve the accuracy of 3D key point detection.

[0271] Figure 11 The flowchart of another method for detecting key points of a human body provided in an embodiment of the present application is exemplified.

[0272] Here, the electronic device 100 uses each frame image to determine a set of 3D key points as an example. Figure 11 As shown, the method may include steps S301 to S310. In which:

[0273] S301. The electronic device 100 may capture images through a camera.

[0274] S302: The electronic device 100 can determine the tth frame image corresponding to the nth frame image based on the nth frame image captured by the camera. n Group 2D key points and according to the t nA set of 2D key points, estimating the tth corresponding user to the nth frame image n Group 3D key points.

[0275] S303, the electronic device 100 can use the t n A set of 3D key points is used to detect whether the user's legs are upright.

[0276] S304: If it is detected that the user's legs are upright, the electronic device 100 may determine the t n The angle between the straight line containing the 3D key points of the legs in the group of 3D key points and the image plane, where the image plane is a plane perpendicular to the optical axis of the camera.

[0277] S305: The electronic device 100 may determine whether the angle is smaller than a preset angle.

[0278] The above steps S301 to S305 can refer to the above Figure 8 Steps S101 to S106 in the method shown are not described in detail here.

[0279] S306: If the angle is smaller than the preset angle, the electronic device 100 may determine the angle as the tth angle. n Compensation angle.

[0280] S307, the electronic device 100 can use the t n The compensation angle corrects the tth n The position information of the group 3D key points.

[0281] The above steps S301 to S307 are that when the user's legs are detected to be standing upright, the electronic device 100 can calculate the angle between the straight line where the 3D key points of the legs are located in the 3D key points determined based on the 2D key points and the image plane. The electronic device 100 can use this angle as a compensation angle and use this compensation angle to correct the position information of the 3D key points determined based on the 2D key points when the user's legs are standing upright. In other words, if the user's legs are detected to be standing upright, the electronic device 100 does not need to use the initial compensation angle or the compensation angle determined when the user's legs were most recently detected to be standing upright as a basis to determine the compensation angle corresponding to the 3D key points at this time.

[0282] S308: If the electronic device 100 detects that the user's legs are not upright in the above step S303, or if the electronic device 100 determines in the above step S305 that the angle is greater than or equal to the preset angle, the electronic device 100 can determine the displacement Δy of the lower edge of the human body rectangular frame from the n-1 frame image to the n frame image based on the n-1 frame image and the n frame image captured by the camera. n , the change in the height of the human body rectangle Δh n and the height h of the human body rectangle in the nth frame imagen .

[0283] S309, the electronic device 100 can be based on the above Δy n , the above Δh n and the above h n , determine from the t n-1 Group 3D key points to tth n The compensation angle change of the group 3D key points, the tth n-1 The set of 3D key points is obtained based on the n-1 frame image captured by the camera.

[0284] The above steps S308 and S309 can refer to the above steps respectively. Figure 10 Steps S204 and S205 in the method shown are not described in detail here.

[0285] S310, the electronic device 100 can be configured according to the t n-1 The compensation angle and the above compensation angle change are used to determine the t n compensation angle, tth n-1 The compensation angle is used to correct the t n-1 The position information of the 3D key points of the group, the tth n-1 The compensation angle is the tth n-1 The angle between the straight line where the 3D key points of the legs are located and the image plane, or the angle between the tth p The compensation angle and the p Group 3D key points to tth n-1 The sum of the compensation angle changes of all two adjacent groups of 3D key points between the tth group of 3D key points p The compensation angle is calculated by the electronic device 100 when it detects the user's legs standing upright for the last time before the camera captures the n-1th frame image.

[0286] The electronic device 100 can be calculated according to the above formula (2): n =α n-1 +Δα n To determine. Among them, α n For the tth n Compensation angle. n-1 For the tth n-1 Compensation angle. n-1 The calculation method is the same as above Figure 10 The methods shown are different.

[0287] Specifically, if the electronic device 100 uses the t n-1 The set of 3D key points detects that the user's legs are upright, and the tth n-1 If the angle between the straight line where the 3D key points of the legs are located and the image plane is less than the preset angle, then the tn-1 Compensation angle α n-1 Can be t n-1 The angle between the straight line containing the 3D keypoints of the leg in the group of 3D keypoints and the image plane.

[0288] or, p is a positive integer less than n-1. That is, the p-th frame image is the image captured by the camera before the n-1-th frame image. p For the tth p Compensation angle. p Can be used to correct the t p The position information of the group 3D key points. p It is calculated when the electronic device 100 detects the user's legs standing upright for the last time before the camera captures the n-1th frame image. p For the tth p The angle between the straight line where the 3D key point of the leg in the group of 3D key points is located and the image plane is smaller than the preset angle. From the t p Group 3D key points to tth n-1 The compensation angle change between all two adjacent groups of 3D key points. It can reflect the change in the degree of perspective deformation of the character in all two adjacent frames from the p-th frame image to the n-1-th frame image.

[0289] When the above t n Compensation angle α n , the electronic device 100 may execute step S307, i.e., using α n To correct the t n The position information of the group 3D key points.

[0290] In some embodiments, the electronic device 100 can determine the amount of change in the compensation angle of the previous group of 3D key points and the next group of 3D key points. There may be one or more groups of 3D key points between the previous group of 3D key points and the next group of 3D key points. The amount of change in the compensation angle can be determined based on the height of the user and the distance between the user and the camera when the previous frame of image is collected, and the height of the user and the distance between the user and the camera when the next frame of image is collected. The key points of the user in the previous frame of image are the previous group of 3D key points. The key points of the user in the next frame of image are the next group of 3D key points. The electronic device 100 can determine the compensation angle of the next group of 3D key points based on the compensation angle of the previous group of 3D key points and the amount of change in the compensation angle.

[0291] Optionally, the previous frame image may be captured when the electronic device 100 detects that the user's legs are standing upright. For example, the previous frame image may be captured when the electronic device 100 last detected that the user's legs are standing upright, before the capture time of the subsequent frame image.

[0292] In the above Figure 11 In the method shown, the electronic device 100 may not detect whether the user's legs are upright for each set of 3D key points. Optionally, the electronic device 100 may detect whether the user's legs are upright regularly or irregularly. Alternatively, the electronic device 100 may monitor whether the user's legs are upright when the fitness course is played to the point where the user is instructed to complete the action of standing upright. When the fitness course is played to the point where the user is instructed to complete the action of standing upright, the electronic device 100 may determine the user's 3D key points using the currently captured image. The electronic device 100 may determine the user's 3D key points based on the current set of 3D key points. Figure 11 The compensation angle is determined in steps S308 to S310, and the compensation angle is used to correct the above-mentioned 3D key points.

[0293] Depend on Figure 11 As can be seen from the method shown, when the user's legs are detected to be upright, the electronic device 100 can follow the above Figure 8 Otherwise, the electronic device 100 can use the initial compensation angle or the compensation angle determined when the user's legs are most recently detected to be upright as a basis, and then determine the compensation angle according to the above method. Figure 10 The compensation angle is determined by the method shown. This can reduce the accumulated error when determining the compensation angle by accumulating the compensation angle changes. Figure 8 and Figure 10 The method shown in the above embodiment can further reduce the error in calculating the degree of perspective deformation of the character in the image and improve the accuracy of 3D key point detection.

[0294] The method of determining the compensation angle when detecting that the user's legs are upright in the aforementioned embodiment is not limited to that described above. The electronic device 100 may also determine the compensation angle when detecting that the user's posture matches a preset posture. The preset posture may be a posture with the upper body upright and / or legs upright. In other words, detecting whether the user's legs are upright is one way to detect whether the user's posture matches the preset posture.

[0295] If it is detected that the user's upper body is standing upright, the electronic device 100 may determine the angle between the straight line where the 3D key points of the upper body are located and the image plane as the compensation angle, and use the compensation angle to correct the 3D key points.

[0296] If it is detected that the user's upper body and legs are both upright, the electronic device 100 can determine the angle between the straight line where the upper body 3D key points and / or the leg 3D key points are located and the image plane as the compensation angle, and use the compensation angle to correct the 3D key points.

[0297] Figure 12 Another example of a position distribution diagram of key points of a human body is shown. Figure 12 As shown, the key points of the human body may include the head point, the first neck point, the second neck point, the left shoulder point, the right shoulder point, the right elbow point, the left elbow point, the right hand point, the left hand point, the first chest and abdomen point, the second chest and abdomen point, the third chest and abdomen point, the right hip point, the left hip point, the middle point between the left and right hips, the right knee point, the left knee point, the right ankle point, and the left ankle point. The key points are not limited to the above key points. Other key points may also be included in the embodiments of the present application, which are not specifically limited here.

[0298] Compared to Figure 1 The electronic device 100 can identify more key points of the human body. Figure 12 The position information of each key point in the 2D plane and the position information in the 3D space are determined. Figure 12 The 2D keypoints and 3D keypoints corresponding to each keypoint are shown.

[0299] Figure 13 The flowchart of another method for detecting key points of a human body provided in an embodiment of the present application is exemplified.

[0300] like Figure 13 As shown, the method may include steps S401 to S408.

[0301] S401. The electronic device 100 may capture images through a camera.

[0302] S402 : The electronic device 100 may determine, based on the m frames of images captured by the camera, m groups of 2D key points corresponding to the m frames of images.

[0303] S403 : The electronic device 100 may estimate a set of 3D key points corresponding to the user and the m frames of images based on the m sets of 2D key points.

[0304] The above steps S401 to S403 can refer to the above Figure 8 Steps S101 to S103 in the above are not described in detail here.

[0305] S404: The electronic device 100 may use a set of 3D key points corresponding to the m frames of images to determine whether the user's posture matches a preset posture, where the preset posture is a posture with the upper body and / or legs standing upright.

[0306] In some embodiments, the m frames of images can be any m frames of images captured by the camera. That is, the electronic device 100 can determine whether the user's posture determined by each set of 3D key points matches a preset posture. If the user's posture matches the preset posture, the electronic device 100 can execute the following step S405. If the user's posture does not match the preset posture, the electronic device 100 can execute the following step S408.

[0307] In some embodiments, the above-mentioned preset posture may be a posture included in a fitness course or a somatosensory game. The electronic device 100 may obtain the action that the fitness course or somatosensory game instructs the user to complete. Among them, the electronic device 100 may store the moment when the above-mentioned preset posture is played in the fitness course or somatosensory game and the 3D key points corresponding to the above-mentioned preset posture. At the moment when the fitness course or somatosensory game is played to the above-mentioned preset posture, the electronic device 100 may compare the set of 3D key points corresponding to the above-mentioned m-frame images with the 3D key points corresponding to the above-mentioned preset posture. If the posture of the user indicated by the set of 3D key points corresponding to the above-mentioned m-frame images matches the above-mentioned preset posture, the electronic device 100 may execute the following step S405. The above-mentioned m-frame images are collected by the camera at the moment when the fitness course or somatosensory game is played to the above-mentioned preset posture.

[0308] Optionally, the electronic device 100 may store the moment when the preset posture is played in a fitness course or somatosensory game. When the fitness course or somatosensory game plays to the moment when the preset posture is played, the electronic device 100 may obtain the 3D key points corresponding to the preset posture. Furthermore, the electronic device 100 may compare the set of 3D key points corresponding to the m frames of image with the 3D key points corresponding to the preset posture.

[0309] In some embodiments, the m frames of images are captured by the camera at a time when the user is in a fitness class or somatosensory game that is not in the preset posture. The electronic device 100 can use the compensation angle determined when the user was in the most recent fitness class or somatosensory game that was in the preset posture to correct the 3D key points determined in the m frames of images. In other words, the electronic device 100 does not need to determine whether the user's posture determined by each set of 3D key points matches the preset posture. This can save computing resources of the electronic device 100.

[0310] The preset postures are not limited to those in the fitness courses or somatosensory games. The preset postures may be those in the action library of other applications. The action library may store information indicating the user to perform various actions (such as image information of the action, audio information of the action, etc.).

[0311] In some embodiments, some time before a fitness course or somatosensory game starts playing, or during the fitness course or somatosensory game, the electronic device 100 may instruct the user to complete the action corresponding to the preset posture. That is, the above-mentioned preset posture may not be included in the fitness course or somatosensory game. When applications such as fitness courses or somatosensory games are running, the electronic device 100 may regularly or irregularly instruct the user to complete the action corresponding to the preset posture. When instructing the user to complete the action corresponding to the preset posture, the electronic device 100 may compare the set of 3D key points corresponding to the above-mentioned m-frame images with the set of 3D key points corresponding to the above-mentioned preset posture. If the posture of the user indicated by the set of 3D key points corresponding to the above-mentioned m-frame images matches the above-mentioned preset posture, the electronic device 100 may perform the following step S405. The above-mentioned m-frame images are collected by the camera when the electronic device 100 instructs the user to complete the action corresponding to the preset posture.

[0312] Optionally, when determining whether the user's posture matches a preset posture, the electronic device 100 may use a portion of the set of 3D key points (such as key points of the upper body or key points of the legs) for comparison. For example, if the preset posture is an upper body upright posture, the electronic device 100 may compare the 3D key points of the upper body in the set of 3D key points of the user with the 3D key points of the upper body in the set of 3D key points corresponding to the preset posture. If the position information of the two sets of upper body 3D key points is the same or the difference is less than a threshold, the electronic device 100 may determine that the user's posture matches the preset posture. If the preset posture is a leg upright posture, the electronic device 100 may compare the 3D key points of the legs in the set of 3D key points of the user with the 3D key points of the legs in the set of 3D key points corresponding to the preset posture. If the position information of the two sets of leg 3D key points is the same or the difference is less than a threshold, the electronic device 100 may determine that the user's posture matches the preset posture.

[0313] S405. The electronic device 100 may determine the angle between the straight line where some of the 3D key points in the above set of 3D key points are located and the image plane, where the some of the 3D key points include the 3D key points of the upper body and / or the 3D key points of the legs, and the image plane is a plane perpendicular to the optical axis of the camera.

[0314] In some embodiments, the above-mentioned preset posture is a posture of the upper body standing upright. The electronic device 100 detects that the user's posture matches the above-mentioned preset posture, which can be represented by the neck point (such as Figure 12 The first and second neck points shown) and chest and abdomen points (as shown Figure 12The first chest and abdomen point, the second chest and abdomen point, and the third chest and abdomen point shown in the above set of 3D key points are approximately on a straight line. The electronic device 100 can determine the angle between the straight line on which any two 3D key points in the above set of 3D key points lie, namely, the first neck point, the second neck point, the first chest and abdomen point, the second chest and abdomen point, and the third chest and abdomen point, and the image plane. For example, the electronic device 100 determines the angle between the straight line on which the first neck point and the third chest and abdomen point in the above set of 3D key points lie, and the image plane.

[0315] Optionally, when it is detected that the user's posture matches the above-mentioned preset posture, the electronic device 100 can calculate the average of the angles between the image plane and multiple straight lines where any two 3D key points are located, among the first neck point, the second neck point, the first chest and abdomen point, the second chest and abdomen point, and the third chest and abdomen point.

[0316] In some embodiments, the above-mentioned preset posture is a posture of the legs standing upright. The electronic device 100 can determine the angle between the straight line where the leg 3D key point is located and the image plane according to the method of the above-mentioned embodiment. No further details will be given here.

[0317] In some embodiments, the preset posture is a posture with the upper body upright and the legs upright. The electronic device 100 detecting that the user's posture matches the preset posture can indicate that, in a set of 3D key points corresponding to the m frames of image, the first neck point, the second neck point, the first chest and abdomen point, the second chest and abdomen point, the third chest and abdomen point, the right hip point, the right knee point, the right ankle point, the left hip point, the left knee point, and the left ankle point are approximately on the same plane. The electronic device 100 can determine the angle between the straight line containing any two of the first neck point, the second neck point, the first chest and abdomen point, the second chest and abdomen point, the third chest and abdomen point, the right hip point (or left hip point), the right knee point (or left knee point), and the right ankle point (or left ankle point) in the set of 3D key points and the image plane. For example, the electronic device 100 determines the angle between the straight line containing the first neck point and the right ankle point in the set of 3D key points and the image plane.

[0318] Optionally, when it is detected that the user's posture matches the above-mentioned preset posture, the electronic device 100 can calculate the average of the angles between the image plane and multiple straight lines including the first neck point, the second neck point, the first chest and abdomen point, the second chest and abdomen point, the third chest and abdomen point, the right hip point (or left hip point), the right knee point (or left knee point), and the right ankle point (or left ankle point).

[0319] S406: The electronic device 100 may determine whether the angle is smaller than a preset angle.

[0320] S407: If the angle is smaller than the preset angle, the electronic device 100 may use the angle to update the compensation angle stored in the electronic device 100. After the update, the compensation angle stored in the electronic device 100 is the angle.

[0321] S408: The electronic device 100 may use the compensation angle stored in the electronic device 100 to correct the position information of the group of 3D key points.

[0322] The above steps S406 to S408 can refer to the above Figure 8 Steps S106 to S108 in the above are not described in detail here.

[0323] The foregoing Figure 10 The method for determining the initial compensation angle in the method flow chart shown can be as described above. Figure 13 In the aforementioned Figure 11 The method flow chart shown in FIG. 1 is not limited to determining the compensation angle when detecting the user's legs standing upright. The electronic device 100 can also use the above-mentioned Figure 13 The method shown is used to determine the compensation angle.

[0324] From the above Figure 13 As can be seen from the method shown, when it is detected that the user's posture matches the preset posture, the electronic device 100 can use the 3D key points determined based on the image to determine the degree of perspective deformation of the character in the image. The above-mentioned preset posture can be a posture with the upper body upright and / or the legs upright. In other words, when the user's upper body is upright and / or the legs are upright, the electronic device 100 can use the 3D key points determined based on the image to determine the degree of perspective deformation of the character in the image. Furthermore, the electronic device 100 can correct the position information of the 3D key points determined based on the image, reduce the error caused by the image perspective deformation to the position information of the 3D key points, and improve the accuracy of the detection of the position information of the 3D key points. This method only requires one camera, which not only saves costs, but also has a low computational complexity for key point detection.

[0325] Additionally, each time a user's posture is detected to match a preset posture, if the angle between the line containing the 3D key points of the upper body and / or the 3D key points of the legs and the image plane can be used to correct the position information of the 3D key points, the electronic device 100 can update the compensation angle. The updated compensation angle can more accurately reflect the degree of perspective distortion of the person in the image captured by the camera at the user's current position. This can reduce the impact of changes in the position between the user and the camera on the correction of the 3D key point position information, thereby improving the accuracy of the corrected 3D key point position information. In other words, the electronic device 100 can use the updated compensation angle to correct the position information of the 3D key points. The human body model determined by the corrected 3D key points can more accurately reflect the user's posture, improving the accuracy of posture detection. In this way, during fitness or somatosensory gaming, the electronic device 100 can more accurately determine whether the user's posture is correct and whether the amplitude of the user's movements meets the requirements, thereby providing the user with a better experience while playing fitness or somatosensory gaming.

[0326] In an embodiment of the present application, the electronic device can determine the first moment based on the first multimedia information, and the first moment can be the moment when the first multimedia information instructs the user to perform an action that satisfies the first condition. The above-mentioned first multimedia information can be relevant content in a fitness course or a somatosensory game application. The above-mentioned multimedia information may include one or more types of content in the form of video, animation, voice, text, etc. The above-mentioned action that meets the first condition can be an action of standing upright the upper body and / or an action of standing upright the legs. The electronic device can determine whether the user performs an action that meets the first condition by comparing the user's 3D key points with the 3D key points corresponding to the above-mentioned action that meets the first condition. Optionally, if the above-mentioned action that meets the first condition is an action of standing upright the legs, the electronic device can determine whether the user performs an action that meets the first condition by comparing the user's 3D key points with the 3D key points corresponding to the above-mentioned action that meets the first condition. Figure 7A The method shown is used to determine whether the user performs an action that satisfies the first condition.

[0327] In an embodiment of the present application, the electronic device can determine the compensation angle variation according to a first model. The first model is the compensation angle determination model in the aforementioned embodiment.

[0328] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting key points of a human body, characterized in that: The method is applied to an electronic device including one or more cameras, and the method includes: acquiring a first image of a first user through the camera; determining a first set of 3D key points of the first user based on the first image; determining whether multiple 3D key points in the first set of 3D key points satisfy a first condition, the first condition comprising: multiple 3D key points in the first set of 3D key points match 3D key points corresponding to a first action, the first action being at least one upright action of the upper body or the legs; If the first condition is met, a first compensation angle is determined based on the multiple 3D key points, wherein if the first action is an upper body upright action, the first compensation angle is the angle between a straight line including the neck point and the chest and abdomen points in the first group of 3D key points and the image plane of the first image; if the first action is an upright leg action, the first compensation angle is the angle between a straight line including any two 3D key points among the hip point, knee point, and ankle point in the first group of 3D key points and the image plane of the first image; performing rotation correction on the first group of 3D key points using the first compensation angle; If the first condition is not met, the first group of 3D key points is rotationally corrected using a second compensation angle, where the second compensation angle is determined based on a second group of 3D key points, which is the most recent group of 3D key points that meets the first condition before acquiring the first image.

2. The method according to claim 1, characterized in that The method of acquiring a first image of a first user by using the camera specifically includes: determining a first moment according to the first multimedia information, where the first moment is a moment when the first multimedia information instructs the user to perform an action that satisfies the first condition; The first image of the first user is acquired by the camera within a first time period starting from the first moment.

3. The method according to claim 2, characterized in that The method further comprises: If the 3D key points corresponding to the action performed by the user indicated by the first multimedia information at the second moment do not meet the first condition, a third compensation angle is used to rotationally correct the fourth group of 3D key points determined based on the fourth image captured within the second time period starting from the second moment, where the third compensation angle is determined based on the third group of 3D key points, and the third group of 3D key points is the most recent group of 3D key points that meet the first condition before the second moment.

4. The method according to any one of claims 1 to 3, characterized in that The multiple 3D key points in the first group of 3D key points include a hip point, a knee point, and an ankle point. The method of determining whether the multiple 3D key points in the first group of 3D key points meet the first condition specifically includes: Calculating a first angle between a straight line between the left hip point and the left knee point, and a straight line between the left knee point and the left foot point in the first set of 3D key points, and a second angle between a straight line between the right hip point and the right knee point, and a straight line between the right knee point and the right foot point in the first set of 3D key points; Determining whether multiple 3D key points in the first group of 3D key points meet the first condition by detecting whether a difference between the first angle and 180° is less than a first difference and whether a difference between the second angle and 180° is less than the first difference; Among them, if the first condition is met, the first compensation angle is the angle between the straight line where any two 3D key points among the hip point, knee point, and ankle point in the first group of 3D key points are located and the image plane of the first image.

5. The method according to claim 4, characterized in that The situation where multiple 3D key points in the first group of 3D key points meet the first condition includes: the difference between the first angle and 180° is less than the first difference and / or the difference between the second angle and 180° is less than the first difference.

6. The method according to any one of claims 1 to 3 and 5, characterized in that The second compensation angle is the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the second group of 3D key points are located and the image plane of the image used to determine the second group of 3D key points.

7. The method according to claim 4, characterized in that The second compensation angle is the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the second group of 3D key points are located and the image plane of the image used to determine the second group of 3D key points.

8. The method according to claim 3, characterized in that The third compensation angle is the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the third group of 3D key points are located and the image plane of the image used to determine the third group of 3D key points.

9. The method according to any one of claims 1-3, 5, 7 and 8, characterized in that The second compensation angle is the sum of the first compensation angle change and the third angle; the third angle is the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the second group of 3D key points are located and the image plane of the image used to determine the second group of 3D key points; the first compensation angle change is determined based on a first height H1, a first distance Y1, a second height H2, and a second distance Y2, wherein H1 and Y1 are respectively the height of the first user when the first image is captured and the distance between the first user and the camera, and H2 and Y2 are respectively the height of the first user when the second image is captured and the distance between the first user and the camera; the key points of the first user in the first image are the first group of 3D key points, and the key points of the first user in the second image are the second group of 3D key points; Among them, during the time period from capturing the second image to capturing the first image, the more the height of the first user decreases, the more the second compensation angle increases compared to the third angle; the more the distance between the first user and the camera decreases, the more the second compensation angle increases compared to the third angle; and the smaller H1 is, the more the second compensation angle increases compared to the third angle.

10. The method according to claim 4, characterized in that The second compensation angle is the sum of the first compensation angle change and the third angle; the third angle is the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the second group of 3D key points are located and the image plane of the image used to determine the second group of 3D key points; the first compensation angle change is determined based on a first height H1, a first distance Y1, a second height H2, and a second distance Y2, wherein H1 and Y1 are respectively the height of the first user when the first image is captured and the distance between the first user and the camera, and H2 and Y2 are respectively the height of the first user when the second image is captured and the distance between the first user and the camera; the key points of the first user in the first image are the first group of 3D key points, and the key points of the first user in the second image are the second group of 3D key points; Among them, during the time period from capturing the second image to capturing the first image, the more the height of the first user decreases, the more the second compensation angle increases compared to the third angle; the more the distance between the first user and the camera decreases, the more the second compensation angle increases compared to the third angle; and the smaller H1 is, the more the second compensation angle increases compared to the third angle.

11. The method according to claim 6, characterized in that The second compensation angle is the sum of the first compensation angle change and the third angle; the third angle is the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the second group of 3D key points are located and the image plane of the image used to determine the second group of 3D key points; the first compensation angle change is determined based on a first height H1, a first distance Y1, a second height H2, and a second distance Y2, wherein H1 and Y1 are respectively the height of the first user when the first image is captured and the distance between the first user and the camera, and H2 and Y2 are respectively the height of the first user when the second image is captured and the distance between the first user and the camera; the key points of the first user in the first image are the first group of 3D key points, and the key points of the first user in the second image are the second group of 3D key points; Among them, during the time period from capturing the second image to capturing the first image, the more the height of the first user decreases, the more the second compensation angle increases compared to the third angle; the more the distance between the first user and the camera decreases, the more the second compensation angle increases compared to the third angle; and the smaller H1 is, the more the second compensation angle increases compared to the third angle.

12. The method according to claim 3, characterized in that The third compensation angle is the sum of the second compensation angle change and the fourth angle; the fourth angle is the angle between the straight line where the upper body 3D key points and / or the leg 3D key points in the third group of 3D key points are located and the image plane of the image used to determine the third group of 3D key points; the second compensation angle change is determined based on a third height H3, a third distance Y3, a fourth height H4, and a fourth distance Y4, wherein H3 and Y3 are respectively the height of the first user when the fourth image is captured and the distance between the first user and the camera, and H4 and Y4 are respectively the height of the first user when the third image is captured and the distance between the first user and the camera; the key points of the first user in the third image are the third group of 3D key points; Among them, in the time period between capturing the third image and the fourth image, the more the height of the first user decreases, the more the third compensation angle increases compared to the fourth angle; the more the distance between the first user and the camera decreases, the more the third compensation angle increases compared to the fourth angle; and the smaller H3 is, the more the third compensation angle increases compared to the fourth angle.

13. The method according to claim 9, characterized in that The first compensation angle change is determined by a first model, and the first model is obtained by training multiple groups of training samples. One group of training samples includes: a change in the lower edge of the human body's position in the image between an image acquired earlier and an image acquired later, a change in the human body's height in the image, the human body's height in the image acquired later, and a change in the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in two groups of 3D key points determined based on the image acquired earlier and the image acquired later and the image plane; multiple 3D key points in the two groups of 3D key points all meet the first condition.

14. The method according to claim 10 or 11, characterized in that The first compensation angle change is determined by a first model, and the first model is obtained by training multiple groups of training samples. One group of training samples includes: a change in the lower edge of the human body's position in the image between an image acquired earlier and an image acquired later, a change in the human body's height in the image, the human body's height in the image acquired later, and a change in the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in two groups of 3D key points determined based on the image acquired earlier and the image acquired later and the image plane; multiple 3D key points in the two groups of 3D key points all meet the first condition.

15. The method according to claim 12, characterized in that The second compensation angle change is determined by a first model, and the first model is obtained by training multiple groups of training samples, where one group of training samples includes: a change in the lower edge of the human body's position in the image between an image acquired earlier and an image acquired later, a change in the human body's height in the image, the human body's height in the image acquired later, and a change in the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in two groups of 3D key points determined based on the image acquired earlier and the image acquired later and the image plane; multiple 3D key points in the two groups of 3D key points all meet the first condition.

16. The method according to any one of claims 1-3, 5, 7, 8, 10-13, and 15, characterized in that Before determining the first compensation angle according to the plurality of 3D key points, the method further includes: determining whether an angle between a straight line including the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and an image plane of the first image is less than a fifth angle; If the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the image plane of the first image is greater than the fifth angle, the first group of 3D key points is rotationally corrected using the second compensation angle.

17. The method according to claim 4, characterized in that Before determining the first compensation angle according to the plurality of 3D key points, the method further includes: determining whether an angle between a straight line including the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and an image plane of the first image is less than a fifth angle; If the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the image plane of the first image is greater than the fifth angle, the first group of 3D key points is rotationally corrected using the second compensation angle.

18. The method according to claim 6, characterized in that Before determining the first compensation angle according to the plurality of 3D key points, the method further includes: determining whether an angle between a straight line including the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and an image plane of the first image is less than a fifth angle; If the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the image plane of the first image is greater than the fifth angle, the first group of 3D key points is rotationally corrected using the second compensation angle.

19. The method according to claim 9, characterized in that Before determining the first compensation angle according to the plurality of 3D key points, the method further includes: determining whether an angle between a straight line including the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and an image plane of the first image is less than a fifth angle; If the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the image plane of the first image is greater than the fifth angle, the first group of 3D key points is rotationally corrected using the second compensation angle.

20. The method according to claim 14, wherein Before determining the first compensation angle according to the plurality of 3D key points, the method further includes: determining whether an angle between a straight line including the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and an image plane of the first image is less than a fifth angle; If the angle between the straight line containing the upper body 3D key points and / or the leg 3D key points in the first group of 3D key points and the image plane of the first image is greater than the fifth angle, the first group of 3D key points is rotationally corrected using the second compensation angle.

21. An electronic device, characterized in that: The electronic device includes a camera, a display screen, a memory and a processor, wherein: the camera is used to capture images, the memory is used to store computer programs, and the processor is used to call the computer program so that the electronic device executes the method described in any one of claims 1-20.

22. A computer storage medium, characterized in that include: Computer instructions; when the computer instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 20.

23. A computer program product, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Human body behavior recognition method and device based on depth camera and basic posture

    CN108305283A