Method and apparatus for determining parameter information in human body image, and electronic device
The human body image is analyzed through the human body detection network and state detection model, the target linear equation and its constraints are determined, the parameter information in the human body image is optimized, and the problem of slow inverse kinematics solving of human body and joint angle solving in the prior art is solved, achieving efficient and natural motion capture effect.
Patent Information
- Application Number
- PCT/CN2024/124483
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-10-12
- Publication Date
- 2025-06-19
AI Technical Summary
In the prior art, the method of automatically deriving the human body inverse kinematics solution (IK) method in calculating the Jacobian matrix leads to a slow solution speed, introducing a lot of time overhead, and real-time motion capture cannot be achieved; at the same time, joint angle solution does not conform to the physiological structure of the human body, and the internal parameter calibration of the monocular camera is not flexible enough, which introduces errors.
The human body detection network and state detection model are used to identify the nodes and postures of the human body images collected by the image acquisition device, determine the target linear equation and its constraints, and optimize the parameter information in the human body image by solving the target linear equation.
By optimizing and constraining relevant human body joint nodes and posture information, the natural and efficient capture of human body movement data is achieved, the real-time and accuracy of the motion capture process is improved, and the problem of poor robustness is solved.
Smart Images

Figure CN2024124483_19062025_PF_FP_ABST
Abstract
Description
Method, device and electronic equipment for determining parameter information in human body image
[0001] This application claims priority to a Chinese patent application filed with the Patent Office of China on December 12, 2023, with application number 2023117068424, entitled “Method, device and electronic device for determining parameter information in human body images”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of computer vision, and more specifically, to a method, device, and electronic device for determining parameter information in a human body image. Background Art
[0003] Human motion capture is a technology that uses cameras, inertial sensors and other devices to capture the movements of human limbs, torso, hands, and head, and generates parameters that can be used by rendering engines to drive 3D virtual digital humans.
[0004] Considering the cost and flexibility of motion capture solutions, the related art usually uses a monocular camera as a motion capture sensor. After detecting human joints, the human inverse kinematics (IK) algorithm is used to solve the motion capture parameters. However, there are certain technical problems in the human inverse kinematics (IK) process. For example, the human inverse kinematics (IK) method based on iterative optimization in the related art mainly uses an automatic derivation method when calculating the Jacobian matrix. This method is slow to solve and will introduce a large time overhead, resulting in the entire motion capture process not being able to achieve real-time. During the human inverse kinematics (IK) process, due to the limitation of the accuracy of human joint detection, the solved human joint angles often do not conform to the physiological structure of the human body. In the calibration of the internal parameters of the monocular camera, the camera internal parameters are pre-calibrated offline or solved at the initial stage of motion capture, resulting in the motion capture solution being inflexible and introducing errors. There is a lot of room for improvement in the naturalness of motion capture, especially in the interaction between the human and the ground and the rationality of the body posture.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a method, device, and electronic device for determining parameter information in a human body image, so as to at least solve the technical problem of poor robustness when capturing human body movements by using human inverse kinematics to solve human joint parameters in related technologies.
[0008] According to one aspect of an embodiment of the present application, a method for determining parameter information in a human body image is provided, comprising: using a human body detection network to perform joint point recognition on a human body image captured by an image acquisition device to obtain human body joint point coordinates; using a state detection model to recognize a human body posture in the human body image to obtain human body posture information in the human body image; determining a target linear equation based on the human body joint point coordinates, and determining constraints of the target linear equation based on the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image; solving the target linear equation based on the constraints to obtain parameter information of the human body image.
[0009] In some embodiments, the method further includes: before using a human body detection network to perform joint point recognition on a human body image captured by an image acquisition device, using a target detection network to recognize the human body image to obtain a human body envelope box corresponding to each human body in the human body image, wherein the target detection network is used to identify the position of the human body in the human body image; when the number of human body envelope boxes corresponding to the human body image is greater than 1, comparing the area of each human body envelope box in the human body image to obtain a comparison result; and using the human body envelope box with the largest area in the comparison result as the output of the target detection network.
[0010] In some embodiments, a human body detection network is used to perform joint point recognition on a human body image captured by an image acquisition device to obtain human body joint point coordinates, including: inputting a human body image containing a human body envelope frame into the human body detection network to obtain 2D joint point coordinates of the human body joints in the human body image; determining 3D joint point coordinates of the human body joints in the human body image based on the 2D joint point coordinates of the human body joints and the human body image, wherein the 3D joint point coordinates are used to reflect the real physical scale of the human body joints in the human body image, and the human body joint point coordinates include the 2D joint point coordinates of the human body joints and the 3D joint point coordinates of the human body joints.
[0011] In some embodiments, a target linear equation is determined based on the coordinates of human joint points, and constraints of the target linear equation are determined based on human posture information, including: determining 2D joint reprojection error and 3D joint position error based on human joint point coordinates; determining the center of mass projection constraint of the target linear equation based on human posture information; determining the target linear equation based on 2D joint reprojection error, 3D joint position error, center of mass projection constraint, human joint angle constraint, ground plane constraint and intrinsic parameter regularization constraint of image acquisition equipment.
[0012] In some embodiments, the human body posture information includes the grounding status of the human feet, wherein the grounding status of the human feet includes the grounding status of the left forefoot, the grounding status of the left rear foot, the grounding status of the right forefoot and the grounding status of the right rear foot.
[0013] In some embodiments, the center of mass projection constraint of the target linear equation is determined based on the human body posture information, including: when the human foot is in a grounded state and there are three points in contact with the ground, determining that the projection of the human body's center of mass to the ground plane is within the triangle formed by the three points, wherein the three points are any three of the left forefoot, the left hindfoot, the right forefoot and the right hindfoot; when the human foot is in a grounded state and there are four points in contact with the ground, determining that the projection of the human body's center of mass to the ground plane is within the quadrilateral formed by the four points, wherein the four points are the left forefoot, the left hindfoot, the right forefoot and the right hindfoot.
[0014] In some embodiments, the derivative formula corresponding to the human kinematics formula in the target linear equation is used for the joint angle of any node in any single-chain structure in the human skeleton structure, and the derivative formula is determined by the posture information of the human root node in the world coordinate system, the joint angle information in the single-chain structure, the posture information of the joint in the root node coordinate system, and the 3D coordinates of the joint in the root node coordinate system.
[0015] According to another aspect of an embodiment of the present application, a device for determining parameter information in a human body image is also provided, including: a first recognition module, used to use a human body detection network to identify the joints of a human body image captured by an image acquisition device to obtain the coordinates of the human body joints; a second recognition module, used to use a state detection model to identify the human body posture in the human body image to obtain the human body posture information in the human body image; a determination module, used to determine the target linear equation based on the coordinates of the human body joints, and to determine the constraints of the target linear equation based on the human body posture information, wherein the target linear equation is used to optimize the parameter information captured from the human body image; and a solution module, used to solve the target linear equation based on the constraints to obtain the parameter information of the human body image.
[0016] According to another aspect of the embodiments of the present application, an electronic device is also provided, including: a memory for storing program instructions; a processor, connected to the memory, for executing program instructions to implement the following functions: using a human body detection network to identify joints of a human body image captured by an image acquisition device to obtain coordinates of human body joints; using a state detection model to identify human body postures in the human body image to obtain human body posture information in the human body image; determining a target linear equation based on the coordinates of human body joints, and determining constraints of the target linear equation based on the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image; solving the target linear equation based on the constraints to obtain parameter information of the human body image.
[0017] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned method for determining parameter information in a human body image by running the computer program.
[0018] In an embodiment of the present application, a human body detection network is used to identify the joints of a human body image captured by an image acquisition device to obtain the coordinates of the human body joints; a state detection model is used to identify the human body posture in the human body image to obtain the human body posture information in the human body image; a target linear equation is determined based on the coordinates of the human body joints, and constraints of the target linear equation are determined based on the human body posture information, wherein the target linear equation is used to optimize the parameter information captured from the human body image; the target linear equation is solved based on the constraints to obtain parameter information of the human body image, thereby achieving the purpose of determining accurate human body image parameter information by optimizing and constraining relevant human body joints and posture information, thereby realizing the technical effect of capturing human body motion data more naturally and efficiently, and further solving the technical problem of poor robustness in capturing human body motion due to the use of human inverse kinematics to solve human joint parameters in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] 1 is a hardware structure diagram of a computer terminal for implementing a method for determining parameter information in a human body image according to an embodiment of the present application;
[0021] FIG2 is a flow chart of a method for determining parameter information in a human body image according to an embodiment of the present application;
[0022] FIG3 is a structural diagram of a device for determining parameter information in a human body image according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] First, some nouns or terms that appear in the process of explaining the embodiments of this application are subject to the following explanations:
[0026] Human inverse kinematics (IK): refers to solving the angle parameters of each joint of the human body when the positions of each joint of the human body are known.
[0027] Human forward kinematics (FK): refers to solving the position of each joint angle of the human body when the parameters of each joint angle of the human body are known.
[0028] SMPL (Skinned Multi-Person Linear model): A parametric modeling method for the human body that simplifies the human skeleton into 24 major joints. The parameters of this model are the parameters output by the motion capture method.
[0029] The embodiment of the method for determining parameter information in a human body image provided in the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal for implementing a method for determining parameter information in a human body image. As shown in Figure 1, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used in the figure to illustrate) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than those shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0030] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0031] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for determining parameter information in a human body image in the embodiments of the present application. The processor executes the software programs and modules stored in the memory 104 to perform various functional applications and data processing, thereby implementing the above-mentioned method for determining parameter information in a human body image. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0032] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0033] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .
[0034] It should be noted that, in some optional embodiments, the computer terminal shown in FIG1 may include hardware components (including circuits), software components (including computer code stored on a computer-readable medium), or a combination of hardware components and software components. It should be noted that FIG1 is only an example of a specific embodiment and is intended to illustrate the types of components that may be present in the computer terminal.
[0035] In the above-mentioned operating environment, an embodiment of the present application provides an embodiment of a method for determining parameter information in a human body image. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0036] FIG2 is a flow chart of a method for determining parameter information in a human body image according to an embodiment of the present application. As shown in FIG2 , the method includes the following steps:
[0037] Step S202: using a human body detection network to perform joint point recognition on the human body image captured by the image acquisition device to obtain the coordinates of the human body joint points.
[0038] In the above step S202, the input human body image is detected and analyzed through the human body detection network, and the 2D joint point coordinates and 3D joint point coordinates of the human body joints with real physical scale under the SMPL protocol can be obtained, wherein the detected joint points may include joint points such as the head, shoulders, arms, wrists, knees, and ankles.
[0039] Step S204: using a state detection model to identify the human body posture in the human body image to obtain human body posture information in the human body image.
[0040] In the above step S204, the human body posture information includes the grounding status of the human foot. By detecting the human body posture in the human body image through the state detection model, the grounding status of the left forefoot of the human foot, the grounding status of the left rear sole of the human foot, the grounding status of the right forefoot of the human foot and the grounding status of the right rear sole of the human foot can be obtained.
[0041] Step S206, determining a target linear equation based on the coordinates of the human body joint points, and determining constraints of the target linear equation based on the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image.
[0042] In the above step S206, when determining the constraint conditions of the target linear equation, human posture information can be used for constraint to improve the robustness and rationality of the target linear equation optimization process, so that the obtained target linear equation is more consistent with the movement characteristics and physiological structure of the human body.
[0043] Step S208: solving the target linear equation according to the constraint conditions to obtain parameter information of the human body image.
[0044] In the above step S208, the parameter information of the human body image is obtained according to the solution of the target linear equation, that is, the final human motion capture parameter information, including the human body root node pose, camera intrinsic parameters, camera fourth-order distortion parameters, and human body joint angle parameters.
[0045] Through the above steps S202 to S208, the purpose of determining accurate human body image parameter information by optimizing and constraining relevant human body joint points and posture information is achieved, thereby realizing the technical effect of capturing human body motion data more naturally and efficiently, and further solving the technical problem of poor robustness when capturing human body motion due to the use of human inverse kinematics to solve human joint parameters in related technologies.
[0046] In some embodiments, the method further includes: before using a human body detection network to perform joint point recognition on a human body image captured by an image acquisition device, using a target detection network to recognize the human body image to obtain a human body envelope box corresponding to each human body in the human body image, wherein the target detection network is used to identify the position of the human body in the human body image; when the number of human body envelope boxes corresponding to the human body image is greater than 1, comparing the area of each human body envelope box in the human body image to obtain a comparison result; and using the human body envelope box with the largest area in the comparison result as the output of the target detection network.
[0047] In an embodiment of the present application, a yolo-tiny network (i.e., the above-mentioned target detection network) can be used to locate and identify images captured by a monocular camera, determine the position of the human body in the human body image and the human body envelope corresponding to each human body; when there are multiple people in the human body image, that is, when the number of recognized envelopes is greater than 1, compare the area sizes of multiple envelopes corresponding to multiple people in the human body image, and select the human body envelope with the largest area as the target envelope. Among them, the yolo-tiny network is a real-time target detection network that can directly identify and locate the human body envelope from the image in a single forward propagation process, so as to realize the recognition and positioning of the human body position in the image captured by the monocular camera, and provide basic data for the subsequent analysis and application of human body images.
[0048] In some embodiments, a human body detection network is used to perform joint point recognition on a human body image captured by an image acquisition device to obtain human body joint point coordinates, including: inputting a human body image containing a human body envelope frame into the human body detection network to obtain 2D joint point coordinates of the human body joints in the human body image; determining 3D joint point coordinates of the human body joints in the human body image based on the 2D joint point coordinates of the human body joints and the human body image, wherein the 3D joint point coordinates are used to reflect the real physical scale of the human body joints in the human body image, and the human body joint point coordinates include the 2D joint point coordinates of the human body joints and the 3D joint point coordinates of the human body joints.
[0049] In an embodiment of the present application, the input human body image containing the human body envelope is detected and analyzed by the human body detection network, and the 2D joint point coordinates of the human body joints and the 3D joint point coordinates of the human body joints under the SMPL protocol can be obtained. Specifically: the human body image containing the human body envelope is detected by the human body detection network, and the human body features inside the human body image are extracted; on the extracted human body features, a key point recognition algorithm (such as a convolutional neural network) is used to detect the human body joints to obtain 2D joint point coordinates; according to the 2D joint point coordinates and human body image information (such as human body posture information and camera internal parameters, etc.), the 3D joint point coordinates are obtained, wherein the 3D joint point coordinates can reflect the real physical scale of the human body joints in the human body image; the obtained 2D joint point coordinates and 3D joint point coordinates are stored to form a complete human body joint point coordinates.
[0050] The data generated by the 2D and 3D joint detection modules often contains outliers and abnormal values due to occlusion and other factors. Therefore, data cleaning of the detected 2D and 3D human joints is required to remove outliers and noise. Specifically, a Kalman filter based on the CA model is used to smooth the same 3D joints between frames to eliminate jitter and sudden changes in joint position. Kalman filtering can effectively suppress the occurrence of these outliers, improve the accuracy and stability of joint position estimation, and provide more reliable input data for subsequent tasks such as action recognition.
[0051] In some embodiments, a target linear equation is determined based on the coordinates of human joint points, and constraints of the target linear equation are determined based on human posture information, including: determining 2D joint reprojection error and 3D joint position error based on human joint point coordinates; determining the center of mass projection constraint of the target linear equation based on human posture information; determining the target linear equation based on 2D joint reprojection error, 3D joint position error, center of mass projection constraint, human joint angle constraint, ground plane constraint and intrinsic parameter regularization constraint of image acquisition equipment.
[0052] In the embodiment of the present application, the target linear equation is determined by the coordinates of the human joint points, and different constraints are introduced according to the human posture information to improve the fitting effect. The constraints can include 2D joint point reprojection error, 3D joint point position error, center of mass projection constraint, human joint angle constraint, ground plane constraint and intrinsic parameter regularization constraint of the image acquisition device. The following describes the above six constraints respectively:
[0053] 1.2D joint point reprojection error: The 2D joint point reprojection error can be obtained by projecting the 3D joint point positions onto the image plane and comparing them with the corresponding 2D joint point coordinates.
[0054] 2.3D joint point position error: By comparing the 3D joint point position with its coordinate in the camera coordinate system, the 3D joint point position error can be obtained.
[0055] 3. Center of mass projection constraint: The human foot grounding detector detects four key points of the human foot (left forefoot, left hindfoot, right forefoot, and right hindfoot). When three or more key points are detected in contact with the ground, the center of mass projection constraint is triggered. This center of mass projection constraint ensures that the relationship between the human posture and the ground conforms to the laws of mechanics.
[0056] 4. Human Joint Angle Constraints: Based on the feasible range of human joints, constraints on human joint angles can be determined, as shown in the table below. This prevents postures that are inconsistent with human anatomy. The table below shows the coordinates and constraints corresponding to each human joint.
[0057] 5. Ground plane constraint: The ground plane equation is fitted by the positions of the key points of the foot in contact with the ground detected at each moment (left forefoot, left hindfoot, right forefoot and right hindfoot); each time the plane equation is fitted, the key points of the foot in contact with the ground in all historical frames are added, and the random sampling consensus (RANSAC) method is used to obtain the optimal plane equation; after obtaining the plane equation, additional constraints are added to the key points of the foot and added to the overall optimization problem; among them, constraint 1: the z value of the key points of the foot in contact with the ground should be near the ground plane; constraint 2: the movement speed of the key points of the foot in contact with the ground should be close to 0.
[0058] 6. Regularization constraints on the intrinsic parameters of the image acquisition device: The camera intrinsic parameters (using a pinhole camera model) and camera distortion parameters (using a fourth-order distortion model) are added to the overall optimization problem, and the camera-related parameters are optimized for each frame during the motion capture process. Among them, the initial optimization values of the camera focal length fx and fy are set to 1000, the initial values of the optical center deviations cx and cy are set to half the image width and height, respectively, and the fourth-order distortion parameters are set to 0.
[0059] Furthermore, based on the above-mentioned 2D joint point reprojection error, 3D joint point position error, center of mass projection constraint, human joint angle constraint, ground plane constraint, and intrinsic parameter regularization constraint of the image acquisition device, the target linear equation is determined as follows:
[0060] Where θ is the human joint angle; P is the 3D joint point, in meters; x is the 2D joint point, in pixels; K is the pinhole camera internal parameter; D is the pinhole camera fourth-order distortion parameter; T is the root node pose; N is the number of human joints; F is the number of ground joints; X foot_concat is the grounding state of the joint point; Π is the pinhole camera projection equation; Reg(α,n)=(α-n) 2 is the regularization equation, where α is the optimization parameter and n is the threshold; Floor(P) is the ground plane constraint equation, which represents the square of the distance between the 3D point P and the plane; ω is a scalar, which represents the weight of each error term; C(x) is the center of mass projection equation, where x represents the ground state of the joint point; R(θ) is the joint angle constraint; FK(θ) is the forward kinematics formula, where the input joint angle is θ and the output 3D joint point position is P.
[0061] In the above formula, represents the 3D joint position error, Represents the 2D joint reprojection error. This formula also covers camera self-calibration. C(xfoot_concat)*ωfoot_concat represents the center of mass projection constraint. Represents the regularization constraint of the internal parameters of the image acquisition device (camera), represents the ground plane constraint, and R(θ) represents the human joint angle constraint.
[0062] In some embodiments, the human body posture information includes the grounding status of the human feet, wherein the grounding status of the human feet includes the grounding status of the left forefoot, the grounding status of the left rear foot, the grounding status of the right forefoot and the grounding status of the right rear foot.
[0063] In some embodiments, the center of mass projection constraint of the target linear equation is determined based on the human body posture information, including: when the human foot is in a grounded state and there are three points in contact with the ground, determining that the projection of the human body's center of mass to the ground plane is within the triangle formed by the three points, wherein the three points are any three of the left forefoot, the left hindfoot, the right forefoot and the right hindfoot; when the human foot is in a grounded state and there are four points in contact with the ground, determining that the projection of the human body's center of mass to the ground plane is within the quadrilateral formed by the four points, wherein the four points are the left forefoot, the left hindfoot, the right forefoot and the right hindfoot.
[0064] In an embodiment of the present application, the center of mass projection constraint of the target linear equation is determined by human body posture information. The human body posture information may include the judgment of the grounding status of the human feet, that is, the grounding status of the left forefoot, the grounding status of the left rear foot, the grounding status of the right forefoot and the grounding status of the right rear foot, which is used to determine the support point of the human body on the ground, and then trigger the center of mass projection constraint. Specifically: four key points of the human foot (left forefoot, left hindfoot, right forefoot and right hindfoot) are detected by a human foot grounding detector. When three or more key points are detected to be in contact with the ground, it indicates that the relationship between the human posture and the ground conforms to the laws of mechanics. Among them, when three key points of the human foot are in contact with the ground, that is, when any three of the left forefoot, left hindfoot, right forefoot and right hindfoot are in contact with the ground, the projection of the human body's center of mass onto the ground plane is determined to be within the triangle formed by the three points. When four key points of the human foot are in contact with the ground, that is, when all of the left forefoot, left hindfoot, right forefoot and right hindfoot are in contact with the ground, the projection of the human body's center of mass onto the ground plane is determined to be within the quadrilateral formed by the four points. By constraining the center of mass projection constraint, the center of gravity of the human body can be ensured to be within a reasonable range, thereby increasing the stability and reliability of the target linear equation.
[0065] In some embodiments, the derivative formula corresponding to the human kinematics formula in the target linear equation is used for the joint angle of any node in any single-chain structure in the human skeleton structure, and the derivative formula is determined by the posture information of the human root node in the world coordinate system, the joint angle information in the single-chain structure, the posture information of the joint in the root node coordinate system, and the 3D coordinates of the joint in the root node coordinate system.
[0066] In an embodiment of the present application, a human forward kinematics (FK) formula (i.e., the above-mentioned human kinematics formula) is also provided. Since the human trunk and limbs can be regarded as multiple chain structures (with the human pelvic node as the parent node), such as the pelvis-spine-clavicle-shoulder joint-elbow joint-wrist joint forms a single chain, the derivative formula corresponding to the above-mentioned human kinematics formula can be used to calculate the joint angle of any node in any single chain structure in the human skeleton structure. The derivative formula is determined by the posture information of the human root node in the world coordinate system, the joint angle information in the single chain structure, the posture information of the joint in the root node coordinate system, and the 3D coordinates of the joint in the root node coordinate system to help determine the movement and angle changes of each joint in the human skeleton structure, as follows:
[0067] in, is the rotation matrix of the root node of the human body in the world coordinate system; R i is the rotation matrix form of the i-th joint angle in a single chain; The rotation matrix form of the posture of the i-th joint in the root node coordinate system; is the 3D coordinate of the i-th joint in the root node coordinate system; a^ is an antisymmetric matrix, which represents the expansion of the 3D vector a.
[0068] It should be noted that when the human kinematics formula is used to derive the joint angle of any node in any single chain structure in the human skeleton structure, that is, j>k, that is, j must be the parent node of k; it is interpreted as the joint angle θ in the Lie algebra form of the function of finding the position of the jth node in a single chain using FK k The partial derivative of .
[0069] In the embodiments of this application, the human body is defined as a standard, integrally floating torque control system, in which the torque interaction between each joint and muscle determines the human body's motion. In order to perform secondary optimization on human motion capture, that is, to further optimize the human body's joint angles and make the human body's motion more reasonable, the following equation is solved:
[0070] in, p root is the coordinate of the root joint, r is the angle of each joint of the human body, N=3+3K, K is the number of joints; is the inertia matrix; h is the nonlinear effect term considering gravity, Coriolis force and centripetal force; is the observed value of the joint angular acceleration; J is the joint Jacobian matrix; J c is the Jacobian matrix of the grounded joint; τ is the driving force on the joint; λ is the additional support force on the grounded joint; and p is the position of the joint.
[0071] In an embodiment of the present application, by using the general mathematical form of the FK formula to derive the human joint angle under the human skeleton structure, the human inverse kinematics solution process (IK) can be greatly accelerated, thereby optimizing the solution time from 30ms to 5ms, improving the calculation efficiency and real-time performance; adding the human joint angle constraint term can prevent the appearance of postures that do not conform to the physiological structure of the human body in the motion capture results, improve the accuracy and authenticity of the motion capture results, and enhance the robustness of the system; introducing human kinematic constraints makes the motion capture results more natural, smooth, and realistic, and improves the realism and realism of motion capture; through the full-time camera parameter self-calibration method, the estimation of camera parameters is more accurate, thereby improving the motion capture precision and accuracy, and enhancing the reliability of the system; introducing the human center of mass constraint term and the ground plane estimation method makes the estimation of the human body's global posture more reasonable and natural, avoiding unreasonable situations such as feet hanging in the air or inserted into the ground, and improving the authenticity and rationality of the motion capture results.
[0072] According to an embodiment of the present application, a device for determining parameter information in a human body image is provided. It should be noted that the device for determining parameter information in a human body image according to the embodiment of the present application can be used to execute the method for determining parameter information in a human body image according to the embodiment of the present application. The following describes the device for determining parameter information in a human body image according to the embodiment of the present application.
[0073] FIG3 is a structural diagram of a device for determining parameter information in a human body image according to an embodiment of the present application. As shown in FIG3 , the device includes:
[0074] A first recognition module 30 is configured to use a human body detection network to perform joint recognition on a human body image acquired by an image acquisition device to obtain coordinates of the human body joints;
[0075] A second recognition module 32 is used to recognize the human posture in the human image using a state detection model to obtain human posture information in the human image;
[0076] a determination module 34 for determining a target linear equation based on the coordinates of the human body joint points and determining constraints of the target linear equation based on the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image;
[0077] Solving module 36 is used to solve the target linear equation according to the constraint conditions to obtain the parameter information of the human body image
[0078] Through the first recognition module 30, the second recognition module 32, the determination module 34, and the solution module 36 in the above-mentioned device for determining parameter information in the human body image, the purpose of determining accurate human body image parameter information by optimizing and constraining relevant human body joint points and posture information is achieved, thereby achieving the technical effect of capturing human body motion data more naturally and efficiently, and further solving the technical problem of poor robustness when capturing human body motion due to the use of human inverse kinematics to solve human joint parameters in related technologies.
[0079] In the device for determining parameter information in a human body image provided in an embodiment of the present application, the first recognition module is also used to use a target detection network to identify the human body image to obtain a human body envelope frame corresponding to each human body in the human body image, wherein the target detection network is used to identify the position of the human body in the human body image; when the number of human body envelope frames corresponding to the human body image is greater than 1, the area of each human body envelope frame in the human body image is compared to obtain a comparison result; and the human body envelope frame with the largest area in the comparison result is used as the output of the target detection network.
[0080] In the device for determining parameter information in a human body image provided in an embodiment of the present application, the first recognition module is also used to input a human body image containing a human body envelope frame into a human body detection network to obtain 2D joint point coordinates of human body joints in the human body image; based on the 2D joint point coordinates of the human body joints and the human body image, the 3D joint point coordinates of the human body joints in the human body image are determined, wherein the 3D joint point coordinates are used to reflect the real physical scale of the human body joints in the human body image, and the human joint point coordinates include the 2D joint point coordinates of the human body joints and the 3D joint point coordinates of the human body joints.
[0081] In the device for determining parameter information in a human body image provided in an embodiment of the present application, the second recognition module is also used to identify the grounding status of human feet, including the grounding status of the left forefoot, the grounding status of the left hindfoot, the grounding status of the right forefoot, and the grounding status of the right hindfoot.
[0082] In the device for determining parameter information in a human body image provided in an embodiment of the present application, the determination module is also used to determine the 2D joint point reprojection error and the 3D joint point position error based on the human joint point coordinates; determine the center of mass projection constraint of the target linear equation based on the human body posture information; and determine the target linear equation based on the 2D joint point reprojection error, the 3D joint point position error, the center of mass projection constraint, the human body joint angle constraint, the ground plane constraint, and the intrinsic parameter regularization constraint of the image acquisition device.
[0083] In the device for determining parameter information in a human body image provided in an embodiment of the present application, the determination module is also used to determine that, when three points of the human foot are in contact with the ground, the projection of the human body's center of mass onto the ground plane is within the triangle formed by the three points, wherein the three points are any three of the left forefoot, the left hindfoot, the right forefoot, and the right hindfoot; and when four points of the human foot are in contact with the ground, the projection of the human body's center of mass onto the ground plane is within the quadrilateral formed by the four points, wherein the four points are the forefoot, the left hindfoot, the right forefoot, and the right hindfoot.
[0084] An embodiment of the present application also provides an electronic device, including: a memory for storing program instructions; a processor, connected to the memory, for executing program instructions to implement the following functions: using a human body detection network to identify joint points of a human body image captured by an image acquisition device to obtain coordinates of the human body joint points; using a state detection model to identify the human body posture in the human body image to obtain human body posture information in the human body image; determining a target linear equation based on the coordinates of the human body joint points, and determining the constraints of the target linear equation based on the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image; solving the target linear equation based on the constraints to obtain parameter information of the human body image.
[0085] It should be noted that the above-mentioned electronic device is used to execute the method for determining parameter information in a human body image shown in FIG2 , so the relevant explanations in the above-mentioned method for determining parameter information in a human body image are also applicable to the electronic device and will not be repeated here.
[0086] An embodiment of the present application further provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned method for determining parameter information in a human body image by running the computer program.
[0087] It should be noted that the above-mentioned non-volatile storage medium is used to execute the method for determining parameter information in the human body image shown in Figure 2. Therefore, the relevant explanations in the above-mentioned method for determining parameter information in the human body image are also applicable to the non-volatile storage medium and will not be repeated here.
[0088] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0089] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0090] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0091] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0092] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0093] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0094] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for determining parameter information in a human body image, comprising: A human body detection network is used to identify joint points of human body images collected by an image acquisition device to obtain the coordinates of human body joint points; Using a state detection model to identify the human body posture in the human body image to obtain the human body posture information in the human body image; Determining a target linear equation according to the coordinates of the human body joint points, and determining a constraint condition of the target linear equation according to the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image; The target linear equation is solved according to the constraint conditions to obtain parameter information of the human body image.
2. The method according to claim 1, further comprising: Before using the human body detection network to identify joint points of human body images collected by the image acquisition device, Using a target detection network to identify the human body image, and obtaining a human body envelope corresponding to each human body in the human body image, wherein the target detection network is used to identify the human body position in the human body image; When the number of human body envelope frames corresponding to the human body image is greater than 1, comparing the area of each human body envelope frame in the human body image to obtain a comparison result; The human body envelope with the largest area in the comparison results is used as the output of the target detection network.
3. The method according to claim 2, wherein: The human body detection network is used to identify the joint points of the human body images collected by the image acquisition device to obtain the coordinates of the human body joint points, including: Inputting a human body image including the human body envelope into the human body detection network to obtain 2D joint point coordinates of human body joints in the human body image; Based on the 2D joint point coordinates of the human joints and the human body image, the 3D joint point coordinates of the human joints in the human body image are determined, wherein the 3D joint point coordinates are used to reflect the real physical scale of the human joints in the human body image, and the human joint point coordinates include the 2D joint point coordinates of the human joints and the 3D joint point coordinates of the human joints.
4. The method according to claim 1, wherein: Determining a target linear equation according to the human body joint point coordinates, and determining a constraint condition of the target linear equation according to the human body posture information, including: Determine 2D joint point reprojection error and 3D joint point position error according to the human body joint point coordinates; Determining the centroid projection constraint of the target linear equation according to the human body posture information; The target linear square is determined based on the 2D joint point reprojection error, the 3D joint point position error, the centroid projection constraint, the human body joint angle constraint, the ground plane constraint and the intrinsic parameter regularization constraint of the image acquisition device. Procedure.
5. The method according to claim 4, wherein: The human body posture information includes the grounding state of human feet, wherein the grounding state of human feet includes the grounding state of the left forefoot, the grounding state of the left rear foot, the grounding state of the right forefoot and the grounding state of the right rear foot.
6. The method according to claim 5, wherein: Determining the centroid projection constraint of the target linear equation according to the human body posture information includes: In the case where three points of the human foot are in contact with the ground, determining that the projection of the human body's center of mass onto the ground plane is within a triangle formed by the three points, wherein the three points are any three of the left forefoot, the left rear foot, the right forefoot, and the right rear foot; When four points of the human foot are in contact with the ground, the projection of the center of mass of the human body onto the ground plane is determined to be within the quadrilateral formed by the four points, wherein the four points are the left forefoot, the left rear foot, the right forefoot and the right rear foot.
7. The method according to claim 1, wherein: The derivative formula corresponding to the human kinematics formula in the target linear equation is used for the joint angle of any node in any single-chain structure in the human skeleton structure, and the derivative formula is determined by the posture information of the human root node in the world coordinate system, the joint angle information in the single-chain structure, the posture information of the joint in the root node coordinate system, and the 3D coordinates of the joint in the root node coordinate system.
8. The method according to claim 4, wherein: The determining of 2D joint point reprojection errors and 3D joint point position errors based on the human body joint point coordinates includes: The 2D joint point reprojection error is obtained by projecting the 3D joint point position onto the image plane and comparing it with the corresponding 2D joint point coordinates; The 3D joint point position error is obtained by comparing the 3D joint point position with its coordinates in the camera coordinate system.
9. A device for determining parameter information in a human body image, comprising: The first recognition module is used to use a human body detection network to perform joint point recognition on a human body image acquired by an image acquisition device to obtain coordinates of human body joint points; A second recognition module is used to recognize the human body posture in the human body image by using a state detection model to obtain the human body posture information in the human body image; a determination module, configured to determine a target linear equation according to the coordinates of the human body joint points, and to determine a constraint condition of the target linear equation according to the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image; A solution module is used to solve the target linear equation according to the constraint conditions to obtain parameter information of the human body image.
10. An electronic device, comprising: A memory for storing program instructions; A processor is connected to the memory and is used to execute program instructions to implement the following functions: using a human body detection network to identify joints of a human body image acquired by an image acquisition device to obtain coordinates of human body joints; using a state detection model to identify human body postures in the human body image to obtain human body posture information in the human body image; determining a target linear equation based on the coordinates of the human body joints, and determining constraints of the target linear equation based on the human body posture information, wherein the target linear equation is used to optimize parameter information captured from the human body image; solving the target linear equation based on the constraints to obtain parameter information of the human body image.
11. A non-volatile storage medium comprising a stored computer program, wherein: The device where the non-volatile storage medium is located executes the method for determining parameter information in a human body image as described in any one of claims 1 to 8 by running the computer program.
Citation Information
Patent Citations
Motion capture method and device, equipment and storage medium
CN112381003A
Monocular video-based multi-stage human motion capture method and device, and medium
CN116386141A
Three-dimensional crowd data generation method based on single-view color image
CN117079066A
Method and device for determining parameter information in human body image and electronic equipment
CN117636399A
System and method for image capture device pose estimation
US20170249751A1