A 3D Human Pose Estimation Method and System

By using dual depth cameras for data acquisition and generator-based judgment, combined with a pre-trained neural network, the problems of error amplification and spatial information loss in 3D pose estimation are solved, achieving more accurate 3D human pose estimation.

CN116824619BActive Publication Date: 2026-04-03HUIZHIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies for 3D pose estimation rely too heavily on 2D pose estimation networks, leading to amplified errors and loss of spatial information, resulting in low accuracy.

Method used

A dual-depth camera is used to capture horizontal and vertical perspectives. A generator is used to determine the pose and generate a 3D human pose. The pose analysis is performed using a pre-trained neural network and a neural network unit combination generator.

Benefits of technology

It improves the accuracy of 3D human pose estimation by using dual depth cameras to capture data and a generator to determine the motion sequence, thus generating a more accurate 3D human pose.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824619B_ABST
    Figure CN116824619B_ABST
Patent Text Reader

Abstract

This invention provides a three-dimensional human pose estimation method and system. The method includes: Step 1: acquiring a first horizontal view and a second vertical view of the target human body using dual depth cameras configured at acquisition points, wherein the horizontal and vertical views are related to the height of the target human body; Step 2: performing pose determination on the first and second acquired images based on a pre-trained neural network and a generator composed of neural network units to obtain a human motion sequence; Step 3: constructing a three-dimensional human pose based on the human motion sequence. By acquiring horizontal and vertical data using dual depth cameras and performing pose determination using a generator to obtain a motion sequence, and then generating a three-dimensional human pose, the accuracy of pose acquisition is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a three-dimensional human pose estimation method and system. Background Technology

[0002] Currently, in the process of 3D conversion, one approach is to revert 2D image information to 3D coordinates to obtain 3D information. Another approach is to first obtain 2D information and then upscale it to 3D space to obtain 3D information. However, when these two approaches are applied to 3D pose estimation, they suffer from excessive reliance on 2D pose estimation networks, which leads to error amplification and loss of spatial information, resulting in low accuracy of 3D pose estimation.

[0003] Therefore, this invention proposes a three-dimensional human pose estimation method and system. Summary of the Invention

[0004] This invention provides a three-dimensional human pose estimation method and system, which uses dual depth cameras to acquire horizontal and vertical data, and uses a generator to determine the pose, obtain a motion sequence, and then generate a three-dimensional human pose, effectively improving the accuracy of pose acquisition.

[0005] This invention provides a three-dimensional human pose estimation method, comprising:

[0006] Step 1: Based on the dual depth cameras configured at the acquisition point, perform a first acquisition of the target human body from a horizontal perspective and a second acquisition of the target human body from a vertical perspective, wherein the horizontal and vertical perspectives are related to the height of the target human body.

[0007] Step 2: Based on the pre-trained neural network and the generator composed of neural network units, perform pose determination on the first and second acquired images to obtain the human motion sequence;

[0008] Step 3: Based on the human motion sequence, construct the three-dimensional human posture of the target human body.

[0009] Preferably, the method involves first acquiring a horizontal view of the target human body and second acquiring a vertical view of the target human body using dual depth cameras configured at the acquisition points, including:

[0010] The first distance between the center position of the target human body and the dual depth camera is obtained, and based on the first coordinate system of the dual depth camera facing the center position, the straight line corresponding to the first distance is displayed on the first coordinate system, the first angle between the straight line corresponding to the first distance and the first coordinate system is obtained, and the horizontal and vertical viewing angles of the dual depth camera are determined.

[0011] When the dual-depth camera performs the first acquisition of the target human body from a horizontal perspective, the supplementary lighting device is controlled to provide the first supplementary lighting to the dual-depth camera based on the first acquisition state of the dual-depth camera and the first environmental state.

[0012] When the dual-depth camera performs a second acquisition of the target human body from a vertical perspective, the supplementary lighting device is controlled to provide a second supplementary light to the dual-depth camera based on the second acquisition state of the dual-depth camera and the second environmental state.

[0013] After acquiring the dual-depth camera image based on the first supplementary light and the second supplementary light, a first acquired image and a second acquired image are obtained.

[0014] Preferred options also include:

[0015] Obtain the network construction path of the pre-trained neural network and the unit construction path of the neural network unit;

[0016] Based on the network construction lines and unit construction lines, determine the line matrix, and obtain the attitude estimation template contained in each matrix unit of the line matrix respectively;

[0017] The attitude estimation template is parsed to obtain a patch group for each matrix unit, wherein the patch group contains several two-dimensional attitude structures;

[0018] Perform structural overlap and non-overlap determination on all the patch groups in the matrix unit, and perform a first calibration on the overlapping two-dimensional structures and a second calibration on the non-overlapping two-dimensional structures in the same patch group;

[0019] Based on the frequency of occurrence of the same first label and its importance in each piece group, a first weight is assigned to the corresponding overlapping two-dimensional structure in each piece group;

[0020] Based on the importance of the corresponding piece groups, a second weight is assigned to the corresponding non-overlapping two-dimensional structure.

[0021] The cell setting result is obtained based on the first weight set for the overlapping two-dimensional structures in the same matrix cell and the second weight set for the non-overlapping two-dimensional structures.

[0022] The generator is obtained based on the settings of all units and the constructed circuit.

[0023] Preferably, the pose of the first and second acquired images is determined to obtain a human motion sequence, including:

[0024] Based on the generator, the first acquired image is analyzed for horizontal pixels to obtain a first pose sequence;

[0025] Based on the generator, the second acquired image is analyzed vertically to obtain a second pose sequence;

[0026] Based on the first posture sequence and the second posture sequence, a human motion sequence is obtained.

[0027] Preferably, a human motion sequence is obtained based on the first posture sequence and the second posture sequence, including:

[0028] The first pose sequence and the second pose sequence are combined at the same pixel point to obtain a combination pair of each pixel point;

[0029] Determine whether the first sequence and the second sequence in the combination pair are the same. If they are the same, analyze the next combination pair.

[0030] If they are inconsistent, the first template that matches the first pose sequence and the second template that matches the second pose sequence are matched based on the pose estimation template of the generator.

[0031] Based on the weight setting results of each template behavior in the first template and the weight setting results of each template behavior in the second template, determine whether the sequence difference between the first sequence and the second sequence meets the setting criteria.

[0032] If satisfied, analyze the next pair;

[0033] If the conditions are not met, the combination pair is determined to be abnormal;

[0034] Determine the first number of abnormal pairs and the second number of qualified pairs;

[0035] When the first quantity is less than the second quantity, the abnormal combination pairs are compensated and corrected based on the combination pairs corresponding to the second quantity to obtain qualified combination pairs;

[0036] The human motion sequence is obtained by combining the horizontal and vertical sequences of all qualified pairs.

[0037] Preferably, based on the combination pairs corresponding to the second quantity, abnormal combination pairs are compensated and corrected to obtain qualified combination pairs, including:

[0038] Determine the distribution of the first position points of the combination pairs corresponding to the second quantity;

[0039] Determine the distribution of the second position points of the combination pairs corresponding to the first quantity, and determine the positional relationship between the second position point distribution and the first position point distribution;

[0040] When the positions are intersecting, the abnormal combination pair itself is compensated for by the connecting line between the normal combination pair adjacent to the abnormal combination pair, and the intersecting line is obtained.

[0041] When the position is non-intersecting, the abnormal combination pair is contour compensated according to the behavioral contour of the matched first template and the behavioral contour of the second template to obtain the compensation line.

[0042] Based on the cross lines and compensation lines, qualified combination pairs are obtained.

[0043] Preferably, the three-dimensional human posture of the target human body, based on a human motion sequence, includes:

[0044] The human motion sequence is input into the 3D generation model in sequence to obtain the 3D human posture.

[0045] This invention provides a three-dimensional human pose estimation system, comprising:

[0046] The human body acquisition module is used to acquire the target human body from a horizontal perspective and from a vertical perspective based on dual depth cameras configured at the acquisition point, wherein the horizontal and vertical perspectives are related to the height of the target human body.

[0047] The sequence generation module is used to determine the pose of the first and second acquired images based on a pre-trained neural network and a generator composed of neural network units, and obtain a human motion sequence.

[0048] The posture composition module is used to construct the three-dimensional human posture of the target human body based on the human motion sequence.

[0049] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings.

[0050] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0051] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0052] Figure 1 This is a flowchart of a three-dimensional human pose estimation method in an embodiment of the present invention;

[0053] Figure 2 This is a structural diagram of a three-dimensional human pose estimation system according to an embodiment of the present invention. Detailed Implementation

[0054] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0055] This invention provides a three-dimensional human pose estimation method, such as... Figure 1 As shown, it includes:

[0056] Step 1: Based on the dual depth cameras configured at the acquisition point, perform a first acquisition of the target human body from a horizontal perspective and a second acquisition of the target human body from a vertical perspective, wherein the horizontal and vertical perspectives are related to the height of the target human body.

[0057] Step 2: Based on the pre-trained neural network and the generator composed of neural network units, perform pose determination on the first and second acquired images to obtain the human motion sequence;

[0058] Step 3: Based on the human motion sequence, construct the three-dimensional human posture of the target human body.

[0059] In this embodiment, the generator is a model composed of a pre-trained ResNet-50 neural network and LSTM units, used to generate reasonable human motion sequences. Reasonable human motion postures are determined by a multi-layer LSTM unit and a Transformer-based discriminator.

[0060] In this embodiment, dual depth cameras refer to the presence of two depth cameras.

[0061] In this embodiment, the horizontal viewing angle refers to the first acquisition of the camera's horizontal viewing angle range, and the vertical viewing angle refers to the second acquisition of the camera's vertical viewing angle range. The horizontal viewing angle range is 30-70 degrees, and the vertical viewing angle range is 90-180 degrees.

[0062] In this embodiment, in the dual-depth cameras, one camera is mainly responsible for acquiring the horizontal viewpoint, and the other camera is mainly responsible for acquiring the vertical viewpoint, ensuring that each camera performs its own task.

[0063] In this embodiment, the height of the human body is only used to provide a center point for adjusting the acquisition range of the horizontal and vertical viewing angles.

[0064] In this embodiment, the first acquired image is mainly for the analysis of horizontal pixels, and the second acquired image is mainly for the analysis of vertical pixels. In this way, the generator judges the posture of the image to obtain the human motion sequence.

[0065] In this embodiment, the purpose of acquiring the human motion sequence is to effectively convert it into a three-dimensional human pose.

[0066] The beneficial effects of the above technical solution are: by using dual depth cameras to acquire data horizontally and vertically, and by using a generator to determine the pose, a motion sequence is obtained, which in turn generates a three-dimensional human pose, effectively improving the accuracy of pose acquisition.

[0067] This invention provides a three-dimensional human pose estimation method, which performs a first acquisition of the target human body from a horizontal perspective and a second acquisition of the target human body from a vertical perspective based on dual depth cameras configured at acquisition points, including:

[0068] The first distance between the center position of the target human body and the dual depth camera is obtained, and based on the first coordinate system of the dual depth camera facing the center position, the straight line corresponding to the first distance is displayed on the first coordinate system, the first angle between the straight line corresponding to the first distance and the first coordinate system is obtained, and the horizontal and vertical viewing angles of the dual depth camera are determined.

[0069] When the dual-depth camera performs the first acquisition of the target human body from a horizontal perspective, the supplementary lighting device is controlled to provide the first supplementary lighting to the dual-depth camera based on the first acquisition state of the dual-depth camera and the first environmental state.

[0070] When the dual-depth camera performs a second acquisition of the target human body from a vertical perspective, the supplementary lighting device is controlled to provide a second supplementary light to the dual-depth camera based on the second acquisition state of the dual-depth camera and the second environmental state.

[0071] After acquiring the dual-depth camera image based on the first supplementary light and the second supplementary light, a first acquired image and a second acquired image are obtained.

[0072] In this embodiment, the central body refers to the intersection of the horizontal line and the vertical line constructed based on the target human body, and the first distance is the distance between the intersection and the position point of the dual depth camera.

[0073] In this embodiment, the first coordinate system is based on the horizontal line as the abscissa and the vertical line as the ordinate, with the direction of the abscissa pointing towards the target human body. At this time, the straight line corresponding to the first distance can be drawn on the first coordinate system to obtain the first included angle.

[0074] In this embodiment, the first acquisition state is related to the horizontal viewing angle, and the first environmental state is related to the ambient light around the camera that acquires data based on the horizontal viewing angle. At this time, the ambient lighting is adjusted to achieve the first supplementary lighting. The first supplementary lighting is not only a simple adjustment of brightness, but also an adjustment of the supplementary lighting device. When the supplementary lighting device reaches a certain position, the first supplementary lighting is achieved at that position.

[0075] In this embodiment, the second acquisition state is related to the vertical viewing angle, and the second environmental state is related to the ambient light or darkness around the camera that acquires data based on the vertical viewing angle.

[0076] In this embodiment, the first acquired image is obtained after the camera is illuminated with a first supplementary light, and then the camera is used to acquire the image. The second acquired image is similar to the first acquired image, and will not be described again here.

[0077] The beneficial effects of the above technical solution are: by obtaining the distance between the central body and the dual depth cameras and obtaining the first included angle after constructing the coordinate system, the horizontal and vertical viewing angles can be effectively determined, and by performing horizontal and vertical supplementary lighting, the corresponding images can be effectively acquired, providing a valid basis for subsequent attitude analysis.

[0078] This invention provides a three-dimensional human pose estimation method, which further includes:

[0079] Obtain the network construction path of the pre-trained neural network and the unit construction path of the neural network unit;

[0080] Based on the network construction lines and unit construction lines, determine the line matrix, and obtain the attitude estimation template contained in each matrix unit of the line matrix respectively;

[0081] The attitude estimation template is parsed to obtain a patch group for each matrix unit, wherein the patch group contains several two-dimensional attitude structures;

[0082] Perform structural overlap and non-overlap determination on all the patch groups in the matrix unit, and perform a first calibration on the overlapping two-dimensional structures and a second calibration on the non-overlapping two-dimensional structures in the same patch group;

[0083] Based on the frequency of occurrence of the same first label and its importance in each piece group, a first weight is assigned to the corresponding overlapping two-dimensional structure in each piece group;

[0084] Based on the importance of the corresponding piece groups, a second weight is assigned to the corresponding non-overlapping two-dimensional structure.

[0085] The cell setting result is obtained based on the first weight set for the overlapping two-dimensional structures in the same matrix cell and the second weight set for the non-overlapping two-dimensional structures.

[0086] The generator is obtained based on the settings of all units and the constructed circuit.

[0087] In this embodiment,

[0088] In this embodiment, the second weight = importance × conversion coefficient.

[0089] In this embodiment, the unit setting result refers to the result obtained by determining the weights of all poses involved in the unit and combining them with the pose attributes of the pose itself.

[0090] In this embodiment, the circuit is constructed to provide a path for image analysis, thereby ensuring the rationality of subsequent analysis of the acquired images.

[0091] In this embodiment, the network construction path refers to a construction path of the neural network based on the generator. For example, there are neural networks 01, 02 and 03. The construction path is that 01 and 02 are constructed first, and 03 is constructed later. Similarly, the unit construction path refers to whether the units are constructed together in the same pool or in different pools. The main purpose is to obtain the construction path.

[0092] In this embodiment, the line matrix mainly refers to the poses involved in the construction process of different lines. The pose is determined based on the network construction line and the unit construction line to determine its position. The line matrix is ​​initially a blank matrix with various construction elements set in advance. Based on the network construction line and the unit construction line, the element information of the construction elements involved is obtained and matched to the corresponding blank position in the matrix. The element information is related to the different poses involved in the construction process, thus obtaining the line matrix.

[0093] In this embodiment, the pose estimation template refers to different human poses.

[0094] In this embodiment, the model is analyzed mainly to obtain the two-dimensional structure of different posture behaviors in the corresponding matrix unit. This two-dimensional structure is mainly related to the posture motion trajectory.

[0095] In this embodiment, the matrix unit in the line matrix refers to the unit at the position of n rows and m columns in the matrix. That is, a specified row and column can constitute a unit. For example, there are 3 matrix units. Matrix unit 1 contains structures 1, 2, and 3. Matrix unit 2 contains structures 01, 2, and 3. Matrix unit 3 contains structures 1, 2, and 03. At this time, structures 1, 2, and 3 in matrix unit 1 are first calibrated. Structures 2 and 3 in matrix unit 2 are first calibrated, and 01 is second calibrated. Structures 1 and 2 in matrix unit 3 are first calibrated, and structure 03 is second calibrated.

[0096] In this embodiment, the number of times the same calibration occurs, for example, structure 1 occurs 2 times, structure 2 occurs 3 times, and structure 3 occurs 2 times. Importance refers to the importance of different structures in the corresponding piece group, that is, the importance of the pose.

[0097] The beneficial effects of the above technical solution are: by constructing a matrix and analyzing each element in the matrix, the weight setting results of the corresponding element can be effectively determined, providing a basis for building the generator and indirectly improving the accuracy of subsequent pose recognition.

[0098] This invention provides a three-dimensional human pose estimation method, which performs pose determination on a first acquired image and a second acquired image to obtain a human motion sequence, including:

[0099] Based on the generator, the first acquired image is analyzed for horizontal pixels to obtain a first pose sequence;

[0100] Based on the generator, the second acquired image is analyzed vertically to obtain a second pose sequence;

[0101] Based on the first posture sequence and the second posture sequence, a human motion sequence is obtained.

[0102] In this embodiment, the human motion sequence is obtained by combining the first posture sequence and the second posture sequence, for example, by averaging the two.

[0103] The beneficial effect of the above technical solution is that by acquiring horizontal and vertical sequences, it is easier to improve the accuracy of the acquired motion sequences.

[0104] This invention provides a three-dimensional human pose estimation method, which obtains a human motion sequence based on a first pose sequence and a second pose sequence, including:

[0105] The first pose sequence and the second pose sequence are combined at the same pixel point to obtain a combination pair of each pixel point;

[0106] Determine whether the first sequence and the second sequence in the combination pair are the same. If they are the same, analyze the next combination pair.

[0107] If they are inconsistent, the first template that matches the first pose sequence and the second template that matches the second pose sequence are matched based on the pose estimation template of the generator.

[0108] Based on the weight setting results of each template behavior in the first template and the weight setting results of each template behavior in the second template, determine whether the sequence difference between the first sequence and the second sequence meets the setting criteria.

[0109] If satisfied, analyze the next pair;

[0110] If the conditions are not met, the combination pair is determined to be abnormal;

[0111] Determine the first number of abnormal pairs and the second number of qualified pairs;

[0112] When the first quantity is less than the second quantity, the abnormal combination pairs are compensated and corrected based on the combination pairs corresponding to the second quantity to obtain qualified combination pairs;

[0113] The human motion sequence is obtained by combining the horizontal and vertical sequences of all qualified pairs.

[0114] In this embodiment, the combination pair is: [first pose sequence and second pose sequence].

[0115] In this embodiment, whether they are consistent refers to whether the values ​​of the first attitude sequence and the second attitude sequence are consistent.

[0116] In this embodiment, the first template is obtained by matching from the pose estimation template, and the second template is obtained by matching from the pose estimation template, which is the final matched behavioral pose.

[0117] In this embodiment, the weight setting result refers to the weight setting of the corresponding behavior posture.

[0118] In this embodiment, for example, the weight setting result corresponding to the first template is 0.1, and the weight setting result corresponding to the second template is 0.1. This means that the corresponding behavioral posture is not important, so the size of the sequence difference does not affect the judgment result. In other words, it is considered to meet the setting criteria at this time.

[0119] If the weight setting result corresponding to the first template is 0.2 and the weight setting result corresponding to the second template is 0.2, then the weight setting result affects the behavior posture. If the sequence difference between the two is greater than 0.5, it is considered that the setting standard is not met; otherwise, it is determined that the setting standard is met.

[0120] In this embodiment, for example, the original abnormal combination pair [1, 0.9] is compensated and corrected to become the combination pair [1, 1], thus obtaining a qualified combination pair.

[0121] The beneficial effects of the above technical solution are: by combining the same pixel points, a combination pair is obtained, and by analyzing the consistency of the sequence and combining the weight setting results of each behavior, it is possible to effectively determine whether the sequence difference meets the standard. Reasonable compensation and correction for abnormal combination pairs in advance ensures the accuracy of the obtained sequence and provides an accurate foundation for subsequent acquisition of three-dimensional pose.

[0122] This invention provides a three-dimensional human pose estimation method, which, based on the combination pairs corresponding to a second quantity, compensates and corrects abnormal combination pairs to obtain qualified combination pairs, including:

[0123] Determine the distribution of the first position points of the combination pairs corresponding to the second quantity;

[0124] Determine the distribution of the second position points of the combination pairs corresponding to the first quantity, and determine the positional relationship between the second position point distribution and the first position point distribution;

[0125] When the positions are intersecting, the abnormal combination pair itself is compensated for by the connecting line between the normal combination pair adjacent to the abnormal combination pair, and the intersecting line is obtained.

[0126] When the position is non-intersecting, the abnormal combination pair is contour compensated according to the behavioral contour of the matched first template and the behavioral contour of the second template to obtain the compensation line.

[0127] Based on the cross lines and compensation lines, qualified combination pairs are obtained.

[0128] In this embodiment, the location point distribution is obtained based on the contour position of different points. Therefore, the distribution of qualified combination pairs and the distribution of abnormal combination pairs are obtained.

[0129] In this embodiment, mutual intersection means that the first position point and the second position point intersect each other, and non-intersection means that the first position point and the second position point do not intersect. For example, if position points 10, 11, 12 and position points 21, 22, 23 are considered to not intersect, then position points 10, 21, 22, 11, 23, 12 are considered to intersect.

[0130] In this embodiment, the connecting line refers to the line connecting adjacent normal pairs. The midpoint of the connecting line is used to supplement abnormal pairs. If the midpoint is to the left of the corresponding point of the abnormal pair, the corresponding point of the abnormal pair is moved to the left, but not beyond the midpoint. The same applies to the right.

[0131] In this embodiment, contour compensation involves using the acquired behavioral contours to perform line supplementation on abnormal combination pairs.

[0132] The beneficial effects of the above technical solution are: by obtaining the positional distribution of normal and abnormal combination pairs, and by assessing whether they intersect or not, line compensation or contour compensation is performed on the abnormal combination pairs, thereby obtaining qualified combination pairs and providing accuracy for subsequent attitude analysis.

[0133] This invention provides a three-dimensional human pose estimation method, which, based on a human motion sequence, constructs the three-dimensional human pose of the target human body, including:

[0134] The human motion sequence is input into the 3D generation model in sequence to obtain the 3D human posture.

[0135] In this embodiment, the 3D generation model is pre-trained, and the training samples of the model include different 3D poses and motion sequences that match the 3D poses. Therefore, after inputting the human motion sequence into the model, a 3D human pose can be obtained.

[0136] The beneficial effects of the above technical solution are: by analyzing the sequence through the model, the three-dimensional human posture can be effectively obtained, ensuring the accuracy of the posture.

[0137] This invention provides a three-dimensional human pose estimation system, such as... Figure 2 As shown, it includes:

[0138] The human body acquisition module is used to acquire the target human body from a horizontal perspective and from a vertical perspective based on dual depth cameras configured at the acquisition point, wherein the horizontal and vertical perspectives are related to the height of the target human body.

[0139] The sequence generation module is used to determine the pose of the first and second acquired images based on a pre-trained neural network and a generator composed of neural network units, and obtain a human motion sequence.

[0140] The posture composition module is used to construct the three-dimensional human posture of the target human body based on the human motion sequence.

[0141] The beneficial effects of the above technical solution are: by using dual depth cameras to acquire data horizontally and vertically, and by using a generator to determine the pose, a motion sequence is obtained, which in turn generates a three-dimensional human pose, effectively improving the accuracy of pose acquisition.

[0142] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A three-dimensional human pose estimation method, characterized in that, include: Step 1: Based on the dual depth cameras configured at the acquisition point, perform a first acquisition of the target human body from a horizontal perspective and a second acquisition of the target human body from a vertical perspective, wherein the horizontal and vertical perspectives are related to the height of the target human body. Step 2: Based on the pre-trained neural network and the generator composed of neural network units, perform pose determination on the first and second acquired images to obtain the human motion sequence; Step 3: Based on the human motion sequence, construct the three-dimensional human posture of the target human body; The three-dimensional human pose estimation method also includes: Obtain the network construction path of the pre-trained neural network and the unit construction path of the neural network unit; Based on the network construction lines and unit construction lines, determine the line matrix, and obtain the attitude estimation template contained in each matrix unit of the line matrix respectively; The attitude estimation template is parsed to obtain a patch group for each matrix unit, wherein the patch group contains several two-dimensional attitude structures; Perform structural overlap and non-overlap determination on all the patch groups in the matrix unit, and perform a first calibration on the overlapping two-dimensional structures and a second calibration on the non-overlapping two-dimensional structures in the same patch group; Based on the frequency of occurrence of the same first label and its importance in each piece group, a first weight is assigned to the corresponding overlapping two-dimensional structure in each piece group; Based on the importance of the corresponding piece groups, a second weight is assigned to the corresponding non-overlapping two-dimensional structure. The cell setting result is obtained based on the first weight set for the overlapping two-dimensional structures in the same matrix cell and the second weight set for the non-overlapping two-dimensional structures. The generator is obtained based on the settings of all units and the constructed circuit.

2. The three-dimensional human pose estimation method as described in claim 1, characterized in that, Based on dual depth cameras configured at the acquisition points, the target human body is subjected to a first acquisition of horizontal perspective and a second acquisition of vertical perspective, including: The first distance between the center position of the target human body and the dual-depth camera is obtained, and based on the first coordinate system of the dual-depth camera facing the center position, the straight line corresponding to the first distance is displayed on the first coordinate system, the first angle between the straight line corresponding to the first distance and the first coordinate system is obtained, and the horizontal and vertical viewing angles of the dual-depth camera are determined. When the dual-depth camera performs the first acquisition of the target human body from a horizontal perspective, the supplementary lighting device is controlled to provide the first supplementary lighting to the dual-depth camera based on the first acquisition state of the dual-depth camera and the first environmental state. When the dual-depth camera performs a second acquisition of the target human body from a vertical perspective, the supplementary lighting device is controlled to provide a second supplementary light to the dual-depth camera based on the second acquisition state of the dual-depth camera and the second environmental state. After acquiring the dual-depth camera based on the first supplementary light and the second supplementary light, a first acquired image and a second acquired image are obtained.

3. The three-dimensional human pose estimation method as described in claim 1, characterized in that, Pose determination is performed on the first and second acquired images to obtain a human motion sequence, including: Based on the generator, the first acquired image is analyzed for horizontal pixels to obtain a first pose sequence; Based on the generator, the second acquired image is analyzed vertically to obtain a second pose sequence; Based on the first posture sequence and the second posture sequence, a human motion sequence is obtained.

4. The three-dimensional human pose estimation method as described in claim 3, characterized in that, Based on the first posture sequence and the second posture sequence, a human motion sequence is obtained, including: The first pose sequence and the second pose sequence are combined at the same pixel point to obtain a combination pair of each pixel point; Determine whether the first sequence and the second sequence in the combination pair are the same. If they are the same, analyze the next combination pair. If they are inconsistent, the first template that matches the first pose sequence and the second template that matches the second pose sequence are matched based on the pose estimation template of the generator. Based on the weight setting results of each template behavior in the first template and the weight setting results of each template behavior in the second template, determine whether the sequence difference between the first sequence and the second sequence meets the setting criteria. If satisfied, analyze the next pair; If the conditions are not met, the combination pair is determined to be abnormal; Determine the first number of abnormal pairs and the second number of qualified pairs; When the first quantity is less than the second quantity, the abnormal combination pairs are compensated and corrected based on the combination pairs corresponding to the second quantity to obtain qualified combination pairs; The human motion sequence is obtained by combining the horizontal and vertical sequences of all qualified pairs.

5. The three-dimensional human pose estimation method as described in claim 4, characterized in that, Based on the combination pairs corresponding to the second quantity, the abnormal combination pairs are compensated and corrected to obtain qualified combination pairs, including: Determine the distribution of the first position points of the combination pairs corresponding to the second quantity; Determine the distribution of the second position points of the combination pairs corresponding to the first quantity, and determine the positional relationship between the second position point distribution and the first position point distribution; When the positions are intersecting, the abnormal combination pair itself is compensated for by the connecting line between the normal combination pair adjacent to the abnormal combination pair, and the intersecting line is obtained. When the position is non-intersecting, the abnormal combination pair is contour compensated according to the behavioral contour of the matched first template and the behavioral contour of the second template to obtain the compensation line. Based on the cross lines and compensation lines, qualified combination pairs are obtained.

6. The three-dimensional human pose estimation method as described in claim 1, characterized in that, Based on the human motion sequence, the three-dimensional human posture of the target human body is constituted, including: The human motion sequence is input into the 3D generation model in sequence to obtain the 3D human posture.

7. A three-dimensional human pose estimation system, characterized in that, include: The human body acquisition module is used to acquire the target human body from a horizontal perspective and from a vertical perspective based on dual depth cameras configured at the acquisition point, wherein the horizontal and vertical perspectives are related to the height of the target human body. The sequence generation module is used to determine the pose of the first and second acquired images based on a pre-trained neural network and a generator composed of neural network units, and obtain a human motion sequence. A pose composition module is used to construct the three-dimensional human pose of the target human body based on the human motion sequence. The three-dimensional human pose estimation system also includes a module configured to perform the following operations: Obtain the network construction path of the pre-trained neural network and the unit construction path of the neural network units; determine the path matrix based on the network construction path and the unit construction path, and obtain the pose estimation template contained in each matrix unit of the path matrix; parse the pose estimation template to obtain the patch group of each matrix unit, wherein the patch group contains several pose two-dimensional structures; perform structural overlap judgment and structural non-overlap judgment on the patch groups in all matrix units, and perform a first calibration on the overlapping two-dimensional structures in the same patch group and a second calibration on the non-overlapping two-dimensional structures; set a first weight on the corresponding overlapping two-dimensional structures in each patch group based on the occurrence frequency of the same first calibration and the importance in each patch group; set a second weight on the corresponding non-overlapping two-dimensional structures based on the importance of the corresponding patch group; obtain the unit setting result based on the first weight set on the overlapping two-dimensional structures in the same matrix unit and the second weight set on the non-overlapping two-dimensional structures; obtain the generator based on all unit setting results and the construction path.

Citation Information

Patent Citations

  • Method and apparatus for estimating three-dimensional human body posture

    CN104952105A

  • End-to-end multi-view three-dimensional human body posture estimation method and system and storage medium

    CN112560757A