Two-dimensional-to-three-dimensional attitude skeleton reconstruction system and method thereof

By employing temporal prediction and occlusion semantic compensation methods, the stability and naturalness issues of 3D pose skeleton reconstruction under occlusion and complex motions are solved, enabling the generation of coherent and accurate 3D pose skeletons under occlusion conditions.

CN121074259APending Publication Date: 2025-12-05SQ TECH (SHANGHAI) CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511221947.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies for 3D pose skeleton reconstruction lack stability and naturalness when faced with occlusion, complex movements, or rapid changes in movement, and cannot effectively overcome depth ambiguity and drift problems caused by occlusion.

Method used

A method based on temporal prediction and occlusion semantic compensation is adopted. Through a two-dimensional feature acquisition module, an initial three-dimensional reconstruction module, a temporal prediction compensation module, and an occlusion semantic compensation module, a sequence learning model and an action database are used for posture correction and missing feature point compensation. Combined with biomechanical rationality correction and action smoothing interpolation, a stable and natural three-dimensional posture skeleton is generated.

Benefits of technology

It improves the stability and naturalness of 3D pose skeleton reconstruction, ensuring the generation of coherent and accurate 3D pose skeletons under occlusion and complex motion conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074259A_ABST
    Figure CN121074259A_ABST
Patent Text Reader

Abstract

The invention discloses a two-dimensional-to-three-dimensional attitude skeleton reconstruction system and method based on sequential prediction and occlusion semantic compensation, and the method comprises the steps: obtaining skeleton joint points from two-dimensional image data as a plurality of two-dimensional feature points, and building an initial three-dimensional attitude skeleton according to a skeleton model and a projection inversion technology; thirdly, inputting historical frame data into the sequence learning model to output a predicted attitude, calculating a deviation value between the predicted attitude and an initial three-dimensional attitude skeleton, automatically correcting the stability of the three-dimensional attitude when the detected deviation exceeds a preset threshold value, and automatically correcting the stability of the three-dimensional attitude when two-dimensional feature point missing is found. And finding out the matched similar attitude to estimate and complement the missing feature points of the initial three-dimensional attitude skeleton, and immediately generating a complete and accurate final three-dimensional attitude skeleton, so as to achieve the technical effect of improving the stability and naturalness of the reconstruction of the three-dimensional attitude skeleton.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a two-dimensional to three-dimensional pose skeleton reconstruction system and method thereof, in particular, a two-dimensional to three-dimensional pose skeleton reconstruction system and method thereof based on temporal prediction and occlusion semantic compensation. BACKGROUND

[0002] In recent years, with the popularity and vigorous development of deep learning and computer vision technology, various intelligent image analysis applications have emerged like mushrooms after rain, among which, three-dimensional human pose estimation (3D Human Pose Estimation) is the most concerned and widely used in virtual reality, human-computer interaction, motion analysis and animation production fields.

[0003] Generally speaking, the traditional two-dimensional to three-dimensional pose skeleton reconstruction mostly uses single frame image processing, which estimates the joint points directly from a single two-dimensional image, and then "lifts" the two-dimensional coordinates to three-dimensional space through a neural network model. However, this method only relies on the static information of a single picture, and when the human body in the image is self-occluded or shielded by objects, it will lack context reference and cause depth ambiguity, resulting in incorrect or unreasonable three-dimensional skeleton. Therefore, the traditional single frame reconstruction method has the problem of insufficient stability and accuracy under occlusion.

[0004] Therefore, some manufacturers have proposed a technical means using temporal information, which analyzes the motion information of consecutive image frames to enforce temporal consistency and smoothness. However, this method still cannot fully overcome the problems of drift and cumulative error caused by occlusion, complex motion or rapid changes in motion, etc., resulting in unstable and low naturalness of skeleton reconstruction, so it is still not enough to solve the problems of stability and naturalness of three-dimensional pose skeleton reconstruction.

[0005] In summary, it can be seen that there has been a problem of poor stability and naturalness of three-dimensional pose skeleton reconstruction in the prior art for a long time, so it is necessary to propose an improved technical means to solve this problem. SUMMARY

[0006] The present application discloses a two-dimensional to three-dimensional pose skeleton reconstruction system and method thereof based on temporal prediction and occlusion semantic compensation.

[0007] First, the application discloses a two-dimensional to three-dimensional pose skeleton reconstruction system based on time sequence prediction and occlusion semantic compensation, which comprises a non-transitory computer readable storage medium and a hardware processor. The non-transitory computer readable storage medium is used to store a plurality of instructions; the hardware processor is electrically connected to the non-transitory computer readable storage medium to execute the instructions to realize: a two-dimensional feature acquisition module is used to receive two-dimensional image data, and execute a pose estimation algorithm to obtain skeleton joint nodes from the two-dimensional image data as a plurality of two-dimensional feature points; an initial three-dimensional reconstruction module connected to the two-dimensional feature acquisition module is used to establish an initial three-dimensional pose skeleton based on the obtained two-dimensional feature points according to a skeleton model and a projection inversion technology; and a core prediction compensation module connected to the initial three-dimensional reconstruction module is used to compensate and correct the initial three-dimensional pose skeleton, the core prediction compensation module comprises: a time sequence prediction compensation submodule is used to input historical frame data into a sequence learning model to output a predicted pose, and then calculate the deviation value between the predicted pose and the initial three-dimensional pose skeleton, when the deviation value exceeds a preset threshold, the stability of the motion pose of the initial three-dimensional pose skeleton is corrected based on the predicted pose; and an occlusion semantic compensation submodule connected to the time sequence prediction compensation submodule is used to detect missing feature points in the two-dimensional feature points, and then query a similar pose matching the remaining two-dimensional feature points in an action database according to any one or a combination of Euclidean distance, cosine similarity and multi-joint weighted similarity, and estimate the missing feature points based on the similar pose, and compensate the initial three-dimensional pose skeleton in real time to generate a final three-dimensional pose skeleton, wherein the action database comprises feature point samples of a plurality of motion poses for comparison.

[0008] In addition, the application also discloses a two-dimensional to three-dimensional pose skeleton reconstruction method based on time sequence prediction and occlusion semantic compensation, which comprises the following steps executed by a hardware processor: receiving two-dimensional image data, and executing a pose estimation algorithm to obtain skeleton joint nodes from the two-dimensional image data as a plurality of two-dimensional feature points; establishing an initial three-dimensional pose skeleton based on the obtained two-dimensional feature points according to a skeleton model and a projection inversion technology; inputting historical frame data into a sequence learning model to output a predicted pose, and then calculating the deviation value between the predicted pose and the initial three-dimensional pose skeleton, when the deviation value exceeds a preset threshold, the stability of the motion pose of the initial three-dimensional pose skeleton is corrected based on the predicted pose; and when missing feature points are detected in the two-dimensional feature points, a similar pose matching the remaining two-dimensional feature points is queried in an action database according to any one or a combination of Euclidean distance, cosine similarity and multi-joint weighted similarity, and the missing feature points are estimated based on the similar pose, and the initial three-dimensional pose skeleton is compensated in real time to generate a final three-dimensional pose skeleton, wherein the action database comprises feature point samples of a plurality of motion poses for comparison.

[0009] The system and method disclosed in the present application differ from the prior art in that the present application obtains skeleton joints from two-dimensional image data as a plurality of two-dimensional feature points, and establishes an initial three-dimensional pose skeleton by a skeleton model and a projection inversion technique. Then, historical frame data is input into a sequence learning model to output a predicted pose, and a deviation value between the predicted pose and the initial three-dimensional pose skeleton is calculated. When the detected deviation exceeds a preset threshold, the three-dimensional pose stability is automatically corrected, and when a two-dimensional feature point is missing, a matching similar pose is found to estimate and supplement the missing feature points of the initial three-dimensional pose skeleton, so as to generate a complete and accurate final three-dimensional pose skeleton.

[0010] Through the above technical means, the present application can achieve the technical effect of improving the stability and naturalness of three-dimensional pose skeleton reconstruction. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 The system block diagram of the two-dimensional to three-dimensional pose skeleton reconstruction system based on time sequence prediction and occlusion semantic compensation of the present application.

[0012] Figure 2A And Figure 2B The method flowchart of the two-dimensional to three-dimensional pose skeleton reconstruction method based on time sequence prediction and occlusion semantic compensation of the present application.

[0013] Figure 3A And Figure 3B The schematic diagram of the interactive motion game to which the present application is applied.

[0014] Figure 4 The schematic diagram of the hardware architecture of the present application.

[0015] BRIEF DESCRIPTION OF REFERENCE NUMERALS:

[0016] 110: non-transitory computer readable storage medium

[0017] 120: hardware processor

[0018] 121: two-dimensional feature acquisition module

[0019] 122: initial three-dimensional reconstruction module

[0020] 123: core prediction compensation module

[0021] 124: time sequence prediction compensation submodule

[0022] 125: occlusion semantic compensation submodule

[0023] 131: biomechanical rationality correction module

[0024] 132: motion smoothing interpolation module

[0025] 133: three-dimensional motion output module

[0026] 300: two-dimensional-to-three-dimensional pose skeleton reconstruction system

[0027] 301: user

[0028] 301a: arm

[0029] 302a: dashed trajectory

[0030] 302b: solid trajectory

[0031] 310: camera device

[0032] 320: virtual character

[0033] 330: display device

[0034] 400: hardware architecture

[0035] 411: hardware processor

[0036] 412: memory

[0037] 413: network interface device

[0038] 414: network

[0039] 415: display device

[0040] 416: graphical user interface

[0041] 417: input device

[0042] 418: data storage device

[0043] 419: bus

[0044] Step 210: receiving two-dimensional image data and performing a pose estimation algorithm to obtain skeleton joints as a plurality of two-dimensional feature points from the two-dimensional image data

[0045] Step 220: establishing an initial three-dimensional pose skeleton based on the obtained two-dimensional feature points according to a skeleton model and a projection inversion technique

[0046] Step 230: inputting historical frame data into a sequence learning model to output a predicted pose, calculating a deviation value between the predicted pose and the initial three-dimensional pose skeleton, and performing stability correction on a motion pose of the initial three-dimensional pose skeleton based on the predicted pose when the deviation value exceeds a preset threshold

[0047] Step 240: When at least one missing feature point is detected in the two-dimensional feature points, at least one similar pose matching the remaining two-dimensional feature points is queried in the action database in any one or a combination of the Euclidean distance, cosine similarity, and multi-joint weighted similarity, and the missing feature point is estimated with the similar pose, and the initial three-dimensional pose skeleton is compensated in real time to generate a final three-dimensional pose skeleton, wherein the action database contains feature point samples of various action poses for comparison

[0048] Step 251: According to the human biomechanics constraint condition, the final three-dimensional pose skeleton is corrected to conform to the reasonable joint angle and position of the human structure and motion range

[0049] Step 252: Perform multi-time frame data interpolation on the corrected final three-dimensional pose skeleton, which estimates the feature data at the intermediate time between two adjacent time frames according to the multi-pen feature data at the continuous time points to smooth the time sequence change

[0050] Step 253: Output the smooth and coherent final three-dimensional pose skeleton to the backend system, which includes an animation engine, an external device, and a storage medium DETAILED DESCRIPTION

[0051] The embodiments of the present application will be described in detail below with reference to the drawings and examples, so that the implementation process of how the present application applies technical means to solve technical problems and achieve technical effects can be fully understood and implemented.

[0052] Please refer to Figure 1 , Figure 1 The system block diagram of the two-dimensional to three-dimensional pose skeleton reconstruction system based on temporal prediction and occlusion semantic compensation of the present application, the system includes: a non-transitory computer readable storage medium 110 and a hardware processor 120. Wherein, the non-transitory computer readable storage medium 110 is used to store a plurality of instructions. In actual implementation, the instructions can include special algorithm programs for image data processing, feature acquisition, three-dimensional reconstruction, temporal prediction and compensation, database query, and biomechanics correction, etc. These instructions can be implemented using a programming language and installed in a read-only memory, a hard disk, a solid state disk, a flash memory, or any other type of non-transitory computer readable storage medium 110.

[0053] Then, at the part of the hardware processor 120, it is electrically connected to the non-transitory computer readable storage medium 110 and executes instructions to implement: a two-dimensional feature acquisition module 121, an initial three-dimensional reconstruction module 122, and a core prediction compensation module 123. Among them, the two-dimensional feature acquisition module 121 is used to receive two-dimensional image data, and execute a pose estimation algorithm to obtain skeleton joint nodes from the two-dimensional image data as a plurality of two-dimensional feature points. In actual implementation, the two-dimensional image data can be a single static image or consecutive frames in a video, and the pose estimation algorithm can be an existing mature algorithm based on a convolutional neural network (CNN), such as “OpenPose”, “HRNet”, “CounterPropagation Network (CPN)”, etc., to identify and output two-dimensional pixel coordinates of each key joint node (such as head, neck, shoulder, elbow, wrist, hip, knee, ankle, etc.) of the human body from the image pixels, and output two-dimensional feature points with pixel positions.

[0054] The initial three-dimensional reconstruction module 122 is connected to the two-dimensional feature acquisition module 121, and is used to establish an initial three-dimensional pose skeleton based on the obtained two-dimensional feature points according to a skeleton model and a projection inversion technique. In actual implementation, the skeleton model predefines the topological structure of the human skeleton and the standard length ratio of each bone, and the projection inversion technique is a 2D-to-3D lifting method that can be implemented through a lightweight neural network. This neural network learns to regress the corresponding three-dimensional space coordinates from the input two-dimensional joint node coordinates, especially estimates the depth (i.e., Z-axis) information of each joint node, thereby establishing an initial three-dimensional pose skeleton with depth dimension. For example, a human statistical model (such as SMPL, H36M) can be used for modeling, and a perspective projection inversion or triangulation algorithm can be used to convert multi-view two-dimensional features into joint coordinates in three-dimensional space. At the same time, the “Random Sample Consensus (RANSAC)”, filtering method or similar function algorithm can be used to exclude outlier detection points.

[0055] The core prediction compensation module 123 is connected to the initial three-dimensional reconstruction module 122 to compensate and correct the initial three-dimensional pose skeleton. The core prediction compensation module 123 includes a temporal prediction compensation submodule 124 and an occlusion semantics compensation submodule 125. The temporal prediction compensation submodule 124 inputs historical frame data into a sequence learning model to output a predicted pose, and then calculates a deviation value between the predicted pose and the initial three-dimensional pose skeleton. When the deviation value exceeds a preset threshold, the motion pose of the initial three-dimensional pose skeleton is corrected based on the predicted pose. In actual implementation, the sequence learning model can be a time series network such as a Long Short-Term Memory (LSTM) neural network, a Gated Recurrent Unit (GRU), or a Transformer. The previous N frames of estimated three-dimensional pose data are used as model input to automatically predict the pose at the current time, and the initial three-dimensional result is dynamically corrected according to the deviation between the prediction and the initial three-dimensional result. In addition, the preset threshold can be initially set to twice the average standard deviation of all joints, and can be dynamically adjusted according to the performance of the training data and the validation data of the sequence learning model.

[0056] The occlusion semantics compensation submodule 125 is connected to the temporal prediction compensation submodule 124 to detect missing feature points in the two-dimensional feature points. When missing feature points are detected, one or a combination of Euclidean distance, cosine similarity, and multi-joint weighted similarity is selected to query a similar pose matching the remaining two-dimensional feature points in the action database, and the missing feature points are estimated based on the similar pose to generate a final three-dimensional pose skeleton. The action database includes feature point samples of various motion poses for comparison. In actual implementation, the method of estimating missing feature points based on similar poses can perform one or a combination of weighted average, nearest neighbor search (NNS), K nearest neighbor algorithm (KNN), and probability model based on the corresponding joint angles and positions in the matched similar poses. For example, the remaining two-dimensional feature points are matched with the nearest neighbors, and the corresponding feature points of the samples are used to estimate and fill in the missing feature points, such as the joint coordinates that are blocked or missing. In addition, the action database can be implemented using a dataset that collects joint features of various actions (for example, CMU Mocap). In another embodiment, when the corresponding joint is missing, the Euclidean distance between the N remaining feature points and all feature point samples in the action database is calculated, and the K groups (for example, K = 5) of similar poses with the smallest distance are selected to use their corresponding joint average coordinates as compensation basis.

[0057] In addition, the system of the present application can further comprise a biomechanics plausibility correction module 131 connected to the occlusion semantic compensation sub-module 125, for correcting the final 3D pose skeleton according to human biomechanics constraints, so as to conform to reasonable joint angles and positions of human body structure and motion range. In actual implementation, the human biomechanics constraints can include maximum / minimum bending angles of each joint, joint length ratio, human body connection structure, range of motion, physiological limit, etc., and can refer to model specifications, and if the output result exceeds the legal range, it is automatically adjusted to the reasonable range. More specifically, the human biomechanics constraints can include: (1) bone length constancy, to ensure that the length of bones such as forearm and lower leg remains unchanged in continuous motion; (2) joint angle limitation, to prevent the generation of unreasonable poses such as reverse bending of the knee joint or exceeding the normal range of motion of the shoulder joint; (3) body symmetry constraint, to ensure that the motion of left and right limbs remains coordinated and symmetrical under certain actions (such as walking). For example, biomechanics plausibility can refer to data such as the ±3% variation of bone length recommended by the World Health Organization (WHO) / International Society of Biomechanics (ISB), and the physical limit table of main joint angles, and when exceeding the limit range, it is automatically modified to the nearest critical value.

[0058] Furthermore, the system of the present application can further comprise a motion smoothing interpolation module 132 connected to the biomechanics plausibility correction module 131, for performing multi-time sequence frame data interpolation on the corrected final 3D pose skeleton, the multi-time sequence frame data interpolation being estimated according to a plurality of feature data at consecutive time points to smooth the feature data at intermediate time between two adjacent time sequence frames. In actual implementation, curve interpolation methods such as "B-spline" and "Catmull-Rom" can be used, or weighted moving average (Weighted Moving Average, WMA) can be used for continuous frame smoothing operation, so that the motion of the final 3D pose skeleton is coherent and not jittery. In another embodiment, the multi-time sequence frame data interpolation can use linear interpolation (Linear Interpolation, LERP) to smooth the position trajectory of the joint, or use spherical linear interpolation (Spherical Linear Interpolation, SLERP) to process the rotation amount (represented by quaternion) of the joint, to ensure the continuity of the rotation change. In addition, digital signal processing techniques such as Kalman filter or Savitzky-Golay filter can also be applied to filter the joint trajectory of the entire motion sequence, to eliminate small jitter and improve the smoothness of the overall motion.

[0059] As mentioned above, the system of the present application can further comprise a three-dimensional motion output module 133 connected to the motion smoothing interpolation module 132 for outputting the final three-dimensional pose skeleton in smooth and coherent manner to a backend system, which can include an animation engine, an external device, and a storage medium. In actual implementation, the three-dimensional motion can be standardized to be output in a file format commonly used for 3D model, animation, and scene exchange, such as “BVH”, “FBX”, “glTF”, etc., or transmitted to a 3D engine (e.g. Unity, Unreal, etc.) through an application programming interface (API) for real-time animation demonstration, or read by a virtual character control module to drive the motion of a virtual character, and can also be output to a virtual reality (VR) device, a robot motion control unit, a motion physiological analysis instrument, or a data server for long-term storage and application.

[0060] It is particularly pointed out that, in actual implementation, each module of the present application can be partially or completely realized based on hardware, for example, can be realized through a hardware processor such as an integrated circuit chip, a system on chip (SoC), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), etc. The hardware processor executes a computer program realizing the present application, and the computer program is stored in a computer readable storage medium, i.e., a computer readable storage medium loaded with computer readable program instructions for enabling the hardware processor to realize various aspects of the present application. The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include a hard disk, a random access memory, a read-only memory, a flash memory, an optical disk, a floppy disk, and any suitable combination of the above. The computer readable storage medium used herein is not to be interpreted as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission media (for example, an optical signal propagating through an optical fiber cable), or an electrical signal transmitted through a wire. In addition, the computer readable program instructions described herein can be downloaded from the computer readable storage medium to each computing / processing device, or to an external computer device or an external storage device through a network such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, a fiber optic transmission, a wireless transmission, a router, a firewall, a switch, a hub, and / or a gateway. The network card or network interface in each computing / processing device receives the computer readable program instructions from the network and forwards the computer readable program instructions to the computer readable storage medium stored in each computing / processing device. The computer program instructions for performing the operations of the present application can be combination language instructions, instruction set architecture instructions, machine instructions, machine related instructions, microinstructions, firmware instructions, or original code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Common Lisp, Python, C++, Objective-C, Smalltalk, Delphi, Java, Swift, C#, Perl, Ruby, and PHP, and conventional procedural programming languages such as C language or similar programming languages.The computer program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.

[0061] Referring to Figure 2A and Figure 2B , Figure 2A and Figure 2B is a method flowchart of the present invention based on the time sequence prediction and occlusion semantic compensation two-dimensional to three-dimensional pose skeleton reconstruction method, the following steps are executed by a hardware processor: receiving two-dimensional image data, and executing a pose estimation algorithm to obtain skeleton joints from the two-dimensional image data as a plurality of two-dimensional feature points (step 210); based on the obtained two-dimensional feature points, an initial three-dimensional pose skeleton is established according to a skeleton model and a projection inversion technique (step 220); historical frame data is input into a sequence learning model to output a predicted pose, and a deviation value between the predicted pose and the initial three-dimensional pose skeleton is calculated, when the deviation value exceeds a preset threshold, the motion pose of the initial three-dimensional pose skeleton is corrected based on the predicted pose (step 230); and when a missing feature point is detected in the two-dimensional feature points, a similar pose matching the remaining two-dimensional feature points is queried in a motion database in any one or a combination of Euclidean distance, cosine similarity and multi-joint weighted similarity, and the missing feature point is estimated based on the similar pose, and the initial three-dimensional pose skeleton is compensated in real time to generate a final three-dimensional pose skeleton, wherein the motion database contains feature point samples of various motion poses for comparison (step 240). Through the above steps, the skeleton joints are obtained from the two-dimensional image data as a plurality of two-dimensional feature points, and the initial three-dimensional pose skeleton is established according to the skeleton model and the projection inversion technique. Then, the historical frame data is input into the sequence learning model to output the predicted pose, and the deviation value between the predicted pose and the initial three-dimensional pose skeleton is calculated, when the detected deviation exceeds the preset threshold, the three-dimensional pose stability is automatically corrected, and when the two-dimensional feature points are found to be missing, the matching similar pose is found to estimate and make up the missing feature points of the initial three-dimensional pose skeleton, and a complete and accurate final three-dimensional pose skeleton is generated in real time.

[0062] In addition, as shown in FIG. 2B, after step 240, the final 3D pose skeleton can be corrected according to human biomechanics constraints to comply with reasonable joint angles and positions of human structure and motion range (step 251); multi-time frame data interpolation is performed on the corrected final 3D pose skeleton, the multi-time frame data interpolation is to estimate feature data at intermediate time between two adjacent time frames according to multi-pen feature data at continuous time points to smooth time sequence change (step 252); and the smoothed and coherent final 3D pose skeleton is output to a backend system for use, the backend system includes an animation engine, an external device and a storage medium (step 253). In this way, the final 3D pose skeleton can be more consistent with human biomechanics and the motion pose can be smoother.

[0063] The following Figures 3A to 4 The following description is made by way of example with reference to the accompanying drawings, in which Figure 3A and Figure 3B , Figure 3A and Figure 3B are schematic diagrams of interactive motion games using the present application. In this embodiment, a user 301 can perform dance actions in front of a camera 310, which is a web camera for example, that acquires two-dimensional image data including the user 301 and transmits it to a two-dimensional-to-three-dimensional pose skeleton reconstruction system 300 using the present application. At this time, the two-dimensional-to-three-dimensional pose skeleton reconstruction system 300 can immediately process the acquired two-dimensional image data and reconstruct a three-dimensional pose skeleton corresponding to a virtual character 320 and display it on a display device 330, such as a television screen, so that the virtual character 320 can mimic the actions of the user 301 in synchronization. In this application, two common technical problems can be solved: first, as shown in FIG. 1A, when the user 301 makes a turn or a side turn, causing the arms 301a to be blocked by the body, the camera 310 will not be able to acquire complete hand two-dimensional feature points. At this time, the occlusion semantic compensation submodule 125 will be triggered, and according to the remaining visible joint points (such as shoulders, elbows, torso, etc.), similar dance poses will be matched in the action database, and the most likely position of the arms 301a in three-dimensional space will be accurately estimated, so that the actions of the virtual character 320 remain coherent and natural, avoiding the problem of arm breakage or abnormal instantaneous displacement due to occlusion. Second, as shown in FIG. 1B, when the user 301 performs a dance action, the camera 310 will not be able to acquire complete two-dimensional feature points of the user 301 due to the limitations of the camera angle, resulting in incomplete two-dimensional feature points. At this time, the incomplete two-dimensional feature points will be compensated by the incomplete two-dimensional feature point compensation submodule 130, and the most likely position of the missing feature points in three-dimensional space will be accurately estimated, so that the actions of the virtual character 320 remain coherent and natural, avoiding the problem of incomplete feature points due to incomplete two-dimensional feature points. Figure 3A Figure 3B ​As shown, when the user 301 performs a large-scale motion such as a quick hand waving, the two-dimensional image data acquired by the photographing device 310 can generate dynamic blur, resulting in unreasonable jitter of the path of the initial three-dimensional pose skeleton established by the initial three-dimensional reconstruction module 122 (as shown by the dashed trajectory 302a). At this time, the time-series prediction compensation submodule 124 can predict a relatively smooth trajectory that conforms to the kinematics according to the motion trend of the historical frames (as shown by the solid trajectory 302b), and calculate the deviation value between the dashed trajectory 302a and the solid trajectory 302b, and perform stability correction when the deviation value exceeds a preset threshold to present a smooth motion. It should be noted that when the occlusion semantic compensation submodule 125 completes compensation of missing feature points once, the high-quality, complete final three-dimensional pose skeleton generated thereby can also feed back and update the historical frame data used by the time-series prediction compensation submodule 124, thereby ensuring that the basis for time-series prediction is not contaminated by a single occlusion event, and embodying the synergistic complementarity between the two submodules, so that the final three-dimensional pose skeleton can still stably output natural three-dimensional motion in the case of severe occlusion, short-term shielding, or feature point detection error.

[0064] As shown, Figure 4 As shown, Figure 4 FIG. 4 is a schematic diagram of a hardware architecture of the present application. In the hardware architecture 400, a plurality of computer executable instructions (hereinafter referred to as instructions) are used to drive the machine to perform any one or more methods discussed in the present application. In other embodiments, the machine can be connected (for example, network connected) to other machines in a local area network, an intranet, an extranet, or the Internet. The machine can operate as a server or a client in a client-server network environment, or as a peer machine in a peer-to-peer network environment, and can also operate as a network device, a server, a router, a switch or a bridge, an event generator, a distributed node, a centralized system, or any machine capable of executing a set of instructions (whether sequential or otherwise) that specify actions to be taken by that machine. In addition, although only one machine is shown, the term "machine" should also be considered to include a collection of machines (for example, tablet computers) that can individually or jointly execute a set (or multiple sets) of instructions to perform any one or more methods discussed in the present application.

[0065] The hardware architecture 400, such as a desktop computer, a notebook computer, a tablet computer, a server, a smart phone, etc., includes a hardware processor 411, a memory 412 (such as a read-only memory, a flash memory, a dynamic random access memory, a non-volatile resistive random access memory, an embedded flash memory, or a ferroelectric random access memory (FeRAM)), a network interface device 413, a display device 415, an input device 417, and a data storage device 418 (which can include a fixed or removable computer-readable storage medium), which are in communication with each other by a bus 419.

[0066] The hardware processor 411 can be a microprocessor, a central processing unit (CPU), or a similar device. More specifically, the hardware processor 411 can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a general-purpose instruction set computing (GIPS) processor that implements another instruction set, or a GIPS processor that implements a combination of instruction sets. Further, the hardware processor 411 executes various software elements stored in the memory 412 to perform various functions for the hardware architecture 400. In one embodiment, these software elements include an operating system, a compiler element, and a communication module (or instruction set). The operating system includes various programs, instruction sets, software elements, and / or drivers that are used to control and manage general system tasks and facilitate communication between various hardware and software elements. A compiler is a computer program (or set of programs) that converts source code written in a programming language into another computer language (e.g., an object code). The communication module can communicate with other devices through the network interface device 413. The network interface device 413 interfaces with a network 414, such as a local area network, a wide area network, or a similar network, to communicate with other devices.

[0067] The memory 412 can store program codes and / or data for use by the hardware processor 411. The memory 412 can be implemented using random access memory (e.g., SRAM, DRAM, DDRAM), ROM, magnetic and / or optical storage, flash memory, or any combination thereof. The memory 412 can also include a transmission medium that carries information signals (which can be modulated signals with or without carrier wave) representing instructions or data.

[0068] The display device 415, such as a Liquid Crystal Display (LCD), LED, or Cathode Ray Tube (CRT), is connected to the computer system through the display port and the graphic chipset, and provides a Graphic User Interface (GUI) 416 for the user to operate. In actual implementation, the display device 415 can also be a touch screen or similar device with input and display functions. In addition, the input device 417 (such as a keyboard, mouse, touchpad, etc.) can provide user input instructions and generate trigger signals.

[0069] The data storage device 418 can include a machine-readable storage medium (or more specifically, a computer-readable storage medium) that stores one or more instruction sets embodying any one or more of the methods or functions described herein. The disclosed data storage mechanism can be implemented entirely or at least partially in the memory 412. The data storage device 418 and the memory 412 disclosed in the hardware architecture 400 can be configured to implement a data storage mechanism for performing the operations and steps discussed in the present application.

[0070] In one example, the hardware architecture 400 is an Internet of Things device that can be connected (e.g., networked) to other machines in a local area network, a wide area network, or any network. The hardware architecture 400 can be a distributed system that includes many interconnected computers. The computers can act as servers or clients in a client-server network environment, or as peer machines in a point-to-point (or distributed) network environment.

[0071] In summary, the difference between the present application and the prior art is that the skeleton joint is obtained from the two-dimensional image data as a plurality of two-dimensional feature points, and an initial three-dimensional pose skeleton is established by a skeleton model and a projection inversion technique. Then, historical frame data is input into a sequence learning model to output a predicted pose, and the deviation value between the predicted pose and the initial three-dimensional pose skeleton is calculated. When the detected deviation exceeds a preset threshold, the three-dimensional pose stability is automatically corrected, and when the two-dimensional feature points are missing, the matching similar pose is found to estimate and supplement the missing feature points of the initial three-dimensional pose skeleton, thereby generating a complete and accurate final three-dimensional pose skeleton. Through this technical means, the problems existing in the prior art can be solved, and the technical effect of improving the stability and naturalness of three-dimensional pose skeleton reconstruction is achieved.

[0072] Although the present application has been disclosed with the foregoing embodiments, the disclosure is not to be construed to be limited thereto, and any person skilled in the art can make some changes and modifications without departing from the spirit and scope of the present application, and therefore the patent protection scope of the present application shall be subject to the scope defined by the claims attached to the present specification.

Claims

1. A system for reconstructing a 3D pose skeleton from 2D images based on temporal prediction and occlusion semantic compensation, the system comprising: a non-transitory computer-readable storage medium storing instructions; and a hardware processor electrically connected to the non-transitory computer-readable storage medium and configured to execute the instructions to implement: a 2D feature extraction module configured to receive 2D image data and perform a pose estimation algorithm to obtain skeleton joint points from the 2D image data as a plurality of 2D feature points; and a core prediction compensation module connected to the initial 3D reconstruction module and configured to compensate and correct the initial 3D pose skeleton, the core prediction compensation module comprising: a temporal prediction compensation submodule configured to input historical frame data into a sequence learning model to output a predicted pose, calculate a deviation value between the predicted pose and the initial 3D pose skeleton, and when the deviation value exceeds a preset threshold, perform stability correction on a motion pose of the initial 3D pose skeleton based on the predicted pose; and an occlusion semantic compensation submodule connected to the temporal prediction compensation submodule and configured to, when at least one missing feature point is detected in the 2D feature points, query at least one similar pose matching the remaining 2D feature points in an action database according to a selected one or a combination of a Euclidean distance, a cosine similarity, and a multi-joint weighted similarity, and estimate the missing feature point based on the similar pose to compensate the initial 3D pose skeleton in real time to generate a final 3D pose skeleton, wherein the action database comprises feature point samples of a plurality of motion poses for comparison. 2.The system of claim 1, further comprising a biomechanics plausibility correction module connected to the occlusion semantic compensation submodule and configured to correct the final 3D pose skeleton according to human biomechanics constraints to conform to reasonable joint angles and positions of human structures and motion ranges. 3.The system of claim 2, further comprising a motion smoothing interpolation module connected to the biomechanics plausibility correction module and configured to perform multi-time sequence frame data interpolation on the corrected final 3D pose skeleton, the multi-time sequence frame data interpolation being based on a plurality of feature data at consecutive time points to estimate feature data at an intermediate time between two adjacent time sequence frames to smooth time sequence changes. 4.The system of claim 3, further comprising a 3D motion output module connected to the motion smoothing interpolation module and configured to output the smoothed and coherent final 3D pose skeleton to a backend system, the backend system comprising an animation engine, an external device, and a storage medium. An initial three-dimensional reconstruction module connected to the two-dimensional feature acquisition module is configured to establish an initial three-dimensional posture skeleton based on the acquired two-dimensional feature points according to a skeleton model and a projection inversion technique. ​ ​ ​ ​ ​ ​ ​ ​ 5. The system of claim 1, wherein the similar poses estimate the missing feature points by performing one or a combination of weighted average, nearest neighbor search, K- nearest neighbor algorithm, and probabilistic model based on the corresponding joint angles and positions in the matching similar poses.

6. A method of reconstructing a 3D pose skeleton from 2D images based on temporal prediction and occlusion semantics compensation, comprising the steps of: receiving 2D image data and performing a pose estimation algorithm to obtain skeleton joints as 2D feature points from the 2D image data; establishing an initial 3D pose skeleton based on the obtained 2D feature points according to a skeleton model and a projection inversion technique; inputting historical frame data into a sequence learning model to output a predicted pose, and calculating a deviation value between the predicted pose and the initial 3D pose skeleton, and when the deviation value exceeds a preset threshold, performing stability correction on the motion pose of the initial 3D pose skeleton based on the predicted pose; and the motion database comprises feature point samples of various motion poses for comparison.

7. The method of claim 6, further comprising correcting the final 3D pose skeleton according to human biomechanical constraints to conform to reasonable joint angles and positions of human structure and motion range.

8. The method of claim 7, further comprising performing multi-time sequence frame data interpolation on the corrected final 3D pose skeleton, wherein the multi-time sequence frame data interpolation estimates feature data at an intermediate time between two adjacent time sequences according to multi-pen feature data at consecutive time points to smooth the time sequence change.

9. The method of claim 8, further comprising outputting the smoothed and coherent final 3D pose skeleton to a backend system, wherein the backend system comprises an animation engine, an external device, and a storage medium. When at least one missing feature point is detected in the two-dimensional feature points, at least one similar pose matching the remaining two-dimensional feature points is queried in the action database according to optional one or a combination of Euclidean distance, cosine similarity and multi-joint weighted similarity, and the missing feature point is estimated according to the similar pose, and the initial three-dimensional pose skeleton is compensated in real time to generate a final three-dimensional pose skeleton, wherein, 10. The method of claim 6, wherein the similar poses estimate the missing feature points by performing one or a combination of weighted average, nearest neighbor search, K- nearest neighbor algorithm, and probabilistic model based on the corresponding joint angles and positions in the matching similar poses. ​ ​ ​ ​