Image-based three-dimensional mesh reconstruction method and system therefor
The method addresses the limitations of conventional 2D pose extraction by restoring 3D meshes through unit movement division, smoothing, and PCA, achieving accurate and natural human pose identification and improved mesh restoration.
Patent Information
- Application Number
- PCT/KR2024/018777
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-25
- Filing Date
- 2024-11-25
- Publication Date
- 2025-05-30
AI Technical Summary
Conventional 2D pose extraction methods are inadequate for accurately analyzing human behavior in videos due to size dependencies and inability to capture joint rotation information, leading to low accuracy in mesh restoration and unnatural pose sequences.
A method for restoring 3D meshes by dividing recognizable action sequences into unit movements, extracting 3D pose parameters, applying smoothing for temporal continuity, performing principal component analysis to reduce parameters, and restoring the 3D mesh using the reduced parameters.
This approach enables accurate and natural human pose identification, improving mesh restoration accuracy while ensuring continuous and well-connected action sequences, and allowing for precise behavioral similarity analysis and virtual space mapping.
Smart Images

Figure KR2024018777_30052025_PF_FP_ABST
Abstract
Description
Image-based 3D mesh restoration method and system therefor
[0001] The present invention relates to a method for restoring a 3D mesh based on an image and a system therefor, and more particularly, to a method for distinguishing a unit motion of a person from image content, extracting accurate 3D pose parameters from frames corresponding to the unit motion, and restoring the pose of the person into a 3D mesh, and a system therefor.
[0002] Recognizing or understanding human behavior is essential for analyzing and understanding video content featuring people, such as dramas, movies, entertainment, and sports. The first step toward recognizing or understanding these characters' actions is to track each person and extract their poses. Conventional pose extraction has primarily relied on 2D pose extraction, which displays joint positions (in two dimensions). This method suffers from the drawback that values vary significantly depending on the size of the person displayed on screen, and that joint rotation information cannot be captured.
[0003] To overcome these limitations, methods for extracting 3D poses from 2D images have recently been proposed. Accurately extracting 3D pose information is essential for recognizing and understanding human actions within a video, and the present invention addresses this issue.
[0004] The present invention aims to propose a method for increasing accuracy in 3D mesh restoration while ensuring that the action sequence of a person continues naturally.
[0005] Meanwhile, the technical problems of the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by those skilled in the art from the description below.
[0006] A 3D mesh restoration method according to the present invention is executed by a computing device having a central processing unit and a memory, and includes the steps of: (a) dividing a recognizable action sequence from the image into unit movements of a predetermined length; (b) extracting 3D pose parameters from each frame of the unit movements; (c) applying smoothing to the extracted 3D pose parameters so that they have temporal continuity; (d) performing principal component analysis (PCA) on the smoothed 3D pose parameters to reduce the parameters; and (e) restoring a 3D mesh using the reduced parameters.
[0007] Additionally, in the above method, the length of the unit operation can be defined by the user.
[0008] Additionally, in the above method, the 3D pose parameters may include 3D rotation information of multiple joints and n-dimensional body shape information.
[0009] In addition, in the above method, the smoothing applied in step (c) may utilize a filtering technique that takes temporal continuity into account.
[0010] In addition, the principal component analysis performed in step (d) in the above method may be characterized by reducing the dimension of the 3D pose parameter.
[0011] In addition, the method may further include, after step (e), a step of mapping the restored 3D mesh into a virtual space.
[0012] Additionally, the method may further include, after step (e), a step of analyzing behavioral similarity using the restored 3D mesh. In this case, the behavioral similarity analysis may include at least one of specific behavior detection and dance similarity comparison.
[0013] Additionally, the method may further include a step of extracting initial 3D pose parameters from a 2D image.
[0014] Meanwhile, a 3D mesh restoration system according to another embodiment of the present invention includes a central processing unit and a memory, wherein the central processing unit executes commands for executing a 3D mesh restoration method stored in the memory, wherein the 3D mesh restoration method may include: (a) dividing a recognizable action sequence from the image into unit movements of a predetermined length; (b) extracting 3D pose parameters from each frame of the unit movements; (c) applying smoothing to the extracted 3D pose parameters so as to have temporal continuity; (d) performing principal component analysis (PCA) on the smoothed 3D pose parameters to reduce the parameters; and (e) restoring the 3D mesh using the reduced parameters.
[0015] According to the present invention, there is an effect of being able to extract an accurate 3D pose of a person within an image.
[0016] In particular, according to the present invention, 3D mesh restoration is possible by taking temporal continuity into consideration, so that more accurate and natural human poses can be identified compared to the conventional method.
[0017] In addition, according to the present invention, the restored 3D mesh can be directly utilized in a virtual space, and there is an effect of enabling more accurate control of character movement in the virtual space.
[0018] Furthermore, the present invention enables accurate behavioral similarity analysis using restored 3D meshes, and further facilitates the identification of human behavior, such as when generating highlights of specific actions in sports footage or when removing or blurring scenes featuring specific actions (e.g., smoking, violence, etc.). Furthermore, the present invention enables dance similarity comparisons and can be further utilized for detecting abnormal behavior in work environments.
[0019] Meanwhile, the effects of the present invention are not limited to those mentioned above, and other technical effects not mentioned can be clearly understood by those skilled in the art from the description below.
[0020] Figure 1 conceptually illustrates a 3D mesh restoration method and system according to the present invention.
[0021] Figure 2 illustrates a 3D mesh restoration method according to one embodiment of the present invention step by step.
[0022] Figure 3 illustrates the process by which the action sequence of a person in a video is distinguished into unit movements and the process by which 3D pose parameters are extracted from frames.
[0023] Figure 4 shows examples of how restored 3D meshes are used in various fields.
[0024] FIG. 5 illustrates a 3D mesh restoration method according to another embodiment of the present invention.
[0025] Figure 6 illustrates how the unit movements of a person in an image are defined by user input.
[0026] The purpose, technical configuration, and resulting operational effects of the present invention will be more clearly understood through the following detailed description based on the drawings attached to the specification of the present invention. The following describes embodiments of the present invention in detail with reference to the attached drawings.
[0027] The embodiments disclosed herein should not be construed or used to limit the scope of the present invention. Those skilled in the art will readily appreciate that the descriptions herein, including the embodiments, have a wide range of applications. Therefore, any embodiments described in the detailed description of the present invention are intended to serve as illustrative examples to better illustrate the present invention and are not intended to limit the scope of the present invention to the embodiments.
[0028] The functional blocks depicted in the drawings and described below are merely examples of possible implementations. Other implementations may utilize other functional blocks without departing from the spirit and scope of the detailed description. Furthermore, while one or more functional blocks of the present invention are depicted as individual blocks, one or more of the functional blocks of the present invention may be a combination of various hardware and software configurations that perform the same function.
[0029] Additionally, the expression “including certain components” is an “open” expression, simply indicating the presence of those components, and should not be construed as excluding additional components.
[0030] Furthermore, when it is said that a component is “connected” or “connected” to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.
[0031]
[0032] FIG. 1 conceptually illustrates a 3D mesh restoration method according to the present invention. Referring to the drawing, the present invention relates to a system (100) that extracts a 3D mesh of a specific person as shown on the right side of the drawing by executing the 3D mesh restoration method on an arbitrary image.
[0033] Previously, a method for extracting 3D poses from 2D images was first proposed, and by extending it, methods for extracting 3D poses of people appearing in videos were proposed. Methods for extracting 3D poses from videos can be divided into methods for extracting 3D poses from all frames (2D) and methods for extracting 3D poses using the entire video. The former method extracts poses for each frame, so it restores meshes well but has the disadvantage of not being connected continuously, and the latter uses the entire video, so it restores meshes that are connected continuously but has the problem of low mesh restoration accuracy in each frame.
[0034] The present invention relates to a method for improving mesh restoration accuracy while enabling a character's action sequence to be continuously and well connected to solve these problems.
[0035]
[0036] Figure 2 illustrates each step of a 3D mesh restoration method according to the present invention.
[0037] Referring to the drawing, the 3D mesh restoration method first includes a step (S101) of recognizing a behavior sequence from an image and dividing the behavior sequence into multiple unit actions.
[0038] There may be various methods for recognizing a person's action sequence from an arbitrary image. As an example of a method for recognizing an action sequence, the system (100) can detect 2D joint points in consecutive frames, calculate the amount of change in these joint points between frames, and then recognize the start and end points of a specific motion, i.e., the start and end points of the action sequence, based on whether the amount of change exceeds a threshold.
[0039] Meanwhile, if a single action sequence is recognized from an image, the system (100) can perform an operation to divide and segment the action sequence into multiple unit actions. The operation to divide into unit actions can be performed in various ways. For example, a method to distinguish unit actions by detecting feature points (e.g., joint points, etc.) within an action sequence and analyzing the movement of these feature points (e.g., change in velocity or acceleration of the joint points), a method to distinguish unit actions by dividing the action sequence into regular time intervals, a method to distinguish unit actions based on the time of detection of a specific event (e.g., a goal event in a soccer game, etc.) according to a predetermined rule, etc. can be utilized.
[0040] For reference, in the process of recognizing the above action sequence or dividing the unit actions, the system (100) can utilize an artificial intelligence algorithm, and can train the above artificial intelligence algorithm to automatically distinguish the action sequence or unit actions through learning data.
[0041] After step S101, the system (100) can extract 3D pose parameters from the frame(s) corresponding to each unit motion for each unit motion (S103). One unit motion may correspond to multiple frames in an image, and the system (100) can extract 3D pose parameters from the frames corresponding to each unit motion. The 3D pose parameters refer to a series of variables used to express a person's posture and movement in a three-dimensional space, and the 3D pose parameters in the present invention may include rotation information of multiple (e.g., 24 predetermined joints) joints and body shape information (e.g., 10-dimensional body shape information including the length and width of the torso, etc.).
[0042] For reference, FIG. 3 illustrates how, when a specific person's action sequence (10) is recognized in a video, three unit actions (unit actions 1 to 3) are divided from this action sequence (10). Referring to the drawing, when a scene of a specific person moving is included in the video, the continuous movement behavior can be recognized as one action sequence (10) by the system (100). In addition, if the person takes a walking motion, a running motion, or a crawling motion when moving, the system (100) can distinguish each motion into a unit action as described above, and 3D pose parameters can be extracted from frames corresponding to each unit action.
[0043] After step S103, the system (100) can perform a smoothing operation (S105) on the previously extracted 3D pose parameters. The smoothing step is a step to improve the continuity of the mesh by smoothly connecting the 3D pose parameters in time. In general, when 3D poses are extracted for each frame in an image, the poses of each frame are calculated independently, so that the pose changes between frames may appear abruptly, which may cause the quality of the 3D mesh to be ultimately restored. The above smoothing step is performed to solve this problem, and can be understood as a step to adjust the pose of each frame by considering the 3D pose parameters of temporally adjacent frames. For reference, filters such as a moving average filter and a Gaussian filter that adjust the 3D pose parameters of each frame to match the values of adjacent frames can be utilized in the smoothing step.
[0044] After step S105, the system (100) can reduce the number of parameters required to ultimately restore the 3D mesh by performing principal component analysis (S107) on the smoothed 3D pose parameters. That is, this step is intended to reduce the amount of 3D mesh restoration computation by expressing high-dimensional data in low dimensions, and also to improve the computational performance of the system (100) by removing relatively unimportant parameters. Meanwhile, the principal component analysis mentioned in this detailed description is a known analysis method, so a detailed description will be omitted, but preferably, it may include a step of preprocessing the smoothed 3D pose parameters to convert them into a form (e.g., matrix) that allows principal component analysis, a step of standardizing 3D pose parameter values, a step of calculating a covariance matrix of the standardized data matrix, a step of decomposing the covariance matrix to obtain principal components (eigenvalues and eigenvectors), and a step of selecting principal components in the order of increasing eigenvalues to reduce the dimensionality.
[0045] Finally, the system (100) can perform an operation to restore a 3D mesh by utilizing the dimensionally reduced 3D pose parameters (S109). There may be various methods for the operation to restore a 3D mesh, but step S109 in the present invention may preferably basically include a step of i) loading a template mesh to be used for 3D mesh restoration, and ii) deforming vertices of the template mesh by using the 3D pose parameters, and may additionally further include a texture mapping step.
[0046] Referring to Figure 2 above, a representative embodiment of a 3D mesh restoration method according to the present invention was examined.
[0047]
[0048] Meanwhile, FIG. 4 is intended to show that a 3D mesh extracted and restored from an image can be utilized in various fields. Referring to the drawing, the restored 3D mesh can be mapped to a virtual space and utilized, and can also be utilized to analyze the behavioral similarity of a person or to create a highlight video of players filmed at a sports game.
[0049] Regarding its use in virtual space, the 3D mesh extracted from a video of the user can be used to capture the user's real-time movements and implement the movements of a virtual character in the virtual space. Specifically, by receiving the video of the user as input, reconstructing a 3D mesh, and mapping it to a character in the virtual space, the user's actual movements can be directly reflected as character movements in the virtual space. This virtual space application can be applied to displaying the movements of a user avatar on a metaverse platform or realistically implementing character movements in online games.
[0050] In relation to an embodiment of analyzing the behavioral similarity of a person, a 3D mesh restored from a video can be utilized for comparative analysis with other references (e.g., a 3D mesh extracted and restored from the movements of another user), and for example, can be utilized for comparing a dancer's dance moves, analyzing a baseball player's swing mechanism, etc. In another embodiment, a 3D mesh restored from a video can be compared with a predetermined reference (e.g., a cigarette smoking motion, etc.) to determine whether a motion should be deleted from the video, and this utilization can be applied to editing out unnecessary scenes or controversial people from the video.
[0051] Additionally, the 3D mesh restored from the video can be utilized to automatically generate highlight videos of specific players or specific events (e.g., a goal scene in a soccer game) within a sports video. For example, the system (100) can receive a soccer game video, restore a 3D mesh of a player performing a specific movement, and generate a highlight video by editing only the frames in the section containing the movement.
[0052]
[0053] Fig. 5 illustrates a 3D mesh restoration method according to another embodiment of the present invention. The embodiment illustrated in Fig. 5 is basically similar to the steps described in Fig. 2, but differs in the step of dividing unit motions (S201) and the step of extracting 3D pose parameters from the frames of unit motions (S203).
[0054] First, regarding step S201, the system (100) may perform action sequence recognition and unit motion segmentation operations on its own when an image is input, but may also be implemented to segment unit motions based on input from a user using the system (100). In the case of unit motions segmented on its own by the system (100), the segmentation results may be inaccurate due to various reasons, and to compensate for this, the system (100) may be implemented to reflect the user's input during the unit motion segmentation process.
[0055] FIG. 6 illustrates an example of receiving a drag input from a user to set a specific unit action. Referring to the drawing, the system (100) can output to the user an arbitrary action sequence (10) recognized from an image, as shown in the drawing, and the user can set a specific action as a unit action by making a drag input using the mouse while viewing this screen. The screen provided to the user can display unit actions that the system (100) has independently classified, and the user can set the corresponding unit actions as a single unit action by grouping multiple unit actions into a single drag input in this state. In other words, two different unit actions that were divided by the system (100)'s own classification can be set as a single unit action by the user.
[0056] Next, regarding step S203, the system (100) extracts 3D pose parameters from frames of a unit motion (S203). More specifically, the system (100) may be implemented to extract initial 3D pose parameters from a 2D image (S2031), and then extract 3D pose parameters from frames of the unit motion (S2032). That is, by having the system (100) extract initial 3D pose parameters, i.e., parameters that serve as reference points, from a 2D image, the 3D pose parameters can be more accurately acquired. The initial 3D pose parameters are preferably extracted from the first frame (2D image) among the frames corresponding to the unit motion, but are not necessarily limited thereto. In some cases, the system (100) may be implemented to select the clearest frame among the frames corresponding to the unit motion and extract the initial 3D pose parameters from the corresponding frame. That is, the system (100) according to the present invention extracts 3D pose parameters from frames corresponding to a unit motion, and extracts the initial 3D pose parameters from a 2D image and the 3D pose parameters from the remaining frames separately, thereby increasing the accuracy of the 3D pose parameters and maintaining the continuity of the motion.
[0057] The 3D mesh reconstruction method and system therefor according to the present invention have been described above. Meanwhile, the present invention is not limited to the specific embodiments and applications described above, and various modifications may be made by those skilled in the art without departing from the gist of the present invention as claimed in the claims. Furthermore, such modifications should not be understood as distinct from the technical spirit or scope of the present invention.
Claims
1. A method for restoring a 3D mesh from an image by a computing device having a central processing unit and a memory, (a) a step of dividing a recognizable action sequence from the above image into unit movements of a certain length; (b) a step of extracting 3D pose parameters from each frame of the above unit motion; (c) a step of applying smoothing to the extracted 3D pose parameters to ensure temporal continuity; (d) a step of reducing parameters by performing principal component analysis (PCA) on the smoothed 3D pose parameters; and (e) a step of restoring a 3D mesh using the reduced parameters; Including, How to restore 3D mesh.
2. In paragraph 1, The length of the above unit operation is characterized by being definable by the user. How to restore 3D mesh.
3. In paragraph 1, The above 3D pose parameters include 3D rotation information of multiple joints and n-dimensional body shape information. How to restore 3D mesh.
4. In paragraph 1, The smoothing applied in the above step (c) is characterized by utilizing a filtering technique that takes temporal continuity into account. How to restore 3D mesh.
5. In paragraph 1, The principal component analysis performed in the above step (d) is characterized by reducing the dimension of the 3D pose parameters. How to restore 3D mesh.
6. In paragraph 1, After step (e) above, A step of mapping the restored 3D mesh into virtual space; Including more, How to restore 3D mesh.
7. In paragraph 1, After step (e) above, Step of analyzing behavioral similarity using the restored 3D mesh; Including more, How to restore 3D mesh.
8. In paragraph 7, The above action similarity analysis comprises at least one of specific action detection or dance similarity comparison. How to restore 3D mesh.
9. In paragraph 1, Step of extracting initial 3D pose parameters from 2D images; Including more, How to restore 3D mesh.
10. In a 3D mesh restoration system including a central processing unit and memory, The above central processing unit is characterized by executing commands for executing a 3D mesh restoration method stored in the memory, The above 3D mesh restoration method is, (a) a step of dividing a recognizable action sequence from the above image into unit movements of a certain length; (b) a step of extracting 3D pose parameters from each frame of the above unit motion; (c) a step of applying smoothing to the extracted 3D pose parameters to ensure temporal continuity; (d) a step of reducing parameters by performing principal component analysis (PCA) on the smoothed 3D pose parameters; and (e) a step of restoring a 3D mesh using the reduced parameters; 3D mesh restoration system.
Citation Information
Patent Citations
Method and Die Apparatus for die casting with additional forging effect
KR1020230039619A
Apparatus and method for 4d image reconstruction
KR102181832B1
3D Posed Estimation Apparatus and method based on Joint Interdependency
KR102198470B1
A skeleton-based dynamic point cloud estimation system for sequence compression
KR102577135B1
KR20220137410A