A visual motion capture and 3D digital human skeleton initial pose automatic matching method, system, device, medium and product

CN119785424BActive Publication Date: 2026-10-09TIANXIANG RUIYI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411834581.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2026-10-09
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

[0004]1、为保证渲染端和动捕端的初始姿态一致,一般是手动逐个调整渲染端每个骨骼的初始姿态,使得与动捕端的初始姿态一致,这种方法费时费力且会有肉眼无法分辨的误差

Benefits of technology

[0036] This application provides a method, system, device, medium, and product for automatic matching of visual motion capture and the initial pose of 3D digital human skeleton. This application combines vision-based multi-view motion capture and rendering end, and uses algorithms to automatically identify the initial pose and coordinate system transformation of 3D digital on the rendering end in the motion capture environment, effectively solving the problem of inconsistency between the coordinate system and initial pose of the two, thereby realizing automatic matching of the initial pose of visual motion capture and 3D digital human skeleton.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785424B_ABST
    Figure CN119785424B_ABST
Patent Text Reader

Abstract

The application discloses a visual motion capture and 3D digital human skeleton initial posture automatic matching method, system, device, medium and product, and relates to the technical field of visual rendering. First, in the rendering environment, the posture data of each skeleton of a 3D digital human in a different view and the extrinsic parameters of a virtual camera for taking a photo in each view are acquired; in the motion capture environment, the posture data of each skeleton in the corresponding motion capture end skeleton coordinate system in the photo in each view is captured; then, the initial posture error of the motion capture end relative to the rendering end and the posture data of each skeleton in the motion capture end world coordinate system are calculated; and the conversion relationship between the motion capture end skeleton coordinate system and the rendering end skeleton coordinate system is determined. The application combines the visual multi-view motion capture and the rendering end, and realizes visual motion capture and 3D digital human skeleton initial posture automatic matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual rendering technology, and in particular to a method, system, device, medium and product for automatic matching of visual motion capture and initial pose of 3D digital human skeleton. Background Technology

[0002] Most 3D digital humans are currently rendered in real-time using environments such as Unreal Engine (UE) or Blender. The input is the rotation data of each bone relative to the initial pose. However, due to the different methods of generating digital humans, the default initial pose during production is not the standard T-pose or A-pose. Motion capture ends mostly use wearable or non-wearable visual capture, and their coordinate systems and initial bone poses also vary due to different algorithms. In order to make the two work together, some plugins or algorithms are needed to identify the coordinate system and pose of the digital human bones in rendering ends such as UE and Blender, and then manually adjust them to ensure that they are consistent with the coordinate system and initial pose of the motion capture end so that the two can work together.

[0003] Most existing skeleton recognition algorithms on the rendering side are integrated with the rendering environment as plugins, such as Mixamo in Unreal Engine and Autorig in Blender. The shortcomings of existing technologies include:

[0004] 1. To ensure that the initial poses of the rendering end and the motion capture end are consistent, the initial pose of each bone on the rendering end is usually adjusted manually to make it consistent with the initial pose of the motion capture end. This method is time-consuming and laborious and will have errors that are not discernible to the naked eye.

[0005] 2. The skeletal poses identified by the plugin are based on the rendering environment coordinate system, which is inconsistent with the coordinate system definition of the motion capture terminal. It is also necessary to calculate the coordinate system transformation between the two for each bone. This method is time-consuming, laborious, and has errors that are not discernible to the naked eye. Summary of the Invention

[0006] The purpose of this application is to provide a method, system, device, medium and product for automatic matching of the initial pose of visual motion capture and 3D digital human skeleton, so as to realize the automatic identification of the initial pose error of the motion capture end relative to the rendering end and the transformation relationship between the skeleton coordinate system of the motion capture end and the skeleton coordinate system of the rendering end, thereby realizing the automatic matching of the initial pose of visual motion capture and 3D digital human skeleton.

[0007] To achieve the above objectives, this application provides the following solution.

[0008] In a first aspect, this application provides a method for automatic matching of visual motion capture and initial pose of a 3D digital human skeleton, including:

[0009] In the rendering environment, the 3D digital human is photographed in different views, the pose data of each bone of the 3D digital human in the world coordinate system of the rendering end, and the external parameters of the virtual camera that takes the photos in each view are obtained.

[0010] In the motion capture environment, based on the extrinsic parameters of the virtual camera that takes photos in each view, the pose data of each bone in the photo in each view is captured in its corresponding motion capture end bone coordinate system.

[0011] Based on the pose data of each bone in its corresponding motion capture end bone coordinate system, calculate the initial pose error of the motion capture end relative to the rendering end.

[0012] Based on the pose data of each bone in its corresponding motion capture end bone coordinate system, calculate the pose data of each bone in the motion capture end world coordinate system;

[0013] Based on the pose data of each bone in the motion capture terminal's world coordinate system and the pose data of each bone in the rendering terminal's world coordinate system, as well as the transformation relationship between the motion capture terminal's world coordinate system and the rendering terminal's world coordinate system, the transformation relationship between the motion capture terminal's bone coordinate system and the rendering terminal's bone coordinate system corresponding to each bone is determined.

[0014] Optionally, the formula for calculating the initial pose error of the motion capture end relative to the rendering end is:

[0015] T i,PoseError =inverse(T i,lm );

[0016] Among them, T i,PoseError T represents the initial pose error of skeleton i relative to the rendering end. i,lm Let i be the pose data of skeleton i in the corresponding motion capture end skeleton coordinate system, and inverse() is the inverse function.

[0017] Optionally, the transformation relationship between the motion capture end's bone coordinate system and the rendering end's bone coordinate system is as follows:

[0018] T i,Coordinate =inverse(T i,wr )×T rm ×T i,wm ;

[0019] Among them, T i,Coordinate T represents the transformation relationship between the motion capture end bone coordinate system and the rendering end bone coordinate system corresponding to bone i. i,wr T represents the pose data of skeleton i in the world coordinate system of the rendering end. rm To establish the transformation relationship between the rendering endpoint's world coordinate system and the motion capture endpoint's world coordinate system, T i,wmThe pose data of skeleton i in the world coordinate system of the motion capture end.

[0020] Optionally, in the motion capture environment, based on the extrinsic parameters of the virtual camera that captured the photos in each view, the pose data of each bone in the photos in each view is captured in its corresponding motion capture end bone coordinate system, specifically including:

[0021] In the motion capture environment, the extrinsic parameters of the multi-view visual motion capture module in the motion capture environment are set according to the extrinsic parameters of the virtual camera in each view, and the pose data of each bone in the photo in each view is captured in the corresponding motion capture end bone coordinate system.

[0022] Secondly, this application provides a visual motion capture and 3D digital human skeleton initial pose automatic matching system. The visual motion capture and 3D digital human skeleton initial pose automatic matching system is used to execute the above-mentioned visual motion capture and 3D digital human skeleton initial pose automatic matching method. The visual motion capture and 3D digital human skeleton initial pose automatic matching system includes: a rendering module and a motion capture module.

[0023] The rendering module includes a virtual camera submodule, a first pose acquisition submodule, and a data transmission submodule; the motion capture module includes a second pose acquisition submodule and a data processing submodule.

[0024] The virtual camera submodule is used to acquire photos of the 3D digital human in different views within the rendering environment;

[0025] The first pose acquisition submodule is used to acquire pose data of each skeleton of the 3D digital human in the world coordinate system of the rendering end in the rendering environment;

[0026] The data transmission submodule is used to transmit photos of the 3D digital human in different views in the rendering environment, the posture data of each bone of the 3D digital human in the world coordinate system of the rendering end, and the external parameters of the virtual camera that takes photos in each view to the motion capture module.

[0027] The second pose acquisition submodule is used to capture the pose data of each bone in the photo of each view in the corresponding motion capture end bone coordinate system according to the extrinsic parameters of the virtual camera that takes the photo in each view in the motion capture environment.

[0028] The data processing submodule is used to calculate the initial pose error of the motion capture end relative to the rendering end based on the pose data of each bone in its corresponding motion capture end bone coordinate system; to calculate the pose data of each bone in the motion capture end world coordinate system based on the pose data of each bone in its corresponding motion capture end bone coordinate system; and to determine the transformation relationship between the motion capture end bone coordinate system and the rendering end bone coordinate system based on the pose data of each bone in the motion capture end world coordinate system and the pose data of each bone in the rendering end world coordinate system, as well as the transformation relationship between the motion capture end world coordinate system and the rendering end world coordinate system.

[0029] Optionally, the virtual camera submodule includes multiple virtual cameras;

[0030] In the rendering environment, multiple virtual cameras are set up around the 3D digital human.

[0031] Optionally, in the motion capture environment, a multi-view visual motion capture module is set up, and during motion capture, the extrinsic parameters of the multi-view visual motion capture module are set to be consistent with the extrinsic parameters of each virtual camera.

[0032] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for automatic matching of visual motion capture and initial pose of 3D digital human skeleton.

[0033] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for automatic matching of visual motion capture and initial pose of 3D digital human skeleton.

[0034] Fifthly, this application provides a computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the above-mentioned method for automatic matching of visual motion capture and initial pose of 3D digital human skeleton.

[0035] According to the specific embodiments provided in this application, this application has the following technical effects.

[0036] This application provides a method, system, device, medium, and product for automatic matching of visual motion capture and the initial pose of 3D digital human skeleton. This application combines vision-based multi-view motion capture and rendering end, and uses algorithms to automatically identify the initial pose and coordinate system transformation of 3D digital on the rendering end in the motion capture environment, effectively solving the problem of inconsistency between the coordinate system and initial pose of the two, thereby realizing automatic matching of the initial pose of visual motion capture and 3D digital human skeleton. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating a method for automatic initial pose matching of visual motion capture and 3D digital human skeleton provided in an embodiment of this application.

[0039] Figure 2 Example images of a 3D digital human in different views provided in an embodiment of this application.

[0040] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] In one exemplary embodiment, such as Figure 1 As shown, a method for automatic matching of visual motion capture and initial pose of 3D digital human skeleton is provided, including the following steps 101 to 105.

[0044] Step 101: In the rendering environment, acquire photos of the 3D digital human in different views, the pose data of each bone of the 3D digital human in the world coordinate system of the rendering end, and the extrinsic parameters of the virtual camera that took the photos in each view.

[0045] Step 102: In the motion capture environment, capture the pose data of each bone in the photo of each view in its corresponding motion capture end bone coordinate system according to the extrinsic parameters of the virtual camera that took the photo in each view.

[0046] Step 103: Calculate the initial pose error of the motion capture end relative to the rendering end based on the pose data of each bone in its corresponding motion capture end bone coordinate system.

[0047] Step 104: Calculate the pose data of each bone in the world coordinate system of the motion capture end based on the pose data of each bone in its corresponding motion capture end bone coordinate system.

[0048] Step 105: Based on the pose data of each bone in the world coordinate system of the motion capture terminal and the pose data of each bone in the world coordinate system of the rendering terminal, as well as the transformation relationship between the world coordinate system of the motion capture terminal and the world coordinate system of the rendering terminal, determine the transformation relationship between the motion capture terminal bone coordinate system and the rendering terminal bone coordinate system for each bone.

[0049] By implementing steps 101 to 105 above, the initial pose error of the motion capture end relative to the rendering end and the transformation relationship between the motion capture end skeleton coordinate system and the rendering end skeleton coordinate system are automatically identified, thereby achieving automatic matching of the initial pose of visual motion capture and 3D digital human skeleton.

[0050] In the embodiments of this application, the attitude data is represented in the form of a rotation matrix or a quaternion.

[0051] In another exemplary embodiment of this application, a ring of virtual cameras at different angles is set around the 3D digital human in the rendering environment to obtain photos of the 3D digital human from different views, such as... Figure 2 As shown, in this embodiment of the application, photos of the 3D digital human in six different views were obtained, labeled P1 to P6, and the extrinsic parameters of the virtual camera that took the photos in each view were labeled E1 to E6. Further, the pose of each skeleton of the 3D digital human in the world coordinate system of the rendering terminal was obtained, that is, the pose data of each skeleton of the 3D digital human in the world coordinate system of the rendering terminal, labeled T. i,wr In this context, the subscript i represents the bone sequence number, the subscript r represents the renderer, and the subscript w represents the world coordinate system.

[0052] To achieve the integration of the rendering environment and the motion capture environment, in this embodiment of the application, a data transmission submodule can be set up to transmit photos P1 to P6 and data containing extrinsic parameters E1 to E6, T... i,wr The parameter file is transferred to the motion capture environment.

[0053] In another exemplary embodiment of this application, in the motion capture environment, the extrinsic parameters of the visual motion capture modules at different perspectives in the motion capture environment are replaced with the extrinsic parameters E1 to E6 of the aforementioned virtual camera, and each visual motion capture module is activated to capture the pose data T of each skeleton of the 3D digital human in the images P1 to P6 in its corresponding motion capture end skeleton coordinate system. i,lmIn this embodiment, the subscript m represents the motion capture end, and the subscript l represents the local bone coordinate system. Each bone has a bone coordinate system. The bone coordinate system in the rendering environment is called the rendering end bone coordinate system, and the bone coordinate system in the motion capture environment is called the motion capture end bone coordinate system. The origin of this bone coordinate system can be the starting point of the bone, i.e., the connection point between the bone and its previous bone segment (the previous bone connected to the bone is called the parent bone). The origin of the bone coordinate system can also be the midpoint of the bone; there is no limitation here. In some embodiments, the center point of the hip bone can be set as the origin of the root bone's bone coordinate system. As a preferred option, the root bone's bone coordinate system can coincide with or not coincide with the world coordinate system of its environment, as long as its transformation relationship with the world coordinate system is clear; there is no limitation here.

[0054] In another exemplary embodiment of this application, the pose data of each bone in the photo of each view under its corresponding motion capture end bone coordinate system is captured. This can be obtained by simply setting specific parameters. For details, please refer to the paper "Graph-Based 3D Multi-Person Pose Estimation Using Multi-ViewImages".

[0055] In another exemplary embodiment of this application, after determining the pose data of each skeleton of the 3D digital human in the world coordinate system of the rendering end, and the pose data of each skeleton in the corresponding skeleton coordinate system of the motion capture end, steps 103-105 of this application are executed to automatically identify the initial pose error and transformation relationship, as follows:

[0056] The initial pose error of the motion capture end relative to the rendering end is calculated using the following formula.

[0057] T i,PoseError =inveerse(T i,ln )

[0058] Among them, T i,PoseError T represents the initial pose error of skeleton i relative to the rendering end. i,lm Let i be the pose data of skeleton i in the corresponding motion capture end skeleton coordinate system, and inverse() is the inverse function.

[0059] Calculate the pose T of each bone in the motion capture terminal in the world coordinate system of the motion capture terminal. i,wm In this context, the subscript 'w' represents the world coordinate system, such as the T coordinate system of skeleton 1. 1,wm =T 3,lm ×T 2,lm ×T 1,lmBone 2 is the previous bone of bone 1, and bone 3 is the previous bone of bone 2. Bone 3 is the root bone. In this example, bone 2 is the parent bone of bone 1, and bone 3 is the parent bone of bone 2. The same applies to other bones. Examples will not be given here.

[0060] Since the world coordinate systems of the motion capture end and the rendering end are defined, (assuming the transformation relationship between the two is T) rm Then, the pose of the motion capture end skeleton in the rendering world coordinate system can be calculated: T rm ×T i,wm Therefore, the transformation relationship between the motion capture end bone coordinate system and the rendering end bone coordinate system can be calculated as follows:

[0061] T i,Coordinate =inverse(T i,wr )×T rm ×T i,wm

[0062] Among them, T i,Coordinate T represents the transformation relationship between the motion capture end bone coordinate system and the rendering end bone coordinate system corresponding to bone i. i,wr T represents the pose data of skeleton i in the world coordinate system of the rendering end. i,rn To establish the transformation relationship between the rendering endpoint's world coordinate system and the motion capture endpoint's world coordinate system, T i,wm The pose data of skeleton i in the world coordinate system of the motion capture end.

[0063] At this point, visual motion capture and automatic matching of the initial pose of the 3D digital human skeleton are completed.

[0064] Based on the same inventive concept, this application also provides a visual motion capture and 3D digital human skeleton initial pose automatic matching system for implementing the aforementioned visual motion capture and 3D digital human skeleton initial pose automatic matching method. The solution provided by this system is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the visual motion capture and 3D digital human skeleton initial pose automatic matching system provided below can be found in the limitations of the visual motion capture and 3D digital human skeleton initial pose automatic matching method described above, and will not be repeated here.

[0065] In an exemplary embodiment, a visual motion capture and 3D digital human skeleton initial pose automatic matching system is provided, including: a rendering module and a motion capture module; the rendering module includes: a virtual camera submodule, a first pose acquisition submodule and a data transmission submodule; the motion capture module includes: a second pose acquisition submodule and a data processing submodule; the virtual camera submodule and the first pose acquisition submodule are connected to the data transmission submodule, and both the data transmission submodule and the second pose acquisition submodule are connected to the data processing submodule.

[0066] The virtual camera submodule is used to acquire photos of the 3D digital human in different views within the rendering environment.

[0067] The first pose acquisition submodule is used to acquire pose data of each skeleton of the 3D digital human in the world coordinate system of the rendering end in the rendering environment.

[0068] The data transmission submodule is used to transmit photos of the 3D digital human in different views in the rendering environment, the posture data of each bone of the 3D digital human in the world coordinate system of the rendering end, and the extrinsic parameters of the virtual camera that took the photos in each view to the motion capture module.

[0069] The second pose acquisition submodule is used in the motion capture environment to capture the pose data of each bone in the photo of each view in its corresponding motion capture end bone coordinate system, based on the extrinsic parameters of the virtual camera that takes the photo in each view.

[0070] The data processing submodule is used to calculate the initial pose error of the motion capture end relative to the rendering end based on the pose data of each bone in its corresponding motion capture end bone coordinate system; to calculate the pose data of each bone in the motion capture end world coordinate system based on the pose data of each bone in its corresponding motion capture end bone coordinate system; and to determine the transformation relationship between the motion capture end bone coordinate system and the rendering end bone coordinate system based on the pose data of each bone in the motion capture end world coordinate system and the pose data of each bone in the rendering end world coordinate system, as well as the transformation relationship between the motion capture end world coordinate system and the rendering end world coordinate system.

[0071] In another exemplary embodiment, the virtual camera submodule described above includes multiple virtual cameras; in this embodiment, it includes six virtual cameras.

[0072] The motion capture module mentioned above is a multi-view visual motion capture module.

[0073] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for visual motion capture and automatic initial pose matching of a 3D digital human skeleton.

[0074] Those skilled in the art will understand that Figure 3 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0075] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0076] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0077] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0078] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0079] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0080] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for automatic matching of visual motion capture and initial pose of a 3D digital human skeleton, characterized in that, include: In the rendering environment, the 3D digital human is photographed in different views, the pose data of each bone of the 3D digital human in the world coordinate system of the rendering end, and the external parameters of the virtual camera that takes the photos in each view are obtained. In the motion capture environment, the extrinsic parameters of the multi-view visual motion capture module in the motion capture environment are set according to the extrinsic parameters of the virtual camera in each view, and the pose data of each bone in the photo in each view is captured in the corresponding motion capture end bone coordinate system. Based on the pose data of each bone in its corresponding motion capture end bone coordinate system, calculate the initial pose error of the motion capture end relative to the rendering end. Based on the pose data of each bone in its corresponding motion capture end bone coordinate system, calculate the pose data of each bone in the motion capture end world coordinate system; Based on the pose data of each bone in the motion capture terminal's world coordinate system and the pose data of each bone in the rendering terminal's world coordinate system, as well as the transformation relationship between the motion capture terminal's world coordinate system and the rendering terminal's world coordinate system, the transformation relationship between the motion capture terminal's bone coordinate system and the rendering terminal's bone coordinate system corresponding to each bone is determined.

2. The method for automatic matching of visual motion capture and initial pose of 3D digital human skeleton according to claim 1, characterized in that, The formula for calculating the initial pose error of the motion capture end relative to the rendering end is: ; in, The skeleton of the motion capture end relative to the rendering end The initial attitude error. For bones The pose data in the corresponding motion capture end skeleton coordinate system. To find the inverse function.

3. The method for automatic matching of visual motion capture and initial pose of 3D digital human skeleton according to claim 1, characterized in that, The transformation relationship between the motion capture end's bone coordinate system and the rendering end's bone coordinate system is as follows: ; in, For bones The transformation relationship between the motion capture end bone coordinate system and the rendering end bone coordinate system. For bones Pose data in the world coordinate system of the rendering end. To establish the transformation relationship between the rendering endpoint's world coordinate system and the motion capture endpoint's world coordinate system, For bones Attitude data in the world coordinate system of the motion capture terminal To find the inverse function.

4. A visual motion capture and 3D digital human skeleton initial pose automatic matching system, characterized in that, The visual motion capture and 3D digital human skeleton initial pose automatic matching system is used to execute the visual motion capture and 3D digital human skeleton initial pose automatic matching method according to any one of claims 1-3, and the visual motion capture and 3D digital human skeleton initial pose automatic matching system includes: a rendering module and a motion capture module; The rendering module includes a virtual camera submodule, a first pose acquisition submodule, and a data transmission submodule; the motion capture module includes a second pose acquisition submodule and a data processing submodule. The virtual camera submodule is used to acquire photos of the 3D digital human in different views within the rendering environment; The first pose acquisition submodule is used to acquire pose data of each skeleton of the 3D digital human in the world coordinate system of the rendering end in the rendering environment; The data transmission submodule is used to transmit photos of the 3D digital human in different views in the rendering environment, the posture data of each bone of the 3D digital human in the world coordinate system of the rendering end, and the external parameters of the virtual camera that takes photos in each view to the motion capture module. The second pose acquisition submodule is used to capture the pose data of each bone in the photo of each view in the corresponding motion capture end bone coordinate system according to the extrinsic parameters of the virtual camera that takes the photo in each view in the motion capture environment. The data processing submodule is used to calculate the initial pose error of the motion capture end relative to the rendering end based on the pose data of each bone in its corresponding motion capture end bone coordinate system; to calculate the pose data of each bone in the motion capture end world coordinate system based on the pose data of each bone in its corresponding motion capture end bone coordinate system; and to determine the transformation relationship between the motion capture end bone coordinate system and the rendering end bone coordinate system based on the pose data of each bone in the motion capture end world coordinate system and the pose data of each bone in the rendering end world coordinate system, as well as the transformation relationship between the motion capture end world coordinate system and the rendering end world coordinate system.

5. The visual motion capture and 3D digital human skeleton initial pose automatic matching system according to claim 4, characterized in that, The virtual camera submodule includes multiple virtual cameras; In the rendering environment, multiple virtual cameras are set up around the 3D digital human.

6. The visual motion capture and 3D digital human skeleton initial pose automatic matching system according to claim 5, characterized in that, In the motion capture environment, a multi-view visual motion capture module is set up. During motion capture, the extrinsic parameters of the multi-view visual motion capture module are set to be consistent with the extrinsic parameters of each virtual camera.

7. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the visual motion capture and 3D digital human skeleton initial pose automatic matching method according to any one of claims 1-3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the visual motion capture and 3D digital human skeleton initial pose automatic matching method as described in any one of claims 1-3.

9. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the visual motion capture and 3D digital human skeleton initial pose automatic matching method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Inertial motion capture pose transient calibration method and inertial motion capture system

    CN106648088A

  • Calibration method and device for limb motion capture, electronic equipment and storage medium

    CN111681281A