Motion capture method and system

By using neural network methods to capture motion from images acquired by cameras, the high cost and clothing dependence of traditional motion capture methods are solved, achieving efficient and accurate motion capture results.

CN114973404BActive Publication Date: 2025-09-05BEIJING JULI DIMENSION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210490709.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2025-09-05
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

Traditional motion capture methods rely on expensive optical equipment and specific clothing, and inertial motion capture has low accuracy, which cannot meet the needs of ordinary consumers.

Method used

By employing a neural network-based approach, images of motion capture personnel are captured through a camera, latent variables of clothing morphology features are extracted, and the motion capture network is trained. This enables motion capture without optical equipment or specific clothing, improving accuracy and generalization ability.

Benefits of technology

It achieves high-precision motion capture, reduces costs, and improves the efficiency and robustness of motion capture, making it suitable for ordinary consumers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973404B_ABST
    Figure CN114973404B_ABST
Patent Text Reader

Abstract

The present invention discloses a motion capture method and system. The method includes: collecting a set motion image and an image to be captured of a person performing motion capture; inputting the set motion image into a calibration feature network to obtain clothing morphological feature latent variables; inputting the clothing morphological feature latent variables into a motion capture network to train the motion capture network; and inputting the image to be captured into the motion capture network to obtain a pose estimation result. This method efficiently performs motion capture and improves the accuracy of motion capture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of motion capture technology, and specifically to a motion capture method and system. Background Art

[0002] Traditional motion capture methods use optical motion capture or inertial motion capture, which are used in film and television, virtual anchors, games and other fields. Optical capture is expensive and has extremely high requirements for the motion capture site. A lot of time needs to be spent on setting up and calibrating the motion capture site before motion capture. In addition, motion capture personnel need to wear specific motion capture clothing when performing motion capture. Price, cumbersome installation, and high clothing requirements are all major obstacles that prevent traditional optical motion capture from reaching ordinary consumers. Although inertial motion capture has significantly improved in price and installation compared to optical capture, it still has high clothing requirements for motion capture personnel. Motion capture personnel need to wear a certain number of sensors to successfully perform motion capture. Due to the defects of its own sensors, inertial motion capture also has the problem of low capture accuracy.

[0003] Therefore, traditional motion capture solutions cannot solve the dependence of motion capture personnel on clothing and at the same time ensure the accuracy of motion capture. Summary of the Invention

[0004] To this end, the embodiments of the present application provide a motion capture method and system that does not require the purchase of expensive traditional optical capture equipment or the wearing of inertial capture sensors. Motion capture can be achieved with only a camera, thereby improving the generalization ability and accuracy of motion capture.

[0005] In order to achieve the above objectives, the embodiments of the present application provide the following technical solutions:

[0006] According to a first aspect of an embodiment of the present application, a motion capture method is provided, the method comprising:

[0007] Collect the set action images and the images to be captured of the motion capture personnel;

[0008] Inputting the set action image into the calibration feature network to obtain clothing morphological feature latent variables;

[0009] Inputting the clothing morphological feature latent variables into a motion capture network to train the motion capture network;

[0010] The image to be motion captured is input into the motion capture network to obtain a posture estimation result.

[0011] Optionally, when collecting the set action image and the to-be-action-captured image of the motion capture personnel, the method further includes:

[0012] Collect random motion images of motion capture personnel;

[0013] Annotate the correspondence between images and joint rotations based on random motion images of motion capture personnel;

[0014] The annotation results are input into the motion capture network to train the motion capture network.

[0015] Optionally, inputting the clothing morphological feature latent variables into a motion capture network to train the motion capture network includes:

[0016] Performing a linear transformation on the clothing morphological feature latent variables to obtain a feature map; the feature map includes the clothing morphological features of the motion capture person and prior information of the motion capture environment;

[0017] The feature map is input into the motion capture network for training.

[0018] Optionally, performing a linear transformation on the clothing morphological feature latent variables to obtain a feature map includes:

[0019] Normalizing the clothing morphological feature latent variables;

[0020] The normalized result, the mean value and the standard deviation of the clothing morphological feature latent variable are multiplied together to obtain the feature map.

[0021] Optionally, inputting the set action posture image into a calibration feature network to obtain clothing morphological feature latent variables includes:

[0022] Extracting clothing and posture features of the motion capture person in the set action image;

[0023] The clothing posture features are input into the calibration feature network to obtain clothing morphological feature latent variables.

[0024] According to a second aspect of an embodiment of the present application, a motion capture system is provided, the system comprising:

[0025] An image acquisition module is used to acquire the set action images and the images to be captured of the motion capture personnel;

[0026] A feature calibration module, configured to input the set action image into a calibration feature network to obtain clothing morphological feature latent variables;

[0027] A motion capture module training module, used for inputting the clothing morphological feature latent variables into a motion capture network to train the motion capture network;

[0028] The motion capture module is used to input the image to be motion captured into the motion capture network to obtain a posture estimation result.

[0029] Optionally, the system further comprises:

[0030] The image acquisition module is further used to acquire random action images of the motion capture personnel;

[0031] The fine-tuning module is used to annotate the correspondence between images and joint rotations based on random motion images of motion capture personnel; and is also used to input the annotated results into the motion capture network to train the motion capture network.

[0032] Optionally, the motion capture module training module is used to:

[0033] Performing a linear transformation on the clothing morphological feature latent variables to obtain a feature map; the feature map includes the clothing morphological features of the motion capture person and prior information of the motion capture environment;

[0034] The feature map is input into the motion capture network for training.

[0035] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.

[0036] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to implement the method described in the first aspect above.

[0037] In summary, the present invention provides a motion capture method and system. These methods collect a set motion image and an image to be captured of a person performing the motion capture; input the set motion image into a calibration feature network to obtain clothing morphological latent variables; input the clothing morphological latent variables into a motion capture network to train the network; and input the image to be captured into the motion capture network to obtain pose estimation results. This method allows for efficient motion capture and improves the accuracy of motion capture. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other implementation drawings based on the provided drawings without inventive effort.

[0039] The structures, proportions, sizes, etc. illustrated in this specification are intended solely to complement the contents disclosed herein and to facilitate understanding and reading by persons skilled in the art. They are not intended to limit the conditions under which the present invention may be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportions, or adjustments in sizes, without affecting the efficacy and objectives of the present invention, shall remain within the scope of the technical contents disclosed herein.

[0040] Figure 1 A flow chart of a motion capture method provided in an embodiment of the present application;

[0041] Figure 2 A schematic diagram of a motion capture embodiment provided in an embodiment of the present application;

[0042] Figure 3 A block diagram of a motion capture system provided in an embodiment of the present application;

[0043] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown;

[0044] Figure 5 A schematic diagram of a computer-readable storage medium provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0045] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0046] The purpose of the embodiments of the present application is to reduce the high cost of traditional motion capture solutions and solve the problem of motion capture personnel's dependence on clothing, and to provide an image motion capture method based on neural network multi-motion calibration, so that motion capture personnel do not need to carry out complex venue matching for optical motion capture, do not need to wear specific motion capture clothing or equipment, and can achieve high-precision motion capture with just daily clothing and a camera. While ensuring the accuracy of motion capture, the efficiency of the entire capture process is significantly improved.

[0047] Figure 1 A motion capture method provided in an embodiment of the present application is shown, the method comprising:

[0048] Step 101: collecting the set action image and the image to be captured of the motion capture person;

[0049] Step 102: inputting the set action image into the calibration feature network to obtain clothing morphological feature latent variables;

[0050] Step 103: inputting the clothing morphological feature latent variables into a motion capture network to train the motion capture network;

[0051] Step 104: Input the image to be motion captured into the motion capture network to obtain a pose estimation result.

[0052] In a possible embodiment, when collecting the set action images and the images to be action captured of the motion capture personnel in step 101, the method further includes: collecting random action images of the motion capture personnel; annotating the correspondence between the images and joint rotations based on the random action images of the motion capture personnel; and inputting the annotation results into the motion capture network to train the motion capture network.

[0053] In a possible implementation, in step 102, inputting the set action posture image into a calibration feature network to obtain clothing morphological feature latent variables includes:

[0054] Extracting clothing posture features of the motion capture personnel in the set action image; inputting the clothing posture features into the calibration feature network to obtain clothing morphological feature latent variables.

[0055] In a possible implementation, in step 103, inputting the clothing morphological feature latent variables into a motion capture network to train the motion capture network includes:

[0056] The clothing morphological feature latent variables are linearly transformed to obtain a feature map; the feature map includes the clothing morphological features of the motion capture personnel and prior information of the motion capture environment; the feature map is input into the motion capture network for training.

[0057] In a possible implementation, performing a linear transformation on the clothing morphological feature latent variables to obtain a feature map includes: normalizing the clothing morphological feature latent variables; and multiplying the normalized result, the mean value, and the standard deviation of the clothing morphological feature latent variables to obtain the feature map.

[0058] As can be seen, the motion capture method proposed in the present embodiment first extracts latent variables of the current subject's clothing morphological characteristics by having the subject perform standard motion poses. This latent variable is then used as the feature input for the fixed layer of the motion capture network, which is then used to change the feature statistics to automatically calibrate the motion capture network to adapt to the current subject and scene characteristics. After obtaining the current subject's characteristics, the network achieves a more accurate pose estimation for the subject, ultimately driving and controlling the body movements of the virtual character. This enables motion capture through RGB image acquisition, improves the accuracy and robustness of the entire motion capture system, and enhances the fluidity and scene adaptability of virtual character motion capture.

[0059] The following is combined with Figure 2 The solutions provided in the embodiments of the present application are further explained.

[0060] During the image acquisition phase, the acquired images are divided into three categories:

[0061] 1. Image Capture for Feature Calibration: The motion capture operator imitates specified motion standards. The camera captures the imitated motions to extract the operator's body features for the second step of feature calibration. These standard motions include ten types: front T-pose, back T-pose, left and right side T-pose, front-hand cross, and left and right leg high-leg raises. These standard calibration motions are derived from unsupervised computation using influence analysis of the training data.

[0062] 2. Collect images for fine-tuning the motion capture module: Motion capture network predictions are prone to offset, so the complex random movements of the current motion capture personnel are collected and manually processed by k frames for fine-tuning the motion capture module.

[0063] 3. Collect images for motion capture: Collect the posture and action images of the person in front of the camera for real-time 3D posture estimation as the input of the motion capture module.

[0064] Among them, the first and third images need to be collected in the same motion capture scene. The first image is used as the input of the calibration feature network, and the output feature vector is obtained as the input of the feature layer of the calibration motion capture module. The particularity of the motion capture personnel and the particularity of the environment are used as prior information in the motion capture module, so that the motion capture module does not need to input continuous frames in the subsequent formal capture, and the output posture obtained by inputting a single frame of human body image into the motion capture module will not jitter, which improves the computational efficiency of the formal motion capture, thereby realizing real-time and high-precision posture estimation.

[0065] It should be noted that the motion capture scene in the embodiment of the present application does not have any special requirements on the clothing of the motion capture personnel, does not require the wearing of traditional motion clothing or motion sensors, and does not have any special requirements on the lighting or object arrangement of the motion capture scene.

[0066] Furthermore, during the feature calibration phase, a convolutional neural network dedicated to feature calibration is used to forward-propagate a latent variable describing the current motion capture subject's clothing and posture. This latent variable, which contains the current subject's clothing and body features, serves as the feature input for the fixed layer of the pose estimation network, used to calibrate and adapt to the current subject and the motion capture environment. This effectively reduces the motion capture module's requirements for generalizing the subject and scene, allowing it to focus more on generalizing the 3D estimation direction of the motion pose. By decoupling the motion pose from the clothing and environment, the overall motion capture capability and accuracy are improved.

[0067] The calibration feature network is run only once before formal pose estimation (the motion capture network). The resulting latent variables correspond to the current pose capture subject's attire and the current motion capture environment. These latent variables are then used as the feature layer input for the pose estimation network during subsequent formal pose estimation, eliminating the need to recapture new images for recalibration. Because calibration is only run once, the computational resources used by the calibration feature network are negligible in the overall computational resources, improving overall operational efficiency and enabling real-time motion capture.

[0068] The calibration feature module is jointly trained with the motion capture module, eliminating the need for explicit labeling of the current motion capture subject's appearance and posture, or the motion capture environment, during the data labeling process. Data labeling refers to the large amount of "image-joint rotation correspondences" used during overall training, while subsequent manual labeling refers to the minimal amount of "image-joint rotation correspondences" used for fine-tuning training.

[0069] This training method makes full use of the feature learning ability of neural networks, and indirectly solves the dimensionality disaster problem of environment description and the complexity of data labeling through system design, reducing the labeling cost of motion capture.

[0070] Furthermore, in the fine-tuning stage of the motion capture network, the corresponding posture labels of the complex postures of the motion capture personnel are manually marked to solve the problem that the motion capture network prediction of some actions is prone to offset, and to perform fine-tuning training of the motion capture module.

[0071] Some movements are so complex that they require annotation and subsequent fine-tuning to ensure the network can handle them. Motion capture network predictions are prone to bias, so we collect the complex, random movements of the person being captured and manually perform k-frame processing. This can be understood as manually processing a small amount of complex data before training the network. K-frame processing refers to manually annotating the "image-joint rotation correspondences." This step significantly improves the performance of complex motion capture, effectively achieving film-quality image motion capture with minimal human intervention.

[0072] Furthermore, during the motion capture phase, the corresponding gestures are captured based on the captured images. The motion capture module is influenced by the latent features extracted by the motion capture subject calibration module and automatically calibrates to the current subject and capture environment, achieving a higher-precision, dedicated motion capture network.

[0073] The calibration feature results are used as additional inputs to the feature layer of the motion capture module, altering the feature layer's statistics. The image used for motion capture is then used as input to the motion capture module, which then outputs the motion capture results. This step, by combining the calibration features, decouples the motion pose from the motion capture clothing and environment. This allows the motion capture module to focus solely on the motion capture problem, automatically calibrating to each individual motion capture subject, their clothing, and the different motion capture environments, thereby improving the accuracy and robustness of the entire motion capture system.

[0074] The motion capture module takes as input the captured image and calibration features, and outputs quaternions for 3D pose estimation. The calibration features serve as constant inputs to the motion capture module, unchanged by the input captured image. These are implemented by the Adain module, a component of the motion capture network that converts the outputs of the calibration feature network into inputs for the motion capture network.

[0075] The adain module inputs the output feature map of the previous convolutional network layer, followed by a linear transformation of the current adain module to calibrate the feature mean and variance, and outputs a new feature map. The adain module first normalizes the input feature map to a new feature map with a mean of 0 and a standard deviation of 1. This feature map is then multiplied by the mean and standard deviation obtained from the calibration feature map. The resulting feature map, which incorporates the current motion capture subject's posture and appearance characteristics as well as prior information about the current motion capture environment, serves as the input for the next convolutional layer.

[0076] By calibrating the feature motion capture module, the pressure to decouple posture and appearance features from motion capture environment characteristics is reduced, allowing for greater focus on the core task of motion estimation. This achieves the splitting and aggregation of the motion capture system's generalization capabilities, improving the system's overall motion capture accuracy, computational efficiency, and generalization capabilities. This system can be applied to a variety of motion capture scenarios, including virtual anchors, VR games, and film and television production.

[0077] Once the motion capture environment or the motion capture personnel are changed, motion capture needs to be repeated from the first step.

[0078] In summary, the present invention provides a motion capture method that collects a set motion image and an image to be captured of a person performing the motion capture; inputs the set motion image into a calibration feature network to obtain clothing morphological feature latent variables; inputs the clothing morphological feature latent variables into a motion capture network to train the motion capture network; and inputs the image to be captured into the motion capture network to obtain a pose estimation result. This method allows for efficient motion capture and improves the accuracy of motion capture.

[0079] Based on the same technical concept, the embodiment of the present application also provides a motion capture system, such as Figure 3 As shown, the system includes:

[0080] The image acquisition module 301 is used to acquire the set action image and the image to be captured by the motion capture personnel;

[0081] A feature calibration module 302 is configured to input the set action image into a calibration feature network to obtain clothing morphological feature latent variables;

[0082] A motion capture module training module 303 is used to input the clothing morphological feature latent variables into a motion capture network to train the motion capture network;

[0083] The motion capture module 304 is configured to input the image to be motion captured into the motion capture network to obtain a pose estimation result.

[0084] In a possible implementation, the system further includes: the image acquisition module 301 is further configured to acquire random motion images of the motion capture personnel;

[0085] The fine-tuning module is used to annotate the correspondence between images and joint rotations based on random motion images of motion capture personnel; and is also used to input the annotated results into the motion capture network to train the motion capture network.

[0086] In a possible implementation, the motion capture module training module 303 is configured to:

[0087] The clothing morphological feature latent variables are linearly transformed to obtain a feature map; the feature map includes the clothing morphological features of the motion capture personnel and prior information of the motion capture environment; the feature map is input into the motion capture network for training.

[0088] The present application also provides an electronic device corresponding to the method provided in the above embodiment. Figure 4, which shows a schematic diagram of an electronic device provided in some embodiments of the present application. The electronic device 20 may include: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected via the bus 202. The memory 201 stores a computer program executable on the processor 200. When the processor 200 executes the computer program, it executes the method provided in any of the aforementioned embodiments of the present application.

[0089] The memory 201 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. The system network element and at least one other network element are connected via at least one physical port 203 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.

[0090] The bus 202 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 201 is used to store programs. The processor 200 executes the programs upon receiving execution instructions. The methods disclosed in any of the aforementioned embodiments of the present application may be applied to or implemented by the processor 200.

[0091] The processor 200 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 200 or by software instructions. The above processor 200 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 201 , and the processor 200 reads the information in the memory 201 and completes the steps of the above method in combination with its hardware.

[0092] The electronic device provided in the embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented by them.

[0093] The present application also provides a computer-readable storage medium corresponding to the method provided in the above embodiment. Figure 5 The computer-readable storage medium shown is a CD 30 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the method provided by any of the aforementioned embodiments is executed.

[0094] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.

[0095] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0096] It should be noted that:

[0097] The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other device. Various general-purpose devices may also be used in conjunction with the teachings herein. Based on the above description, it is apparent that the structure required for constructing such devices is suitable. In addition, the present application is not directed to any specific programming language. It should be understood that various programming languages ​​may be utilized to implement the present application described herein, and the above description of specific languages ​​is provided for the purpose of disclosing the best mode of implementation of the present application.

[0098] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0099] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the claims below, inventive aspects lie in fewer than all the features of the individual embodiments disclosed above. Accordingly, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim itself serving as a separate embodiment of the present application.

[0100] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition may be divided into multiple submodules or subunits or subcomponents. All features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed herein may be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0101] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims below, any of the claimed embodiments may be used in any combination.

[0102] The various component embodiments of the present application can be implemented in hardware, or implemented in a software module running on one or more processors, or implemented in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the creation device of the virtual machine according to an embodiment of the present application. The application can also be implemented as a part or all of the equipment or device program (for example, computer program and computer program product) for performing the method described herein. Such a program realizing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0103] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbols placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0104] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A motion capture method, characterized in that: The method comprises: Collect the set action images and the images to be captured of the motion capture personnel; Inputting the set action image into the calibration feature network to obtain clothing morphological feature latent variables; the clothing morphological feature latent variables include clothing and body features of the motion capture person; Inputting the clothing morphological feature latent variables into a motion capture network to train the motion capture network; Inputting the image to be motion captured into the motion capture network to obtain a posture estimation result; The method of inputting the clothing morphological feature latent variables into the motion capture network to train the motion capture network includes: calculating the mean and standard deviation of the clothing morphological feature latent variables; inputting the output feature map of the upper convolutional network of the motion capture network into the adain module, calibrating the mean and variance of the features through linear transformation, and outputting a new feature map; the new feature map includes the clothing morphological features of the motion capture personnel and the prior information of the motion capture environment; and inputting the new feature map into the motion capture network for training.

2. The method according to claim 1, wherein When collecting the set action image and the image to be captured of the motion capture person, the method further includes: Collect random motion images of motion capture personnel; Annotate the correspondence between images and joint rotations based on random motion images of motion capture personnel; The annotation results are input into the motion capture network to train the motion capture network.

3. The method according to claim 1, wherein The step of inputting the set action image into the calibration feature network to obtain clothing morphological feature latent variables includes: Extracting clothing and posture features of the motion capture person in the set action image; The clothing posture features are input into the calibration feature network to obtain clothing morphological feature latent variables.

4. A motion capture system, characterized in that: The system comprises: An image acquisition module is used to acquire the set action images and the images to be captured of the motion capture personnel; A feature calibration module is used to input the set action image into the calibration feature network to obtain clothing morphological feature latent variables; the clothing morphological feature latent variables include clothing and body features of the motion capture personnel; A motion capture module training module is used to input the clothing morphological feature latent variables into the motion capture network to train the motion capture network; the inputting the clothing morphological feature latent variables into the motion capture network to train the motion capture network includes: calculating the mean and standard deviation of the clothing morphological feature latent variables; inputting the output feature map of the upper convolutional network of the motion capture network into the adain module, calibrating the mean and variance of the features through linear transformation, and outputting a new feature map; the new feature map includes the clothing morphological features of the motion capture personnel and prior information of the motion capture environment; and inputting the new feature map into the motion capture network for training. The motion capture module is used to input the image to be motion captured into the motion capture network to obtain a posture estimation result.

5. The system according to claim 4, wherein: The system further comprises: The image acquisition module is further used to acquire random action images of the motion capture personnel; The fine-tuning module is used to annotate the correspondence between images and joint rotations based on random motion images of motion capture personnel; and is also used to input the annotated results into the motion capture network to train the motion capture network.

6. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method according to any one of claims 1 to 3.

7. A computer-readable storage medium, characterized in that Computer-readable instructions are stored thereon, and the computer-readable instructions can be executed by a processor to implement the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Action guidance method and device based on action capture

    CN110045823A

  • Motion capture method and device, electronic equipment and computer readable storage medium

    CN113079136A