Visual motion capture method and device, equipment and storage medium

By performing topological correction on the motion capture parameters in the visual motion capture method, the problem of poor motion coherence caused by directly assembling motion capture results from different human body parts is solved, thereby improving the consistency and visual effect of human motion capture.

CN121789265APending Publication Date: 2026-04-03BEIJING WATER DROP INTERACTIVE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-06-19
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, vision-based motion capture methods suffer from poor motion coherence due to the different recognition model requirements for different human body parts and inconsistent stability of output motion capture parameters, which fails to meet users' requirements for the coherence of human motion capture.

Method used

By acquiring scene capture images, motion capture images of each human body part are determined and input into the corresponding recognition model. The motion capture parameters are corrected using the topological connection relationship of the human body parts to generate motion capture correction parameters, and finally the motion capture result is determined.

Benefits of technology

It improves the consistency and visual effect of human motion capture, solves the problem of poor motion continuity, and achieves more natural and smooth human motion stitching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789265A_ABST
    Figure CN121789265A_ABST
Patent Text Reader

Abstract

The invention discloses a visual motion capture method and device, equipment and a storage medium. The method comprises the following steps: acquiring a scene capture image, and determining motion capture images of at least two human body parts according to the scene capture image; respectively inputting each motion capture image into a recognition model matched with each human body part, and determining motion capture parameters of each human body part; according to the topological connection relation of the human body parts, correcting the motion capture parameters of the human body parts to obtain motion capture correction parameters of the human body parts; and determining a motion capture result according to the motion capture correction parameter of each human body part. According to the technical scheme, topology correction is carried out on the motion capture parameters of all the human body parts, the problem of poor motion linkage caused by direct splicing of the motion capture results of all the human body parts is solved, and the consistency and visual effect of human body motion capture are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual processing technology, and in particular to a visual motion capture method, apparatus, device, and storage medium. Background Technology

[0002] Currently, vision-based motion capture typically involves using separate recognition models for each body part to identify its movements and determine motion capture parameters. Based on these parameters, motion capture results for each body part are generated. These results are then concatenated and assembled to obtain the final motion capture result for the entire human body.

[0003] However, different body part recognition models have different requirements for input data, and the stability of the output motion capture parameters is also different. The prediction results of different body part models do not take into account the overall coordination of human body movements. Directly piecing together the motion capture results of each body part makes it difficult to achieve smooth connection and cannot meet the user's need for the continuity of human body motion capture. Summary of the Invention

[0004] This invention provides a visual motion capture method, device, equipment, and storage medium to solve the problem of poor motion coherence caused by directly assembling motion capture results from different human body parts. By performing topological correction on the motion capture parameters of each human body part, the consistency and visual effect of human motion capture can be improved.

[0005] According to one aspect of the present invention, a visual motion capture method is provided, the method comprising:

[0006] Acquire scene capture images, and based on the scene capture images, determine motion capture images of at least two human body parts;

[0007] Each motion capture image is input into the recognition model that matches each human body part to determine the motion capture parameters for each human body part.

[0008] Based on the topological connection relationship of each human body part, the motion capture parameters of each human body part are corrected to obtain the motion capture correction parameters of each human body part.

[0009] The motion capture results are determined by adjusting the motion capture parameters for each part of the human body.

[0010] According to another aspect of the present invention, a visual motion capture device is provided, the device comprising:

[0011] The motion capture image acquisition module is used to acquire scene capture images and determine motion capture images of at least two human body parts based on the scene capture images.

[0012] The motion capture parameter determination module is used to input each motion capture image into the recognition model matching each human body part to determine the motion capture parameters of each human body part.

[0013] The correction parameter determination module is used to correct the motion capture parameters of each human body part according to the topological connection relationship of each human body part, so as to obtain the motion capture correction parameters of each human body part.

[0014] The motion capture result determination module is used to determine the motion capture result based on the motion capture correction parameters for each part of the human body.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a visual motion capture method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement a visual motion capture method according to any embodiment of the present invention.

[0020] The technical solution of this invention involves acquiring scene capture images and determining motion capture images of at least two human body parts based on these images. Each motion capture image is then input into a matching recognition model for each human body part to determine motion capture parameters. These parameters are then corrected based on the topological connections of the human body parts to obtain corrected motion capture parameters. Finally, the motion capture result is determined based on these corrected parameters. This technical solution addresses the poor motion coherence caused by directly assembling motion capture results from different human body parts. By performing topological correction on the motion capture parameters of each human body part, the consistency and visual effectiveness of human motion capture are improved.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a visual motion capture method provided in Embodiment 1 of the present invention;

[0024] Figure 2A This is a flowchart of a visual motion capture method provided in Embodiment 2 of the present invention;

[0025] Figure 2B This is a schematic diagram of human joint points provided in Embodiment 2 of the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of a visual motion capture device according to Embodiment 3 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device that implements a visual motion capture method according to an embodiment of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations.

[0030] Example 1

[0031] Figure 1 This is a flowchart illustrating a visual motion capture method according to Embodiment 1 of the present invention. This embodiment is applicable to visual motion capture scenarios in fields such as 3D animation production and game development. The method can be executed by a visual motion capture device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0032] S110. Acquire scene capture images and determine motion capture images of at least two human body parts based on the scene capture images.

[0033] This solution can be executed by electronic devices such as computers and smartphones. These devices can be equipped with one or more image acquisition devices, such as cameras. The image acquisition devices can be used to capture images of the scene. Simply put, a single image acquisition device can achieve single-view image acquisition, while multiple image acquisition devices can achieve multi-view image acquisition.

[0034] Electronic devices can perform human detection on scene-captured images, identify human body parts in the scene-captured images, and locate the positions of these human body parts within the scene-captured images. These human body parts can include the torso, upper limbs, lower limbs, hands, feet, head, and face. Based on the human detection results of the scene-captured images, the electronic device can determine the motion capture images of each human body part within the scene-captured images.

[0035] It should be noted that the number of scene capture images can be one or multiple. Multiple scene capture images can be arranged in chronological order or can be video frames extracted from scene capture video.

[0036] S120. Input each motion capture image into the recognition model that matches each human body part to determine the motion capture parameters of each human body part.

[0037] Electronic devices can be equipped with recognition models for various human body parts. These models can be used to determine motion capture parameters for human body parts based on motion capture images. The recognition models can be built using deep learning algorithms, and the motion capture parameters can include parameters such as the angle, velocity, and acceleration of the human body part's movement.

[0038] After obtaining motion capture images of various body parts, the electronic device can input each motion capture image into a corresponding recognition model for that body part. The recognition model can extract motion features from each motion capture image of that body part, as well as motion features between multiple motion capture images of that body part. Based on these motion features and / or movement characteristics, the electronic device can determine the motion capture parameters for that body part.

[0039] S130. Based on the topological connection relationship of each human body part, the motion capture parameters of each human body part are corrected to obtain the motion capture correction parameters of each human body part.

[0040] Electronic devices can determine the topological connections between different parts of the human body. Based on the motion capture parameters of these topologically connected parts, the devices can correct the motion capture parameters for each part, thus obtaining corrected motion capture parameters for each part. For example, if the upper limb and hand are topologically connected, the electronic devices can correct the motion capture parameters based on the motion capture parameters of the upper limb and hand, thus obtaining corrected motion capture parameters for both the upper limb and hand.

[0041] S140. Determine the motion capture result based on the motion capture correction parameters for each body part.

[0042] After correcting the motion capture parameters for each body part, the electronic device can generate motion capture results for each body part based on these corrected parameters. By stitching together the motion capture results for each body part, the final motion capture result for the entire human body can be obtained.

[0043] This technical solution acquires scene capture images and determines motion capture images of at least two human body parts based on these images. Each motion capture image is then input into a matching recognition model for that specific human body part to determine its motion capture parameters. These parameters are then corrected based on the topological connections between the human body parts, resulting in corrected motion capture parameters. Finally, the motion capture result is determined based on these corrected parameters. This solution addresses the problem of poor motion coherence caused by directly assembling motion capture results from different human body parts. By performing topological correction on the motion capture parameters for each human body part, the consistency and visual quality of human motion capture are improved.

[0044] Example 2

[0045] Figure 2 is a flowchart of a visual motion capture method provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiment. As shown in Figure 2, the method includes:

[0046] S210, Acquire scene capture image.

[0047] S220. Based on a pre-trained human detection model, perform human body part detection on scene capture images and determine the calibration positions of at least two human body parts.

[0048] Understandably, due to the differences in features among different human body parts, electronic devices can pre-train human detection models for each body part. Based on these matched models, the scene-captured image can be used to detect each body part. The structures of the human detection models for each body part can be the same or different; they can be 3D landmark-based models. By inputting the scene-captured image into the human detection models for each body part, the calibrated position of each body part, i.e., its 3D landmark, can be either the absolute or relative position of the body part within the scene-captured image.

[0049] This solution uses different human body detection models to detect different parts of the human body, enabling accurate detection of each part.

[0050] S230. Based on the calibration position of each human body part, determine the bounding box of each human body part, and based on the bounding box of each human body part, determine the motion capture image of each human body part.

[0051] Based on the calibrated positions of each human body part, the electronic device can calculate the bounding box of each part. The shape of the bounding box can be regular, such as a rotated rectangular bounding box, or irregular; this embodiment does not limit the shape of the bounding box. After obtaining the bounding boxes of each human body part, the electronic device can extract the scene capture image of the area covered by the bounding box of each human body part as the motion capture image of each human body part.

[0052] S240. Input each motion capture image into the recognition model that matches each human body part to determine the motion capture parameters of each human body part.

[0053] In one feasible approach, the human body part includes non-head parts; the motion capture parameters include a rotation matrix that matches human joints associated with the human body part.

[0054] The step of inputting each motion capture image into a recognition model that matches each human body part to determine the motion capture parameters for each human body part includes:

[0055] If the motion capture image is a non-head area motion capture image, then the non-head area motion capture image is input into the matching human body part recognition model to determine the rotation matrix of the human body joints associated with the non-head area.

[0056] Motion capture parameters for non-head parts of the human body, such as the torso, limbs, hands, and feet, can be represented by a rotation matrix matching human joints. Figure 2B This is a schematic diagram of human joints provided in Embodiment 2 of the present invention. Based on the SMPL (Skinned Multi-Person Linear) model, it can be achieved through, for example... Figure 2B The motion capture parameters of the 24 joints shown are used to describe human movement. The matched human body part recognition model can predict the rotation matrix of the human joints associated with each human body part based on the motion capture image. Each human joint can be matched with one rotation matrix.

[0057] In another feasible embodiment, the human body part also includes the head part; the motion capture parameters also include a rotation matrix matching reference points associated with the head part and facial expression shape parameters;

[0058] The step of inputting each motion capture image into a recognition model that matches each human body part to determine the motion capture parameters for each human body part includes:

[0059] If the motion capture image is a motion capture image of the head, then the motion capture image of the head is input into the head recognition model to determine the rotation matrix of the reference point matching associated with the head and the facial expression shape parameters.

[0060] Compared to non-head parts such as the torso and limbs, the head has unique characteristics, exhibiting both head movements and facial expressions. Therefore, head motion capture parameters can include a rotation matrix matched with reference points associated with the head and facial expression shape parameters. The reference points associated with the head can be used to describe the locations of head movement changes; for example, the center of the head or the position of the eyeballs can be used as reference points.

[0061] Electronic devices can input motion-captured images of the head into a head recognition model to obtain a rotation matrix for matching reference points associated with the head and facial expression shape parameters. Specifically, the head recognition model can include a first recognition branch and a second recognition branch. The first recognition branch is used to determine the rotation matrix for matching reference points associated with the head, and the second recognition branch is used to determine the facial expression shape parameters. The second recognition branch can be constructed based on a deep model.

[0062] S250. Based on the topological connection relationship of each human body part, the motion capture parameters of each human body part are corrected to obtain the motion capture correction parameters of each human body part.

[0063] In this solution, optionally, the step of correcting the motion capture parameters of each human body part based on the topological connection relationship of each human body part to obtain the motion capture correction parameters of each human body part includes:

[0064] The rotation matrices associated with each human body part are converted into quaternions, and the quaternions matched with each human body part are aligned in coordinate system.

[0065] Based on the topological connection relationship of each human body part, determine the quaternion to be decomposed from the quaternions matched by each human body part, and determine the decomposition direction of the quaternion to be decomposed.

[0066] The quaternion to be decomposed is decomposed according to the decomposition direction, and the axial rotation terms that match the quaternion to be decomposed are determined.

[0067] Based on the axial rotation terms matched by the quaternions to be decomposed, determine the motion capture correction parameters for each human body part.

[0068] Since motion capture parameters for different body parts are predicted by different recognition models, directly driving the 3D motion generator can easily lead to inconsistencies in the connection between body parts. Therefore, the electronic device can convert the rotation matrices associated with each body part into quaternions and align the coordinate systems of the quaternions matched by each body part. Based on the topological connections of each body part, the electronic device can determine the quaternion to be decomposed from the quaternions matched by each body part and determine the decomposition direction of the quaternion to be decomposed. By decomposing the quaternion to be decomposed according to the decomposition direction, the electronic device can determine the axial rotation terms that match the quaternion to be decomposed. Based on the axial rotation terms matched by the quaternion to be decomposed, the electronic device can determine the motion capture correction parameters for each body part.

[0069] In a specific example, the upper limb and hand have a topological connection. The electronic device can decompose the quaternion of the hand along the elbow direction and transfer the decomposed axial rotation terms to the upper limb quaternion according to a pre-set constraint ratio, obtaining the upper limb corrected quaternion, and thus the motion capture correction parameters for the upper limb. The hand quaternion is then corrected based on the remaining terms except for the axial rotation terms, resulting in the corrected hand quaternion, and thus the hand motion capture correction parameters. Through the fusion of motion capture parameters, the generated motion capture result after stitching is more natural and smooth.

[0070] S260. Determine the motion capture result based on the motion capture correction parameters for each body part.

[0071] In a preferred embodiment, after determining the calibration locations of at least two human body parts, the method further includes:

[0072] Based on the calibrated positions of each body part, the motion capture images are corrected and updated.

[0073] Before predicting motion capture parameters from motion capture images of various body parts, electronic devices can correct each motion capture image based on the calibrated position of each body part to obtain high-quality motion capture images for the recognition model, thereby achieving reliable prediction of motion capture parameters.

[0074] This solution improves the stability and accuracy of motion capture model predictions by correcting human motion capture images.

[0075] Based on the above scheme, optionally, the step of correcting and updating the motion capture images of each human body part according to the calibrated position of each human body part includes:

[0076] Based on the relative positional relationship of the calibrated positions of each human body part, determine the perspective transformation matrix matching each human body part;

[0077] Based on the perspective transformation matrix matched to each human body part, the motion capture images are corrected and updated.

[0078] Specifically, electronic devices can calculate the perspective transformation matrix matching each human body part based on the relative positional relationship of the calibrated positions, and then use the perspective transformation matrix to correct the motion capture image. For example, for motion capture images of the hand, electronic devices can use the calibrated positions of human joints such as the base of the index finger, the base of the ring finger, and the base of the palm to calculate the perspective transformation matrix, and then use the perspective transformation matrix to correct the motion capture image of the hand.

[0079] This technical solution acquires scene capture images and determines motion capture images of at least two human body parts based on these images. Each motion capture image is then input into a matching recognition model for that specific human body part to determine its motion capture parameters. These parameters are then corrected based on the topological connections between the human body parts, resulting in corrected motion capture parameters. Finally, the motion capture result is determined based on these corrected parameters. This solution addresses the problem of poor motion coherence caused by directly assembling motion capture results from different human body parts. By performing topological correction on the motion capture parameters for each human body part, the consistency and visual quality of human motion capture are improved.

[0080] Example 3

[0081] Figure 3 This is a schematic diagram of the structure of a visual motion capture device provided in Embodiment 3 of the present invention.

[0082] like Figure 3 As shown, the device includes:

[0083] The motion capture image acquisition module 310 is used to acquire scene capture images and determine motion capture images of at least two human body parts based on the scene capture images.

[0084] The motion capture parameter determination module 320 is used to input each motion capture image into the recognition model matching each human body part to determine the motion capture parameters of each human body part.

[0085] The correction parameter determination module 330 is used to correct the motion capture parameters of each human body part according to the topological connection relationship of each human body part, so as to obtain the motion capture correction parameters of each human body part.

[0086] The motion capture result determination module 340 is used to determine the motion capture result based on the motion capture correction parameters of each human body part.

[0087] In this solution, optionally, the motion capture image acquisition module 310 is specifically used for:

[0088] Based on a pre-trained human detection model, human body parts are detected in scene-captured images to determine the calibration locations of at least two human body parts.

[0089] Based on the calibrated positions of each human body part, the bounding box of each human body part is determined, and based on the bounding box of each human body part, the motion capture image of each human body part is determined.

[0090] In one feasible approach, the human body part includes non-head parts; the motion capture parameters include a rotation matrix that matches human joints associated with the human body part.

[0091] The motion capture parameter determination module 320 is specifically used for:

[0092] If the motion capture image is a non-head area motion capture image, then the non-head area motion capture image is input into the matching human body part recognition model to determine the rotation matrix of the human body joints associated with the non-head area.

[0093] In another feasible embodiment, the human body part also includes the head part; the motion capture parameters also include a rotation matrix matching reference points associated with the head part and facial expression shape parameters;

[0094] The motion capture parameter determination module 320 is specifically used for:

[0095] If the motion capture image is a motion capture image of the head, then the motion capture image of the head is input into the head recognition model to determine the rotation matrix of the reference point matching associated with the head and the facial expression shape parameters.

[0096] Based on the above scheme, optionally, the correction parameter determination module 330 is specifically used for:

[0097] The rotation matrices associated with each human body part are converted into quaternions, and the quaternions matched with each human body part are aligned in coordinate system.

[0098] Based on the topological connection relationship of each human body part, determine the quaternion to be decomposed from the quaternions matched by each human body part, and determine the decomposition direction of the quaternion to be decomposed.

[0099] The quaternion to be decomposed is decomposed according to the decomposition direction, and the axial rotation terms that match the quaternion to be decomposed are determined.

[0100] Based on the axial rotation terms matched by the quaternions to be decomposed, determine the motion capture correction parameters for each human body part.

[0101] In a preferred embodiment, the motion capture parameter determination module 320 is further configured to:

[0102] Based on the calibrated positions of each body part, the motion capture images are corrected and updated.

[0103] Based on the above scheme, the motion capture parameter determination module 320 is specifically used for:

[0104] Based on the relative positional relationship of the calibrated positions of each human body part, determine the perspective transformation matrix matching each human body part;

[0105] Based on the perspective transformation matrix matched to each human body part, the motion capture images are corrected and updated.

[0106] The visual motion capture device provided in this embodiment of the invention can execute a visual motion capture method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0107] Example 4

[0108] Figure 4A schematic diagram of an electronic device 410 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0109] like Figure 4 As shown, the electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory (ROM) 412 or a random access memory (RAM) 413, communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the ROM 412 or loaded from storage unit 418 into the RAM 413. The RAM 413 may also store various programs and data required for the operation of the electronic device 410. The processor 411, ROM 412, and RAM 413 are interconnected via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.

[0110] Multiple components in electronic device 410 are connected to I / O interface 415, including: input unit 416, such as keyboard, mouse, etc.; output unit 417, such as various types of displays, speakers, etc.; storage unit 418, such as disk, optical disk, etc.; and communication unit 419, such as network card, modem, wireless transceiver, etc. Communication unit 419 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0111] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as a visual motion capture method.

[0112] In some embodiments, a visual motion capture method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 410 via ROM 412 and / or communication unit 419. When the computer program is loaded into RAM 413 and executed by processor 411, one or more steps of the visual motion capture method described above may be performed. Alternatively, in other embodiments, processor 411 may be configured to perform a visual motion capture method by any other suitable means (e.g., by means of firmware).

[0113] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0114] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0115] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0117] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0118] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0119] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0120] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A visual motion capture method, characterized in that, The method includes: Acquire scene capture images, and based on the scene capture images, determine motion capture images of at least two human body parts; Each motion capture image is input into the recognition model that matches each human body part to determine the motion capture parameters for each human body part. Based on the topological connection relationship of each human body part, the motion capture parameters of each human body part are corrected to obtain the motion capture correction parameters of each human body part. The motion capture results are determined by adjusting the motion capture parameters for each part of the human body.

2. The method according to claim 1, characterized in that, The step of determining motion capture images of at least two parts of the human body based on scene-captured images includes: Based on a pre-trained human detection model, human body parts are detected in scene-captured images to determine the calibration locations of at least two human body parts. Based on the calibrated positions of each human body part, the bounding box of each human body part is determined, and based on the bounding box of each human body part, the motion capture image of each human body part is determined.

3. The method according to claim 1, characterized in that, The human body parts include non-head parts; the motion capture parameters include a rotation matrix that matches the human joints associated with the human body parts. The step of inputting each motion capture image into a recognition model that matches each human body part to determine the motion capture parameters for each human body part includes: If the motion capture image is a non-head area motion capture image, then the non-head area motion capture image is input into the matching human body part recognition model to determine the rotation matrix of the human body joints associated with the non-head area.

4. The method according to claim 3, characterized in that, The human body parts also include the head; the motion capture parameters also include a rotation matrix that matches reference points associated with the head and facial expression shape parameters. The step of inputting each motion capture image into a recognition model that matches each human body part to determine the motion capture parameters for each human body part includes: If the motion capture image is a motion capture image of the head, then the motion capture image of the head is input into the head recognition model to determine the rotation matrix of the reference point matching associated with the head and the facial expression shape parameters.

5. The method according to claim 4, characterized in that, The process of correcting the motion capture parameters of each human body part based on the topological connection relationship of each part, to obtain the corrected motion capture parameters for each human body part, includes: The rotation matrices associated with each human body part are converted into quaternions, and the quaternions matched with each human body part are aligned in coordinate system. Based on the topological connection relationship of each human body part, determine the quaternion to be decomposed from the quaternions matched by each human body part, and determine the decomposition direction of the quaternion to be decomposed. The quaternion to be decomposed is decomposed according to the decomposition direction, and the axial rotation terms that match the quaternion to be decomposed are determined. Based on the axial rotation terms matched by the quaternions to be decomposed, determine the motion capture correction parameters for each human body part.

6. The method according to claim 2, characterized in that, After determining the calibration locations of at least two human body parts, the method further includes: Based on the calibrated positions of each body part, the motion capture images are corrected and updated.

7. The method according to claim 6, characterized in that, The process of correcting and updating the motion capture images of each body part based on their calibrated locations includes: Based on the relative positional relationship of the calibrated positions of each human body part, determine the perspective transformation matrix matching each human body part; Based on the perspective transformation matrix matched to each human body part, the motion capture images are corrected and updated.

8. A visual motion capture device, characterized in that, include: The motion capture image acquisition module is used to acquire scene capture images and determine motion capture images of at least two human body parts based on the scene capture images. The motion capture parameter determination module is used to input each motion capture image into the recognition model that matches each human body part to determine the motion capture parameters of each human body part. The correction parameter determination module is used to correct the motion capture parameters of each human body part according to the topological connection relationship of each human body part, so as to obtain the motion capture correction parameters of each human body part. The motion capture result determination module is used to determine the motion capture result based on the motion capture correction parameters for each part of the human body.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a visual motion capture method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute a visual motion capture method according to any one of claims 1-7.