Multi-algorithm fusion motion capture algorithm and system based on statistical correction

Through a motion capture solution combining multi-camera and inertial navigation, the data fusion of vision sensors and inertial navigation sensors is used to solve the accuracy and cost problems of existing equipment, and achieve low-cost and high-precision motion capture effect.

CN120279146APending Publication Date: 2025-07-08JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510345452.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing motion capture equipment has problems such as under-precision and excessive cost in the civilian field, which is difficult to meet the needs of CG animation producers.

Method used

Using a solution that combines multi-camera and inertial navigation, video stream data is collected through vision sensors, combined with deep learning algorithms and inertial navigation or indoor positioning sensors, data preprocessing and fusion are performed, and weights are adaptively assigned to eliminate jitter and offset of algorithm results, and improve motion capture accuracy.

Benefits of technology

While reducing costs, it improves the accuracy and ease of use of motion capture, meeting the needs of CG animation makers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279146A_ABST
    Figure CN120279146A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-algorithm fusion motion capture algorithm and system based on statistical correction, starting from the perspective of a target user of a low-cost motion capture device, referring to the idea of multiple cameras of a commercial motion capture device, and using a main vision sensor to collect video stream data; preprocessing the video stream data through an image motion capture algorithm based on deep learning; a combined scheme of vision and inertial navigation is tried, inertial navigation or an indoor positioning sensor is used for collecting motion acceleration data, and a filtering algorithm is used for preprocessing original data; weights are given to data obtained by multiple devices and multiple algorithms in a self-adaptive mode, jitter and offset of certain algorithm results are eliminated, and a more accurate motion capture result is obtained. By providing a set of motion capture solution which is easy to use and low in cost and combines multiple cameras with inertial navigation, an unreliable Z-axis capture result under a single camera is perfected, and the precision of motion capture is ensured while the cost is reduced, so that the requirements of CG animation producers are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communication technologies, and in particular, to an action capture algorithm and system based on statistical correction and multi-algorithm fusion, as well as an electronic device. Background Art

[0002] Action capture technology was first used in the film industry for character motion production at the beginning of this century and has since been widely applied in the CG field. These commercial production fields have extremely high requirements for the accuracy of action capture and are insensitive to costs. With the progress of CG technology and the development of the live broadcast industry, there has also emerged a demand for action capture in the civilian field, such as virtual anchors, VR, and personal CG production. Different from the traditional film and television industry, these industries generally have general requirements for the accuracy of action capture but are very sensitive to prices. The cost of traditional action capture devices, which can easily reach hundreds of thousands of yuan, is unaffordable for these users. In the face of this new demand, there have emerged various products based on different technologies on the market, such as Google's open-source solution Mediapipe using a single camera + AI, ReboCap using wearable device inertial navigation and geomagnetic sensing positioning, and Microsoft's depth camera Kinect... These solutions all have their respective defects, such as unqualified accuracy, unsatisfactory capture frame rate, and still overly high costs, and still cannot well meet the market demand.

[0003] Therefore, it is urgent to provide an easy-to-use and low-cost action capture solution that combines multiple cameras and inertial navigation from the perspective of the target users of low-cost action capture devices to meet the needs of CG animators. Summary of the Invention

[0004] To solve the above problems, the present application proposes an action capture algorithm and system based on statistical correction and multi-algorithm fusion, as well as an electronic device.

[0005] On the one hand, the present application proposes an action capture algorithm based on statistical correction and multi-algorithm fusion, which is implemented based on a multi-modal motion capture device and includes the following steps:

[0006] S1. Obtain the action pose data collected by the multi-modal motion capture device;

[0007] S2. Extract the action state feature values in the action pose data in each modality;

[0008] S3. Statistically analyze the action state feature values in each modality and perform fusion to obtain the 3D coordinate data of the final action state;

[0009] S4. Send the 3D coordinate data of the final action state to the background for backend processing to generate a corresponding 3D action rendering animation.

[0010] As an alternative implementation of the present application, optionally, in S2, extracting the action state feature values from the action posture data in each modality includes:

[0011] Extracting the frame images from the action posture data captured by the motion capture cameras in each modality;

[0012] Using the CNN deep learning algorithm to extract the coordinates of the target action points in the frame images within the preset XY plane, and denoting them as (X pi , Y pi );

[0013] Performing a field of view process on the marked coordinates to expand them into a three-dimensional vector S:

[0014] S = (-X pi * tanθ i , -Y pi * tanθ i , 1);

[0015] Estimating the Z coordinate of the target action point:

[0016]

[0017] Performing a mean calculation on Z Xi and Z Yi to obtain the Z coordinate of the target action point;

[0018] Fusing and calculating the three-dimensional coordinates of the target action point: S·Z, as the visual result of the target action point;

[0019] where i represents the i-th motion capture camera, Pi represents the camera coordinates, D i represents the orientation vector, R i represents the right vector, θ is the camera field of view angle obtained by the camera device software, and A spi is the aspect ratio obtained by the camera device software.

[0020] As an alternative implementation of the present application, optionally, in S2, extracting the action state feature values from the action posture data in each modality includes:

[0021] Collecting the inertial navigation / indoor positioning data generated by the corresponding target action point during motion capture through an inertial navigation / indoor positioning sensor;

[0022] Extracting the inertial result from the inertial navigation / indoor positioning data;

[0023] Calculating the similarity between the inertial result and the visual result at the corresponding time point, and binding the two when the similarity meets the threshold;

[0024] Determine whether the inertial result or the visual result at the corresponding time point needs to be supplemented:

[0025] If so, fuse the inertial result with the visual result at the corresponding time point for result calibration;

[0026] Otherwise, discard it.

[0027] As an optional implementation of the present application, optionally, in S3, statistically calculate the action state eigenvalue in each modality and perform fusion to obtain the 3D coordinate data of the final action state, including:

[0028] Based on the real-time coordinates of the target action point monitored by motion capture, obtain its real-time state parameters, including: real-time position and real-time speed;

[0029] Calculate the position gain based on the real-time position;

[0030] Correct the real-time state parameters according to the position gain and the observation value at the corresponding position;

[0031] Calculate the corrected real-time position;

[0032] Based on the above steps, predict and output the predicted position at the next time point.

[0033] On the other hand, the present application proposes an action capture system based on statistical correction and multi-algorithm fusion for implementing the action capture algorithm based on statistical correction and multi-algorithm fusion described above, including:

[0034] A data acquisition unit for acquiring action posture data collected by a multi-modal motion capture device;

[0035] A data processing unit for extracting the action state eigenvalue in the action posture data in each modality;

[0036] A data fusion unit for statistically calculating the action state eigenvalue in each modality and performing fusion to obtain the 3D coordinate data of the final action state;

[0037] A post-processing unit for sending the 3D coordinate data of the final action state to the background for backend processing to generate a corresponding 3D action rendering animation.

[0038] On the other hand, the present application also proposes an electronic device, including:

[0039] A processor;

[0040] A memory for storing instructions executable by the processor;

[0041] Wherein, when the processor is configured to execute the executable instructions, it implements the action capture algorithm based on statistical correction and multi-algorithm fusion as described above.

[0042] Technical effects of the present invention:

[0043] Based on the perspective of the target users of low-cost action capture devices, this application draws on the idea of multi-cameras in commercial action capture devices, uses the main vision sensor to collect video stream data, and preprocesses the video stream data through an image action capture algorithm based on deep learning. It attempts a combination scheme of vision + inertial navigation, uses an inertial navigation or indoor positioning sensor to collect motion acceleration data, and preprocesses the raw data through a filtering algorithm. It adaptively assigns weights to the data obtained from multiple devices and multiple algorithms, eliminates the jitter and offset of some algorithm results, and obtains a more accurate action capture result. By proposing an easy-to-use and low-cost action capture solution combining multi-cameras and inertial navigation, it improves the unreliable Z-axis capture results under a single camera. While reducing costs, it ensures the accuracy of action capture to meet the needs of CG animation producers.

[0044] Other features and aspects of the present disclosure will become clear from the following detailed description of the exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings included in and constituting a part of this specification, together with the specification, illustrate the exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0046] Figure 1 It is shown as a schematic diagram of the implementation process of the present invention;

[0047] Figure 2 It is shown as a schematic diagram of the application process of the present invention;

[0048] Figure 3 It is shown as a schematic diagram of the system composition structure of the present invention;

[0049] Figure 4 It is shown as an application schematic diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0051] The special term "exemplary" here means "serving as an example, embodiment, or illustrative". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.

[0052] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0053] Embodiment 1

[0054] As Figure 1 shown, on the one hand, the present application proposes an action capture algorithm based on the fusion of multiple algorithms with statistical correction, which is implemented based on a multi-modal motion capture device and includes the following steps:

[0055] S1. Obtain the action posture data collected by the multi-modal motion capture device;

[0056] S2. Extract the action state feature values in the action posture data in each modality;

[0057] S3. Statistically analyze the action state feature values in each modality and perform fusion to obtain the 3D coordinate data of the final action state;

[0058] S4. Send the 3D coordinate data of the final action state to the background for backend processing to generate the corresponding 3D action rendering animation.

[0059] The following will be specifically described in combination with the application flow chart shown in the attached Figure 2 figure.

[0060] The present invention mainly collects action data through a multi-modal motion capture device. After the extraction and fusion of multi-modal motion capture feature data, the state values integrating the feature data of multiple modalities are finally obtained, and these are output as the 3D coordinate data of the motion capture. The above operations are performed on each action point. Finally, through rendering, the corresponding 3D action rendering animation is generated, thereby adaptively weighting the data obtained by multiple devices and multiple algorithms, eliminating the jitter and offset of some algorithm results, and obtaining a more accurate action capture result. The present invention can provide a user-friendly and low-cost action capture solution combining multiple cameras and inertial navigation for users, ensuring the accuracy of action capture while reducing costs.

[0061] Motion capture devices of each modality, such as 3D cameras, collect the action parameters of each action point of the user (it is necessary to select target action points, such as position points of the elbow joint, knee, etc.). For example, inertial navigation devices or indoor positioning sensors can collect the corresponding parameters of the user. The arrangement of these devices can be implemented in combination with the existing motion capture technology.

[0062] In this embodiment, the corresponding device software system, such as the camera device software of a 3D camera, can be matched with the hardware device facilities selected by the user.

[0063] As an alternative implementation of this application, optionally, in S2, extracting the action state feature values in the action pose data in each modality includes:

[0064] Extracting the frame images in the action pose data captured by the motion capture cameras in each modality;

[0065] Using the CNN deep learning algorithm to extract the coordinates of the target action points in the frame images within the preset XY plane, and denoting them as (X pi , Y pi );

[0066] Performing a field of view process on the marked coordinates and expanding them into a three-dimensional vector S:

[0067] S = (-X pi * tanθ i , -Y pi * tanθ i , 1);

[0068] Estimating the Z coordinate of the target action point:

[0069]

[0070] Performing a mean calculation on Z Xi and Z Yi to obtain the Z coordinate of the target action point;

[0071] Fusing and calculating the three-dimensional coordinates of the target action point: S · Z, as the visual result of the target action point;

[0072] where i represents the i-th motion capture camera, Pi represents the camera coordinate, D i represents the orientation vector, R i represents the right vector, θ is the camera field of view angle obtained by the camera device software, and A spi is the aspect ratio obtained by the camera device software.

[0073] The present invention attempts to provide an action capture algorithm for multiple cameras. In existing action capture solutions using vision, most use a single camera + AI. In the tests of the present invention, the results captured by this solution have acceptable accuracy in the projection plane, but the accuracy of its Z-axis is quite unreliable and is basically in an unusable state. The present invention hopes to draw on the idea of multiple cameras in commercial action capture devices and use multiple cameras + existing AI technologies to improve the unreliable Z-axis capture results under a single camera.

[0074] The specific process is as follows:

[0075] It is necessary for the user to set multiple fixed camera devices to measure the relative positional relationship between the devices. Take one of the camera devices as the main camera, which is the origin of a right-handed coordinate system, and record its camera direction as the positive direction of the Z-axis. Calculate the coordinates of the remaining camera devices as P i , with the orientation vector being D i , and the right vector being R i . Here, i represents the i-th camera except the main camera.

[0076] For example: There is a camera device located at the front right of the main camera at the same height as the main camera, and its orientation is pointing 10 m directly in front of the main camera. Then its position P is (5, 0, 5), the forward vector D is and the right vector is

[0077] Obtain its camera field of view angle θ and aspect ratio A through the camera device software spi .

[0078] Perform pose recognition on each frame of each device to obtain the normalized coordinates of each bone point in the camera plane, denoted as (X pi , Y pi ). Process the coordinates of each point of the main camera in the field of view and expand them into a three-dimensional vector S(-X p0 * tanθ0, -Y p0 * tanθ0, 1).

[0079] Denote Z as the estimated value of the Z coordinate in the coordinate system calculated from each bone point. Each camera device can calculate two estimated values of Z based on the X and Y coordinates, namely Z Xi and Z Yi . Estimate the Z coordinate of the target action point:

[0080]

[0081] Perform a mean calculation on Z Xi and Z Yi to obtain the Z coordinate of the target action point;

[0082] Take the average of several Z estimated values calculated by multiple devices as the Z coordinate of each bone point calculated by the system. Its final three-dimensional coordinate is S · Z.

[0083] For the inertial solution devices currently available on the market, every about 30 minutes of use will accumulate a non-negligible error and require recalibration. If combined with the vision solution, the error can be automatically eliminated during use by the calculation results of the video solution, without manual calibration.

[0084] The present invention attempts a combined solution of vision + inertial navigation. Wearable inertial navigation devices are also a popular solution. However, the accuracy of inertial navigation chips under cost constraints is also unreliable. Drift may occur even with a slightly larger movement amplitude, and recalibration is required. Neither the cost, accuracy, nor ease of use is excellent. Therefore, the present invention also attempts a combined solution of vision + inertial navigation. The inertial navigation device is calibrated in real time according to the vision capture results that do not drift, and at the same time, the problems of loss and insufficient frame rate under vision capture occlusion are compensated. The two solutions complement each other, and both the accuracy and ease of use are sufficiently improved, which can be an option for users with a higher budget.

[0085] In this specification, the specific explanation of complementarity is as follows: When using both the vision and inertial solutions at the same time, after starting the capture, according to the similarity between the vision results and the inertial results, the inertial device is automatically bound to the moving limb, and the lengths of each level of the limb are calculated simultaneously, eliminating the trouble of manually binding and manually measuring to fill in the limb length data in the traditional solution.

[0086] As an alternative implementation of this application, optionally, S2. Extract the action state eigenvalue in the action posture data in each modality, including:

[0087] Collect the inertial navigation / indoor positioning data generated by the corresponding target action point during movement during action capture through an inertial navigation / indoor positioning sensor;

[0088] Extract the inertial results from the inertial navigation / indoor positioning data;

[0089] Calculate the similarity between the inertial results and the vision results at the corresponding time point. When the similarity meets the threshold, bind the two;

[0090] Judge whether the inertial results or the vision results at the corresponding time point need to be supplemented:

[0091] If so, fuse the inertial results and the vision results at the corresponding time point for result calibration;

[0092] Otherwise, abandon.

[0093] During the specific operation, when the frame rate of the video solution is lower than the output requirement (for example, the frame rate of the video solution is lower than the preset value of 45), the inertial solution will be used for output to interpolate the results of the video solution to keep up with the output requirement. When some bone points are lost in the video solution (for example, when blocked or out of the camera range), the inertial solution is also used for output to make up for it, avoiding frame loss and incoherence of the action.

[0094] Examples are as follows:

[0095] I. Automatic device binding and limb length calculation

[0096] 1. Inertial device is bound to the moving limb

[0097] Core logic: Automatically associate the device with the limb by matching the similarity of visual and inertial data in the initial motion phase (e.g., when the user performs a standard action).

[0098] Algorithm steps:

[0099] Data synchronization: Align the timestamps of the visual system (30Hz) and the inertial sensor (100Hz), and use a sliding window for dynamic calibration.

[0100] Feature extraction: Extract the displacement vectors of visual joint points (such as shoulders, elbows, wrists), and the acceleration and angular velocity sequences of the inertial sensor.

[0101] Similarity calculation:

[0102] Dynamic Time Warping (DTW): Calculate the minimum path distance between the visual displacement sequence and the acceleration sequence of each inertial device.

[0103] Formula:

[0104] DTW(V, I) = min π ∑ (i,j)∈π / / V i -I j / / 2

[0105] where V is the visual trajectory, I is the inertial data, and π is the alignment path. i and j are the numbers corresponding to the visual camera or inertial sensor respectively.

[0106] Device allocation: Select the inertial device with the minimum DTW distance and bind it to the corresponding limb.

[0107] For example, when the user waves the right arm, the visual system captures the right elbow displacement sequence V = [v1, v2,..., vn], and the data of 3 inertial devices are I1, I2, I3 respectively. The calculated DTW distances are [D1 = 12.3, D2 = 3.2, D3 = 15.7], then device 2 is bound to the right arm.

[0108] 2. Automatic calculation of limb length

[0109] Core logic: Use the positions of visual joint points and the directions of inertial sensors to reverse-derive through a geometric model.

[0110] Algorithm steps:

[0111] Initial frame calibration: The user holds the T-pose, and the visual system measures the absolute coordinates of the joint points (such as the right shoulder Ps = (xs, ys, zs), the right wrist Pw = (xw, yw, zw)).

[0112] Calculation of limb length Larm:

[0113] Formula:

[0114]

[0115] Dynamic verification: During movement, the inertial sensor verifies the consistency of limb length through direction changes. If the deviation exceeds the threshold (e.g., ±2%), recalibration is triggered.

[0116] II. Low-frame-rate video interpolation

[0117] 1. Processing of insufficient frame rate (e.g., video 30Hz → output 45Hz)

[0118] Core logic: Insert the intermediate poses deduced from inertial data between adjacent visual frames.

[0119] Algorithm steps:

[0120] Integration of inertial data: Based on a sampling rate of 100Hz, deduce the bone rotation and displacement through angular velocity ω and acceleration a.

[0121] Pose update:

[0122]

[0123] Position deduction:

[0124] p t and v t are the position and velocity at time point t.

[0125] Generate intermediate frames by interpolation: Insert the intermediate pose t1.5 deduced from inertial data between visual frames t1 and t2.

[0126] Smoothing and optimization: Use spherical linear interpolation (Slerp) to fuse visual and inertial data to avoid sudden changes.

[0127] III. Compensation for lost bone points

[0128] 1. Processing of occlusion or out-of-bounds

[0129] Core logic: When the visual system loses a joint point (e.g., the right hand is occluded), switch to inertial data-driven.

[0130] Algorithm steps:

[0131] Loss detection: Mark as lost when the visual confidence C < 0.2 (in the range of 0 - 1).

[0132] Takeover of inertial data: Use the data of the bound inertial device to deduce the joint position.

[0133] Inverse Kinematics (IK): Combining limb length constraints to calculate the positions of occluded joints.

[0134] Data fusion recovery: After visual recovery, smooth transition through Kalman filtering.

[0135] The present invention attempts to develop a statistics-based fusion algorithm. The present invention attempts to develop a statistics-based fusion algorithm that can adaptively weight the data obtained from multiple devices and multiple algorithms, eliminate the jitter and offset of certain algorithm results, obtain more accurate results, and can also use the statistical data for feedback calibration, such as the real-time calibration of the inertial navigation device mentioned above.

[0136] As an alternative implementation of the present application, optionally, S3. Statistically analyze the action state eigenvalue in each modality and perform fusion to obtain the 3D coordinate data of the final action state, including:

[0137] Based on the real-time coordinates of the target action points monitored by motion capture, obtain its real-time state parameters, including: real-time position and real-time speed;

[0138] Calculate the position gain based on the real-time position;

[0139] Correct the real-time state parameters based on the position gain and the observed value at the corresponding position;

[0140] Calculate the corrected real-time position;

[0141] Based on the above steps, predict and output the predicted position at the next time point.

[0142] The fusion algorithm of the present invention will be described below by taking the data of the IMU inertial sensor as an example.

[0143] Take the data of the IMU inertial sensor as the predicted value and the data of the video stream as the observed value, correct the predicted value with the observed value, perform algorithm fusion filtering, and give the optimal estimate.

[0144] During dynamic capture, three-dimensional parameter calculation is performed according to the coordinates at different time points. Assume that the position of the IMU at time t-1 is p t-1 , the speed is v t-1 , and the acceleration is u. It is easy to obtain the position and speed at time t:

[0145] v t = v t-1 + uΔt (3)

[0146]

[0147] where Δt is the time difference between time t and t-1. The above formula is combined and written in matrix form:

[0148]

[0149] Define the state of the IMU:

[0150]

[0151] Define the state transition matrix:

[0152]

[0153] Define the control matrix:

[0154]

[0155] During the prediction process, a process error will be generated. Denote the process error matrix as w t , and during the observation process, an observation error will be generated. Denote the observation error matrix as o t . Therefore, the standard state quantity expression is:

[0156] x t = Fx t-1 + Bu + w t (9)

[0157] Next, model the observation process. Assume that the observed quantity z t is obtained by linearly transforming the current true state by H and adding the observation error o t . That is:

[0158] z t = Hx t + o t (10)

[0159] Denote the covariance matrices of the errors as Q and R:

[0160] Q = E[ww T (11)

[0161] R = E[oo T (12)

[0162] The IMU itself has systematic and random errors. However, through the previous methods, the present invention has eliminated the errors to the greatest extent. Therefore, the present invention is confident enough to make w t equal to 0 and define R as 1.

[0163] The following is the specific explanation of the fusion algorithm in this specification:

[0164] Denote x t as the true value of the state, as the predicted value of the state, is the estimated value of the state, and the relationship among the three is: Estimated value = Predicted value + Weight * Observation error.

[0165] The predicted value is the same as in Equation (9). Since the estimated value at the previous moment already includes the error, therefore:

[0166]

[0167] Denote K t as the weight at time t, then the relationship is:

[0168]

[0169] Rewrite the above formula as:

[0170]

[0171] For the convenience of subsequent description, the present invention defines the difference between the true value and the predicted value as the prior error: Define the difference between the true value and the estimated value as the posterior error: Therefore, the above formula can be written as:

[0172]

[0173] where I is the identity matrix.

[0174] Next, the present invention respectively defines the covariance matrices of the prior error and the posterior error as and P t :

[0175]

[0176] Substitute the expression of e t (i.e., Equation (16)) into P t to obtain:

[0177]

[0178] Since and o t are uncorrelated and the expectation is 0, the above formula can be simplified to:

[0179]

[0180] Further disassembled into:

[0181]

[0182] In practical applications, will be automatically adjusted in the state update of the algorithm fusion, so directly set the initial value to 1. The fusion algorithm of the present invention is based on the MSE criterion, making the posterior error P tis minimized. According to the MSE criterion, the trace of P t can be selected as the objective function T(P t ), and its derivative is calculated to obtain the minimum value.

[0183] Through the derivative formula:

[0184]

[0185] Let the derivative be 0, and when the posterior error is minimized:

[0186]

[0187] When the above equation holds, K t is the weight that minimizes the posterior error. Substituting K t into Equation 20 gives:

[0188]

[0189] Next, derive the

[0190] required for the next iteration. Subtracting Equation 13 from Equation 9 gives:

[0191]

[0192] Taking the covariance on both sides, and since and o t are uncorrelated, we get:

[0193]

[0194] Thus, it prepares the parameters required for the next derivation.

[0195] In summary, the overall process of fusing different algorithms using the fusion algorithm is as follows:

[0196]

[0197] Obviously, those skilled in the art should understand that to implement all or part of the processes in the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above control embodiments. Those skilled in the art can understand that to implement all or part of the processes in the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above control embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0198] Embodiment 2

[0199] As Figure 3 shown, based on the implementation principle of Embodiment 1, on the other hand, this application proposes a multi-algorithm fusion motion capture system based on statistical correction, which is used to implement the motion capture algorithm based on statistical correction described above, including:

[0200] A data acquisition unit, which is used to obtain the motion posture data collected by the multi-modal motion capture device;

[0201] A data processing unit, which is used to extract the motion state feature values in the motion posture data in each modality;

[0202] A data fusion unit, which is used to statistically analyze the motion state feature values in each modality and perform fusion to obtain the 3D coordinate data of the final motion state;

[0203] A post-processing unit, which is used to send the 3D coordinate data of the final motion state to the background for backend processing to generate the corresponding 3D motion rendering animation.

[0204] For the functions of the above-mentioned units and their interactions, please refer to the corresponding steps in Embodiment 1 for understanding and implementation, and will not be elaborated here.

[0205] Each module or step of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by a computing device. Thus, they can be stored in a storage device and executed by a computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0206] Embodiment 3

[0207] As Figure 4 shown, further, on the other hand, the present application also proposes an electronic device, including:

[0208] A processor;

[0209] A memory for storing instructions executable by the processor;

[0210] Wherein, when the processor is configured to execute the executable instructions, it implements the action capture algorithm of multi-algorithm fusion based on statistical correction described above.

[0211] The electronic device according to the embodiments of the present disclosure includes a processor and a memory for storing instructions executable by the processor. Wherein, when the processor is configured to execute the executable instructions, it implements the action capture algorithm of multi-algorithm fusion based on statistical correction described in any one of the foregoing.

[0212] Here, it should be noted that the number of processors can be one or more. At the same time, in the electronic device according to the embodiments of the present disclosure, an input device and an output device can also be included. Wherein, the processor, the memory, the input device and the output device can be connected through a bus or in other ways, which is not specifically limited herein.

[0213] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs and various modules, such as: the programs or modules corresponding to the action capture algorithm of multi-algorithm fusion based on statistical correction according to the embodiments of the present disclosure. The processor executes various functional applications and data processing of the electronic device by running the software programs or modules stored in the memory.

[0214] The input device can be used to receive input numbers or signals. Wherein, the signal can be a key signal related to the user settings and function control of the device / terminal / server. The output device can include a display device such as a display screen.

[0215] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary technical personnel in the technical field to understand the embodiments disclosed herein.

Claims

1. A motion capture algorithm based on the fusion of multiple algorithms with statistical correction, which is implemented based on a multi-modal motion capture device, and is characterized in that, It includes the following steps: S1. Obtain the action pose data collected by the multi-modal motion capture device; S2. Extract the action state feature values from the action pose data in each modality; S3. Statistically analyze the action state feature values in each modality and perform fusion to obtain the 3D coordinate data of the final action state; S4. Send the 3D coordinate data of the final action state to the background for backend processing to generate the corresponding 3D action rendering animation.

2. The action capture algorithm based on multi-algorithm fusion with statistical correction according to claim 1, wherein S2. Extract the action state feature values from the action pose data in each modality, including: Extract the frame images from the action pose data captured by the motion capture cameras in each modality; Using the CNN deep learning algorithm, extract the coordinates of the target action points in the frame image within the preset XY plane, and denote them as (X pi , Y pi ); Perform a field-of-view process on the marked coordinates and expand them into a three-dimensional vector S: S = (-X pi *tanθ i , -Y pi *tanθ i , 1); Estimate the Z coordinate of the target action point: For Z Xi and Z Yi perform mean calculation to obtain the Z coordinate of the target action point; Fusion-compute the three-dimensional coordinates of the target action point: S·Z, as the visual result of the target action point; Among them, i represents the i-th motion capture camera, Pi represents the camera coordinates, D i represents the orientation vector, R i represents the right vector, θ is the camera field of view angle obtained by the camera device software, A spi is the aspect ratio obtained by the camera device software.

3. The action capture algorithm based on multi-algorithm fusion with statistical correction according to claim 2, wherein S2. Extract the action state feature values from the action pose data in each modality, including: Collect the inertial navigation / indoor positioning data generated by the corresponding target action point during motion capture through the inertial navigation / indoor positioning sensor; Extract the inertial results from the inertial navigation / indoor positioning data; Calculate the similarity between the inertial results and the visual results at the corresponding time points, and bind the two when the similarity meets the threshold; Judge whether the inertial results or the visual results at the corresponding time points need to be supplemented: If so, fuse the inertial results and the visual results at the corresponding time points for result calibration; Otherwise, discard.

4. A motion capture algorithm based on statistical correction and multi-algorithm fusion according to claim 2, characterized in that, S3. Statistically analyze the action state feature values in each modality and perform fusion to obtain the 3D coordinate data of the final action state, including: Obtain the real-time state parameters according to the real-time coordinates of the target action point monitored by the motion capture, including: real-time position and real-time speed; Calculate the position gain according to the real-time position; Correct the real-time state parameters according to the position gain and the observed values at the corresponding positions; Calculate the corrected real-time position; Based on the above steps, predict and output the predicted position at the next time point.

5. A multi-algorithm fusion motion capture system based on statistical correction is used to implement the motion capture algorithm of multi-algorithm fusion based on statistical correction described in any one of claims 1-4, and is characterized in that, It includes: A data acquisition unit for obtaining the action pose data collected by the multi-modal motion capture device; A data processing unit for extracting the action state feature values from the action pose data in each modality; A data fusion unit for statistically analyzing the action state feature values in each modality and performing fusion to obtain the 3D coordinate data of the final action state; A post-processing unit for sending the 3D coordinate data of the final action state to the background for backend processing to generate the corresponding 3D action rendering animation.

6. An electronic device, characterized in that, It includes: A processor; A memory for storing the executable instructions of the processor; Wherein, the processor is configured to implement the action capture algorithm based on multi-algorithm fusion with statistical correction according to any one of claims 1-4 when executing the executable instructions.