Methods, apparatuses, electronic devices and storage media for processing object motion data

By acquiring the image frame sequence and pose data of the target object, and using error constraint processing technology to optimize the predicted action data, the problem of inaccurate action data prediction in monocular vision is solved, and more accurate action data acquisition and driving are achieved.

CN116091622BActive Publication Date: 2026-03-10TSINGHUA UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-03
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Monocular vision suffers from inaccurate prediction of object motion data due to visual occlusion.

Method used

By acquiring image frame sequences containing the target object and their corresponding object pose data, the pose measurement module is used to perform error constraint processing, including key point detection, motion-driven processing, and coordinate system transformation, to optimize the predicted motion data.

Benefits of technology

It improves the accuracy and comprehensiveness of object motion data, compensates for the visual occlusion problem in monocular vision, and achieves more accurate motion representation and driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091622B_ABST
    Figure CN116091622B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, electronic device, and storage medium for processing object motion data. The method includes: acquiring a sequence of image frames containing a target object, and object pose data corresponding to each image frame in the image frame sequence; the object pose data is obtained based on a pose measurement module disposed on the target object; performing motion data prediction on the target object in the target image frames to obtain predicted motion data corresponding to the target image frames; the target image frames are image frames in the image frame sequence that have not undergone motion data prediction; and performing error constraint processing on the predicted motion data corresponding to the target image frames based on the target image frames and the object pose data corresponding to the target image frames to obtain target motion data corresponding to the target image frames. This disclosure can improve the accuracy of determining the motion data of the target object, thereby improving the accuracy of motion driving.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to an object action data processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] Object pose estimation is an important research field of computer vision, and the purpose is to estimate the object pose based on given input information, which can be applied to many application scenarios, such as action recognition, action detection, movies, animation, virtual reality, human-computer interaction, motion analysis, etc.

[0003] In the related art, object action data prediction can generally be performed based on image information obtained by a monocular camera system. Since monocular vision may exist visual occlusion, the object action data prediction is inaccurate. SUMMARY

[0004] The present disclosure provides an object action data processing method and device, electronic equipment and storage medium to at least solve the problem of inaccurate object action data prediction in the related art. The technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, an object action data processing method is provided, comprising:

[0006] obtaining an image frame sequence containing a target object, and object pose data corresponding to each image frame in the image frame sequence; the object pose data is obtained based on a pose measurement module arranged on the target object;

[0007] performing action data prediction on the target object in a target image frame to obtain predicted action data corresponding to the target image frame; the target image frame is an image frame in the image frame sequence that has not been subjected to action data prediction;

[0008] performing error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame to obtain target action data corresponding to the target image frame.

[0009] In an exemplary embodiment, the error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame to obtain target action data corresponding to the target image frame comprises:

[0010] performing action driving processing based on the predicted action data to obtain predicted driving data; the predicted driving data is observation data corresponding to the predicted action data;

[0011] perform error constraint processing on the predicted driving data based on the target image frame and object pose data corresponding to the target image frame, to obtain target driving data;

[0012] determine the target action data based on the target driving data.

[0013] In an example embodiment, the error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and object pose data corresponding to the target image frame, to obtain target action data corresponding to the target image frame, includes:

[0014] perform key point detection based on the target image frame, to obtain detection key point information of the target object;

[0015] perform action driving processing based on the predicted action data, to obtain predicted key point information of the target object;

[0016] perform error constraint processing on the predicted action data based on error information between the detection key point information and the predicted key point information, and object pose data corresponding to the target image frame, to obtain the target action data.

[0017] In an example embodiment, the detection key point information includes detection confidence of each of the plurality of target key points;

[0018] the error constraint processing on the predicted action data based on error information between the detection key point information and the predicted key point information, and object pose data corresponding to the target image frame, to obtain the target action data, includes:

[0019] determine a target confidence threshold corresponding to the target image frame based on detection confidence of each of the plurality of target key points in the target image frame, detection confidence of each of the plurality of target key points in a historical image frame, and a smoothing confidence threshold corresponding to the historical image frame; the historical image frame is an image frame in the image frame sequence that is located in time sequence before the target image frame; the smoothing confidence threshold is determined based on average confidence of the plurality of target key points in the historical image frame;

[0020] determine first weights of each of the plurality of target key points in the target image frame based on a preset confidence threshold, a target confidence threshold corresponding to the target image frame, and detection confidence of each of the plurality of target key points in the target image frame; the preset confidence threshold is less than the target confidence threshold corresponding to the target image frame;

[0021] Based on the first weights corresponding to the multiple target key points, the error information of the detected key point information and the predicted key point information is weighted to generate first weighted error information.

[0022] Based on the first weighted error information and the object pose data corresponding to the target image frame, the predicted action data is subjected to error constraint processing to obtain the target action data.

[0023] In an exemplary embodiment, the step of performing error constraint processing on the predicted motion data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame to obtain target motion data corresponding to the target image frame includes:

[0024] The object pose data corresponding to the target image frame is transformed from the inertial coordinate system to the global coordinate system to obtain the target skeleton orientation data of the target object in the global coordinate system;

[0025] Based on the predicted motion data, motion-driven processing is performed to obtain the predicted bone orientation data of the target object in the global coordinate system.

[0026] Based on the target bone orientation data and the predicted bone orientation data, bone orientation error information is generated;

[0027] Based on the target image frame and the bone orientation error information, the predicted motion data is subjected to error constraint processing to obtain the target motion data.

[0028] In an exemplary embodiment, the step of performing error constraint processing on the predicted motion data based on the target image frame and the bone orientation error information to obtain the target motion data includes:

[0029] Keypoint detection is performed on the target image frame to obtain keypoint information of the target object; the keypoint information includes the detection confidence of each of the multiple target keypoints; the multiple skeletons of the target object in the target image frame correspond to the multiple target keypoints.

[0030] If the detection confidence of the target key point corresponding to any bone is less than the target confidence threshold, a second weight corresponding to the target key point corresponding to the target key point corresponding to the target key point corresponding to the target key point is determined; the second weight is greater than the third weight.

[0031] Based on the second weight corresponding to any one of the bones, the bone orientation error information is weighted to generate second weighted error information.

[0032] Based on the third weight corresponding to the target key point of any skeleton, the error information between the detection information of the target key point and the prediction information of the target key point is weighted to obtain the third weighted error information; the prediction information of the target key point is obtained by performing motion-driven processing on the predicted motion data.

[0033] The predicted action data is subjected to error constraint processing based on the second weighted error information and the third weighted error information to obtain the target action data.

[0034] In one exemplary embodiment, the predicted motion data includes object shape prediction data of the target object;

[0035] The step of performing error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame to obtain target action data corresponding to the target image frame includes:

[0036] Based on the object shape prediction data of the target object in the target image frame and the shape smoothing data corresponding to the historical image frames, shape error information is determined; the historical image frames are image frames in the image frame sequence that are temporally preceding the target image frame; the shape smoothing data is obtained by temporally smoothing the object shape data of the target object in the historical image frames.

[0037] Based on the shape error information, the target image frame, and the object pose data corresponding to the target image frame, error constraint processing is performed on the predicted motion data to obtain the target motion data.

[0038] In one exemplary embodiment, the predicted motion data includes joint prediction data of the target object;

[0039] The step of performing error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame to obtain target action data corresponding to the target image frame includes:

[0040] Based on the joint prediction data corresponding to the target image frame and the joint constraint data corresponding to the previous image frame of the target image frame, joint error information is obtained; the joint constraint data corresponding to the previous image frame is obtained by performing data error constraints on the joint prediction data of the target object in the previous image frame; the previous image frame is an image frame that is temporally preceding the target image frame in the image frame sequence and is adjacent to the target image frame.

[0041] The fourth weight is determined based on the motion data loss information corresponding to the previous image frame; the motion data loss information corresponding to the previous image frame is determined based on the object pose data and the constraint driving data corresponding to the previous image frame; the constraint driving data corresponding to the previous image frame is obtained by constraining the predicted motion data corresponding to the previous image frame with data error.

[0042] The joint error information is weighted based on the fourth weight to generate fourth weighted error information.

[0043] Based on the fourth weighted error information, the joint prediction data is iteratively updated to obtain the joint constraint data corresponding to the target image frame;

[0044] The predicted motion data is obtained by performing error constraint processing on the target image frame, the object pose data corresponding to the target image frame, and the joint constraint data corresponding to the target image frame.

[0045] In an exemplary embodiment, before determining the fourth weight based on the motion data loss information corresponding to the previous image frame, the method further includes:

[0046] Determine the constraint driving data corresponding to the target action data of the previous image frame;

[0047] Based on the object pose data corresponding to the previous image frame and the constraint driving data, determine the motion data loss information corresponding to the previous image frame;

[0048] The determination of the fourth weight based on the motion data loss information corresponding to the previous image frame includes:

[0049] The fourth weight is determined based on the motion data loss information corresponding to the previous image frame and the loss smoothing information corresponding to the previous image frame; the loss smoothing information corresponding to the previous image frame is obtained by performing temporal smoothing on the motion data loss information of historical image frames in the image frame sequence that are temporally preceding the previous image frame.

[0050] In an exemplary embodiment, before performing error constraint processing on the predicted motion data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame to obtain the target motion data corresponding to the target image frame, the method further includes:

[0051] Acquire a calibration image frame containing the calibration object, and object pose data corresponding to the calibration image frame; the pose measurement module on the calibration object is configured in the same way as the pose measurement module on the target object.

[0052] Action data prediction is performed on the calibrated object in the calibrated image frame to obtain predicted action data corresponding to the calibrated image frame;

[0053] Based on the predicted motion data corresponding to the calibration image frame, motion-driven processing is performed to obtain the predicted bone orientation data of the calibration object in the global coordinate system.

[0054] Based on the initial coordinate system transformation parameters, the object pose data corresponding to the calibration image frame is subjected to coordinate system transformation processing to obtain the calibration skeleton orientation data of the calibration object in the global coordinate system.

[0055] Based on the predicted bone orientation data of the calibration object in the global coordinate system and the calibration bone orientation data, calibration orientation error information is obtained;

[0056] The initial coordinate system transformation parameters are iteratively updated based on the calibration orientation error information to obtain the target coordinate system transformation parameters; the target coordinate system transformation parameters are used to perform coordinate system transformation processing on the object pose data corresponding to the target image frame.

[0057] According to a second aspect of the present disclosure, an object motion data processing apparatus is provided, comprising:

[0058] The target data acquisition unit is configured to acquire a sequence of image frames containing a target object, and object pose data corresponding to each image frame in the image frame sequence; the object pose data is obtained based on a pose measurement module set on the target object.

[0059] The first data prediction unit is configured to perform motion data prediction on the target object in the target image frame to obtain predicted motion data corresponding to the target image frame; the target image frame is an image frame in the image frame sequence for which motion data prediction has not been performed.

[0060] The error constraint unit is configured to perform error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame, so as to obtain the target action data corresponding to the target image frame.

[0061] In one exemplary embodiment, the error constraint unit includes:

[0062] The first driving unit is configured to perform action-driven processing based on the predicted action data to obtain predicted driving data; the predicted driving data is observation data corresponding to the predicted action data.

[0063] The first constraint processing unit is configured to perform error constraint processing on the prediction driving data based on the target image frame and the object pose data corresponding to the target image frame, to obtain target driving data.

[0064] The target action data determination unit is configured to determine the target action data based on the target driving data.

[0065] In one exemplary embodiment, the error constraint unit includes:

[0066] The first detection unit is configured to perform key point detection based on the target image frame to obtain the key point information of the target object.

[0067] The second driving unit is configured to perform action-driven processing based on the predicted action data to obtain the predicted key point information of the target object.

[0068] The second constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the error information between the detected key point information and the predicted key point information, and the object pose data corresponding to the target image frame, to obtain the target motion data.

[0069] In one exemplary embodiment, the detection key point information includes the detection confidence scores corresponding to each of the multiple target key points;

[0070] The second constraint processing unit includes:

[0071] The target confidence threshold determination unit is configured to perform a process based on the detection confidence of each of the plurality of target key points in the target image frame, the detection confidence of each of the plurality of target key points in historical image frames, and the smoothing confidence threshold corresponding to the historical image frame, to obtain a target confidence threshold corresponding to the target image frame; the historical image frame is an image frame in the image frame sequence that is temporally preceding the target image frame; the smoothing confidence threshold is determined based on the average confidence of the plurality of target key points in the historical image frames;

[0072] The first weight determination unit is configured to perform an operation based on a preset confidence threshold, a target confidence threshold corresponding to the target image frame, and the detection confidence of each of the plurality of target key points in the target image frame, to determine a first weight corresponding to each of the plurality of target key points in the target image frame; wherein the preset confidence threshold is less than the target confidence threshold corresponding to the target image frame.

[0073] The first weighting unit is configured to perform weighted processing on the error information of the detected key point information and the predicted key point information based on the first weights corresponding to the plurality of target key points, and generate first weighted error information.

[0074] The third constraint processing unit is configured to perform error constraint processing on the predicted action data based on the first weighted error information and the object pose data corresponding to the target image frame, so as to obtain the target action data.

[0075] In one exemplary embodiment, the error constraint unit includes:

[0076] The first coordinate system transformation unit is configured to transform the object pose data corresponding to the target image frame from the inertial coordinate system to the global coordinate system to obtain the target skeleton orientation data of the target object in the global coordinate system.

[0077] The third driving unit is configured to perform motion-driven processing based on the predicted motion data to obtain the predicted bone orientation data of the target object in the global coordinate system.

[0078] The first error information determination unit is configured to generate bone orientation error information based on the target bone orientation data and the predicted bone orientation data.

[0079] The fourth constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the target image frame and the bone orientation error information to obtain the target motion data.

[0080] In one exemplary embodiment, the fourth constraint processing unit includes:

[0081] The second detection unit is configured to perform key point detection based on the target image frame to obtain key point information of the target object; the key point information includes the detection confidence of each of the multiple target key points; the multiple skeletons of the target object in the target image frame correspond to the multiple target key points.

[0082] The second weight determination unit is configured to determine a second weight corresponding to any bone and a third weight corresponding to the target key point of any bone when the detection confidence of the target key point corresponding to any bone is less than a target confidence threshold; the second weight is greater than the third weight.

[0083] The second weighting unit is configured to perform weighted processing on the bone orientation error information based on the second weight corresponding to any one of the bones, and generate second weighted error information.

[0084] The third weighting unit is configured to perform weighted processing on the error information between the detection information and the prediction information of the target key point based on the third weight corresponding to the target key point of any skeleton, to obtain the third weighted error information; the prediction information of the target key point is obtained based on the action-driven processing of the predicted action data.

[0085] The fifth constraint processing unit is configured to perform error constraint processing on the predicted action data based on the second weighted error information and the third weighted error information to obtain the target action data.

[0086] In one exemplary embodiment, the predicted motion data includes object shape prediction data of the target object;

[0087] The error constraint unit includes:

[0088] The shape error information determination unit is configured to perform shape error information determination based on object shape prediction data of the target object in the target image frame and shape smoothing data corresponding to historical image frames; the historical image frames are image frames in the image frame sequence that are temporally preceding the target image frame; the shape smoothing data are obtained by temporally smoothing the object shape data of the target object in the historical image frames.

[0089] The sixth constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the shape error information, the target image frame, and the object pose data corresponding to the target image frame, to obtain the target motion data.

[0090] In one exemplary embodiment, the predicted motion data includes joint prediction data of the target object;

[0091] The error constraint unit includes:

[0092] The joint error information determination unit is configured to perform joint error information based on joint prediction data corresponding to the target image frame and joint constraint data corresponding to the previous image frame of the target image frame; the joint constraint data corresponding to the previous image frame is obtained by performing data error constraints on the joint prediction data of the target object in the previous image frame; the previous image frame is an image frame that is temporally preceding the target image frame in the image frame sequence and is adjacent to the target image frame;

[0093] The third weight determination unit is configured to determine the fourth weight based on the motion data loss information corresponding to the previous image frame; the motion data loss information corresponding to the previous image frame is determined based on the object pose data corresponding to the previous image frame and the constraint driving data corresponding to the previous image frame; the constraint driving data corresponding to the previous image frame is obtained by applying data error constraints to the predicted motion data corresponding to the previous image frame.

[0094] The fourth weighting unit is configured to perform weighting processing on the joint error information based on the fourth weight to generate fourth weighted error information;

[0095] The second iterative update unit is configured to perform iterative updates on the joint prediction data based on the fourth weighted error information to obtain the joint constraint data corresponding to the target image frame.

[0096] The seventh constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the target image frame, the object pose data corresponding to the target image frame, and the joint constraint data corresponding to the target image frame, to obtain the target motion data.

[0097] In one exemplary embodiment, the apparatus further includes:

[0098] The fourth driving unit is configured to execute constraint driving data corresponding to the target action data of the previous image frame;

[0099] The loss information unit is configured to perform an action data loss based on the object pose data corresponding to the previous image frame and the constraint driving data;

[0100] The third weight determination unit includes:

[0101] The weight calculation unit is configured to determine the fourth weight based on the motion data loss information corresponding to the previous image frame and the loss smoothing information corresponding to the previous image frame; the loss smoothing information corresponding to the previous image frame is obtained by temporally smoothing the motion data loss information of historical image frames in the image frame sequence that are temporally preceding the previous image frame.

[0102] In one exemplary embodiment, the apparatus further includes:

[0103] The calibration data acquisition unit is configured to acquire a calibration image frame containing the calibration object and the object pose data corresponding to the calibration image frame; the pose measurement module on the calibration object is configured in the same way as the pose measurement module on the target object.

[0104] The second data prediction unit is configured to perform motion data prediction on the calibration object in the calibration image frame to obtain predicted motion data corresponding to the calibration image frame.

[0105] The fifth driving unit is configured to perform motion-driven processing based on the predicted motion data corresponding to the calibration image frame to obtain the predicted bone orientation data of the calibration object in the global coordinate system.

[0106] The second coordinate system transformation unit is configured to perform coordinate system transformation processing on the object pose data corresponding to the calibration image frame based on the initial coordinate system transformation parameters, so as to obtain the calibration bone orientation data of the calibration object in the global coordinate system.

[0107] The second error information determination unit is configured to execute the predicted bone orientation data and the calibrated bone orientation data based on the calibrated object in the global coordinate system to obtain calibration orientation error information.

[0108] The third iterative update unit is configured to perform iterative updates on the initial coordinate system transformation parameters based on the calibration orientation error information to obtain the target coordinate system transformation parameters; the target coordinate system transformation parameters are used to perform coordinate system transformation processing on the object pose data corresponding to the target image frame.

[0109] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the object action data processing method as described above.

[0110] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of a server, enables the server to perform the object action data processing method as described above.

[0111] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform the above-described object action data processing method.

[0112] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0113] In processing object motion data, this disclosure can process the object motion data of the target object based on a target image frame containing the target object and the object pose data corresponding to the target image frame. This supplements the image data corresponding to the target image frame with the object pose data, compensating for the visual occlusion problem in monocular vision, achieving comprehensive acquisition of the target object's motion data, and improving the accuracy of motion representation of the target object. Furthermore, when predicting the motion data of the target object based on the target image frame, the predicted motion data can be jointly constrained based on the image information in the target image frame and the object pose data corresponding to the target image frame. This reduces the error between the image information in the target image frame and the object pose data corresponding to the target image frame and the constrained predicted motion data. The target motion data is determined based on the predicted motion data after constraint processing, which can improve the accuracy of the predicted motion data determination of the target object. In turn, the accuracy of motion driving can be improved based on the target motion data.

[0114] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0115] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0116] Figure 1 This is a schematic diagram of an implementation environment according to an exemplary embodiment;

[0117] Figure 2 This is a flowchart illustrating an object action data processing method according to an exemplary embodiment;

[0118] Figure 3 This is a flowchart illustrating an action-driven error constraint processing method according to an exemplary embodiment;

[0119] Figure 4 This is a flowchart illustrating an error constraint method based on key point detection according to an exemplary embodiment;

[0120] Figure 5 This is a flowchart illustrating an error constraint method based on key point weights according to an exemplary embodiment;

[0121] Figure 6 This is a flowchart illustrating a method for constraint processing based on bone orientation data according to an exemplary embodiment;

[0122] Figure 7This is a flowchart illustrating a method for constraint processing based on bone weights according to an exemplary embodiment;

[0123] Figure 8 This is a flowchart illustrating a data constraint method based on object shape according to an exemplary embodiment;

[0124] Figure 9 This is a flowchart illustrating a method for error constraint based on loss information of historical image frames, according to an exemplary embodiment.

[0125] Figure 10 This is a flowchart illustrating a method for calculating the fourth weight according to an exemplary embodiment;

[0126] Figure 11 This is a flowchart illustrating a coordinate system transformation parameter optimization method according to an exemplary embodiment;

[0127] Figure 12 This is a schematic diagram illustrating the optimization process of coordinate system transformation parameters according to an exemplary embodiment;

[0128] Figure 13 This is a schematic diagram illustrating the constraint process of predicting action data according to an exemplary embodiment;

[0129] Figure 14 This is a schematic diagram of an object motion data processing device according to an exemplary embodiment;

[0130] Figure 15 This is a block diagram illustrating an electronic device for processing object motion data according to an exemplary embodiment;

[0131] Figure 16 This is a block diagram illustrating another electronic device for processing object motion data according to an exemplary embodiment. Detailed Implementation

[0132] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0133] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0135] Please see Figure 1 The diagram illustrates an implementation environment provided in this embodiment, which may include a data acquisition terminal 110 and a motion data processing terminal 120; the data acquisition terminal 110 and the motion data processing terminal can communicate via a network.

[0136] Specifically, the data acquisition terminal 110 can acquire object data from the target object, including object image data, object posture data, etc. The data acquisition terminal 110 can send the acquired object data to the motion data processing terminal 120. The motion data processing terminal 120 can predict the motion data of the target object based on the received object image data to obtain predicted motion data. Then, it can perform error constraint processing on the predicted motion data based on the object image data and object posture data to obtain the target motion data corresponding to the target object.

[0137] The data acquisition terminal 110 can communicate with the motion data processing terminal 120 based on browser / server (B / S) mode or client / server (C / S) mode. The data acquisition terminal 110 may include: image acquisition device, inertial measurement unit (IMU), smart wearable device, etc.

[0138] The motion data processing terminal 120 and the data acquisition terminal 110 can establish a communication connection via wired or wireless means. The operating system running on the motion data processing terminal 120 can include, but is not limited to, Android, iOS, Linux, Windows, etc. Specifically, the motion data processing terminal 120 can be a physical device such as a smartphone, tablet, laptop, digital assistant, smart wearable device, vehicle terminal, server, etc. When the motion data processing terminal 120 is a server, it can include a standalone server, a distributed server, or a server cluster composed of multiple servers, wherein the server can be a cloud server.

[0139] To address the problem of inaccurate object motion data prediction caused by visual occlusion in monocular vision in related technologies, this disclosure provides an object motion data processing method. The execution entity of this method can be the aforementioned motion data processing terminal; please refer to [link to details]. Figure 2 The method may include:

[0140] S210. Obtain an image frame sequence containing the target object, and object pose data corresponding to each image frame in the image frame sequence; the object pose data is obtained based on a pose measurement module set on the target object.

[0141] In this embodiment, the target object may include objects with changing motion, such as humans and animals. Images of the target object can be acquired using an image acquisition device, resulting in a sequence of image frames containing the target object. The object's pose data can be obtained through a pose measurement module installed on the target object. Specifically, the pose measurement module can be an inertial measurement unit (IMU). Specifically, IMUs can be installed at various skeletal structures of the target object to obtain the corresponding pose data. An IMU is a device that measures the three-axis attitude angles (or angular rates) and acceleration of an object. Gyroscopes and accelerometers are core components of an inertial navigation system. With the help of built-in accelerometers and gyroscopes, the IMU can measure linear acceleration and rotational angular rates from three directions. By calculating the measured linear acceleration and rotational angular rates, information such as the object's attitude, velocity, and displacement can be obtained.

[0142] Since the image frame sequence containing the target object and the object pose data of the target object are obtained through different data acquisition devices, it is necessary to perform time alignment on the data obtained from different data acquisition devices in order to determine the image information of the target object at each time node, as well as the object pose data at the corresponding time node. Based on the image information and object pose data of the target object at each time node, it is possible to reflect the specific action performance information of the target object at each time node from multiple perspectives.

[0143] S220. Perform motion data prediction on the target object in the target image frame to obtain predicted motion data corresponding to the target image frame; the target image frame is an image frame in the image frame sequence that has not undergone motion data prediction.

[0144] In this embodiment, motion data prediction processing can be performed on each image frame sequentially based on the temporal relationship between each image frame in the image frame sequence. The target image frame can be the image frame in the image frame sequence that needs to be constrained for motion prediction.

[0145] When a target image frame contains a target object, deep learning methods can be used to perform image recognition on the target image frame in order to obtain the predicted action data of the target object. Specifically, feature extraction can be performed on the target image frame to obtain the predicted action data of the target object in the target image frame.

[0146] In this embodiment, the motion data for the target object may include shape parameters, pose parameters, etc. The shape parameter represents 10 parameters related to the human body's height, weight, head-to-body ratio, etc., and is a 10-dimensional vector. The shape can be controlled by 10 incremental templates. The pose parameter represents 3 degrees of freedom of global displacement parameters and 72 relative rotation parameters, for a total of 75 parameters. Correspondingly, the predicted motion data for the target image frame may include predicted shape parameters and predicted pose parameters for the target object.

[0147] S230. Based on the target image frame and the object pose data corresponding to the target image frame, perform error constraint processing on the predicted action data corresponding to the target image frame to obtain target action data corresponding to the target image frame.

[0148] Action data constraints on target image frames can refer to predicting the action data of the target object in the target image frame to obtain predicted action data corresponding to the target image frame; and then, based on the image information of the target image frame and the pose data of the object corresponding to the target image frame, applying error constraints to the predicted action data corresponding to the target image frame. Each image frame in the image frame sequence needs to undergo action data constraint processing to make the predicted action data corresponding to each image frame closer to the real action data. This results in more realistic and natural actions of the driven object when using the constrained predicted action data for action driving.

[0149] As described above, image information from the target image frame and the corresponding object pose data can reflect the target object's motion performance information from multiple perspectives, improving the comprehensiveness and accuracy of the motion performance information. However, in the case of predicting motion data for the target object based on the target image frame, two problems arise. First, object occlusion or self-occlusion can lead to inaccurate acquisition of the target image frame, resulting in predicted motion data that may differ from the actual motion data. Second, the predicted motion data is based solely on the target image frame and does not incorporate the corresponding object pose data, which is unaffected by visual occlusion, further contributing to the discrepancy between the predicted and actual motion data. Therefore, to address these issues, error constraint processing can be applied to the predicted motion data corresponding to the target image frame, using the actual acquired object data such as the target image frame and its corresponding object pose data. This optimization of the predicted motion data yields target motion data that closely approximates the actual motion data.

[0150] In processing object motion data, this disclosure can process the object motion data of the target object based on a target image frame containing the target object and the object pose data corresponding to the target image frame. This supplements the image data corresponding to the target image frame with the object pose data, compensating for the visual occlusion problem in monocular vision, achieving comprehensive acquisition of the target object's motion data, and improving the accuracy of motion representation of the target object. Furthermore, when predicting the motion data of the target object based on the target image frame, the predicted motion data can be jointly constrained based on the image information in the target image frame and the object pose data corresponding to the target image frame. This reduces the error between the image information in the target image frame and the object pose data corresponding to the target image frame and the constrained motion data. The target motion data is determined based on the motion data after constraint processing, which can improve the accuracy of determining the target object's motion data. In turn, motion driving based on the target motion data can improve the accuracy of motion driving.

[0151] In this embodiment, the target motion data mentioned above is generally data that cannot be directly observed. The target motion data can be used as input parameters for a motion-driven model. By inputting the target motion data into the motion-driven model, directly observable data corresponding to the target motion data can be obtained, i.e., motion representation data corresponding to the target motion data. The motion-driven model can be an SMPL model (Skinned Multi-Person Linear Model). The input parameters of the SMPL model can include the aforementioned shape data, pose data, etc., and the output parameters are the motion representation data corresponding to the shape data, pose data, etc.

[0152] Accordingly, please refer to Figure 3 It illustrates an action-driven error constraint processing method, which may include:

[0153] S310. Perform action-driven processing based on the predicted action data to obtain predicted driving data; the predicted driving data is the observation data corresponding to the predicted action data.

[0154] S320. Based on the target image frame and the object pose data corresponding to the target image frame, error constraint processing is performed on the prediction driving data to obtain target driving data.

[0155] S330. Determine the target action data based on the target driving data.

[0156] In this embodiment, the image frame data and object pose data contained in the target image frame can both be observable data, while the predicted action data is unobservable data. The predicted action data can be processed based on an action-driven model to obtain the predicted action data, which is the action representation data corresponding to the predicted action data. Furthermore, the target image frame and the object pose data corresponding to it are both observable action representation data directly acquired by an information acquisition device. Therefore, the predicted action data, the target image frame, and the object pose data corresponding to it are all action representation data of the target object. Thus, error constraint processing can be performed on the predicted action data using the target image frame and the object pose data corresponding to it, minimizing the error between the target image frame, the object pose data corresponding to it, and the predicted action data to obtain the corresponding target action data. After obtaining the target action data, error constraint processing can be performed on the predicted action data in reverse to obtain the corresponding target action data.

[0157] Specifically, in the process of error constraint processing of the prediction-driven data, the following steps can be repeatedly executed: perform action-driven processing based on the current prediction action data to obtain the current prediction-driven data; perform error constraint processing on the current prediction-driven data based on the target image frame and the object pose data corresponding to the target image frame to obtain the current constraint-driven data; update the current prediction action data in reverse based on the current constraint-driven data to obtain the current constraint-driven data; determine the current constraint-driven data as the current prediction action data, until the preset number of iterations is met, or until the error information between the target image frame, the object pose data corresponding to the target image frame, and the current prediction-driven data meets the preset error condition.

[0158] Therefore, in the process of constraining the predicted motion data, the predicted motion data that cannot be directly observed can be driven into observable predicted driving data, which facilitates error comparison and calculation with the target video frame and the object pose data corresponding to the target video frame. By constraining the predicted driving data, the error constraint of the predicted motion data can be realized accordingly, further improving the convenience and operability of constraining the error of the predicted motion data.

[0159] When constraining the predicted motion data based on the target image frame and the corresponding object pose data, corresponding observable motion representation data can be further determined based on the target image frame and the corresponding object pose data, such as keypoint information, which can represent the positional information of each bone of the target object; please refer to [link to relevant documentation] for details. Figure 4 It illustrates an error constraint method based on key point detection, which may include:

[0160] S410. Perform key point detection based on the target image frame to obtain the key point information of the target object.

[0161] S420. Based on the predicted action data, perform action-driven processing to obtain the predicted key point information of the target object.

[0162] S430. Based on the error information between the detected key point information and the predicted key point information, and the object pose data corresponding to the target image frame, error constraint processing is performed on the predicted action data to obtain the target action data.

[0163] By inputting the target image frame into the keypoint recognition model, the corresponding keypoint information of the target object can be obtained. After performing action-driven processing on the predicted action data, the resulting predicted driving data may include the predicted keypoint information corresponding to the target image frame. Therefore, error constraint processing can be performed on the predicted keypoint information through the detected keypoint information to achieve error constraint processing on the predicted action data.

[0164] Specifically, the key point information to be detected can be two-dimensional key point information. After the action-driven processing of the predicted action data, the predicted key point information is obtained in three dimensions. Therefore, it is necessary to perform dimensional mapping on the three-dimensional predicted key point information to obtain two-dimensional predicted key point information, so as to facilitate error constraint processing between the two-dimensional detected key point information and the two-dimensional predicted key point information.

[0165] In this embodiment, by performing key point detection on the target image frame, the key point information of the target object can be obtained. The key point information can characterize the actual feature information of each key part of the target object. By using the key point information as a reference and calculating the error with the predicted key point information, the key point error information can be determined. The existence of error in the key point information indicates that there is an error between the corresponding predicted action data and the actual action data. By performing error constraint processing on the predicted action data based on the key point error information, the accuracy of the target action data determination can be improved.

[0166] In one specific embodiment, the key point detection information includes the detection confidence scores corresponding to multiple target key points, and the corresponding weights can be determined based on the detection confidence scores corresponding to each of the multiple target key points; please refer to the relevant documentation. Figure 5 It illustrates an error constraint method based on keypoint weights, which may include:

[0167] S510. Based on the detection confidence scores corresponding to the plurality of target key points in the target image frame, the detection confidence scores corresponding to the plurality of target key points in historical image frames, and the smoothing confidence threshold corresponding to the historical image frame, a target confidence threshold corresponding to the target image frame is obtained; the historical image frame is an image frame in the image frame sequence that is temporally preceding the target image frame; the smoothing confidence threshold is determined based on the average confidence scores of the plurality of target key points in the historical image frame.

[0168] S520. Based on a preset confidence threshold, a target confidence threshold corresponding to the target image frame, and the detection confidence corresponding to each of the plurality of target key points in the target image frame, determine the first weight corresponding to each of the plurality of target key points in the target image frame; the preset confidence threshold is less than the target confidence threshold corresponding to the target image frame.

[0169] S530. Based on the first weights corresponding to the multiple target key points, the error information of the detected key point information and the predicted key point information is weighted and processed to generate first weighted error information.

[0170] S540. Based on the first weighted error information and the object pose data corresponding to the target image frame, the predicted action data is subjected to error constraint processing to obtain the target action data.

[0171] For each image frame in the image frame sequence, the average confidence level of each image frame can be determined based on the detection confidence levels of multiple target keypoints in each image frame. For each image frame, the smoothed confidence level of the historical image frames preceding it can be calculated based on the average confidence level of the historical image frames preceding it. For each image frame, the smoothed confidence threshold of the historical image frames preceding it can be calculated based on the confidence threshold of the historical image frames preceding it. The target confidence threshold of the target image frame can be calculated based on the average confidence level of the target image frame, the smoothed confidence level of the historical image frames preceding it, and the smoothed confidence threshold of the historical image frames preceding it.

[0172] Furthermore, this embodiment can also set a preset confidence threshold, which can be less than the target confidence threshold. The preset confidence threshold is mainly used to filter out target key points with low detection confidence. For example, the preset confidence can be 0.2, 0.25, etc. The first weight corresponding to the target key points with a detection confidence less than the preset confidence threshold is 0.

[0173] The smoothed data in this embodiment can be achieved through methods such as exponential moving average or weighted moving average, and is not limited here.

[0174] In this embodiment, the correspondence between weights and target confidence thresholds can be preset. Once the preset confidence threshold and target confidence threshold of the target image frame are determined, the detection confidence of each target key point can be compared with the preset confidence threshold or target confidence threshold to determine the first weight corresponding to each target key point.

[0175] In one example, both the detected keypoint information and the predicted keypoint information can include the location information of each target keypoint. Therefore, based on the detected and predicted location information corresponding to each target keypoint, error information corresponding to each target keypoint can be obtained. Then, the error information corresponding to each target keypoint is weighted according to a first weight, resulting in first weighted error information corresponding to that target keypoint. Optimizing this first weighted error information to minimize it or to meet preset optimization error conditions allows for the acquisition of corresponding target action data.

[0176] The first weighted error information corresponding to the target image frame can be calculated using equation (1):

[0177]

[0178] Among them, E proj w represents the first weighted error information corresponding to the target image frame. f The first weight corresponding to the target key point, To predict key point information, J 2d,det To detect key information.

[0179] For w f The calculation method can be obtained through equation (2):

[0180]

[0181] Where τ1 is the preset confidence threshold, τ2 is the target confidence threshold, and c is the detection confidence corresponding to the target key point.

[0182] The calculation method for τ2 can be obtained from equation (3):

[0183]

[0184] in, The average confidence level of the target image frame. τ represents the smooth confidence level corresponding to the historical image frames preceding the target image frame. 2,mv The smooth confidence threshold is the historical image frames preceding the target image frame. and τ 2,mv It can be obtained through exponential moving average or weighted moving average.

[0185] Therefore, when calculating the confidence threshold of the target image frame, it is based on the smoothed confidence threshold and the average confidence threshold of historical image frames, thus linking the determination of the confidence threshold corresponding to the target image frame with the confidence information corresponding to historical image frames. Further, the weights of each target key point are determined based on the confidence threshold, which distinguishes the importance of different target key points, increasing the weight of effective information and decreasing the weight of invalid information, thereby accurately obtaining the first weighted error information. Error constraint processing is then performed based on the first weighted error information to optimize the predicted action data, improving the matching between target action data and actual action data, and reducing the error between them. This allows for adaptive adjustment of the average confidence threshold based on historical confidence levels. For example, the detection confidence of complex actions is generally low, so the average confidence threshold can be lowered accordingly to retain key information and avoid the problem of the confidence of complex actions being completely reset to zero when using a fixed confidence threshold.

[0186] In an optional embodiment, the predicted action data is driven, and the resulting predicted driving data may further include the predicted skeleton orientation data of the target object in the global coordinate system; based on the object pose data corresponding to the target image frame, the target skeleton orientation data of the target object in the global coordinate system can be obtained, thereby allowing constraint processing of the predicted action data based on the skeleton orientation data; please refer to [link / reference] for details. Figure 6 It illustrates a method for constraint processing based on bone orientation data, which may include:

[0187] S610. Transform the object pose data corresponding to the target image frame from the inertial coordinate system to the global coordinate system to obtain the target skeleton orientation data of the target object in the global coordinate system.

[0188] S620. Based on the predicted motion data, perform motion-driven processing to obtain the predicted bone orientation data of the target object in the global coordinate system.

[0189] S630. Based on the target bone orientation data and the predicted bone orientation data, generate bone orientation error information.

[0190] S640. Based on the target image frame and the bone orientation error information, perform error constraint processing on the predicted motion data to obtain the target motion data.

[0191] Based on the above content of this embodiment, the object pose data may include the acceleration and orientation of the inertial measurement devices, and then the orientation data of the target object's skeleton can be calculated. Taking the target object as a human body as an example, this embodiment can introduce 5 inertial measurement devices, which are respectively set at the spine, left and right forearms, and left and right lower legs of the target object, thereby obtaining sparse object pose data.

[0192] To facilitate comparison, the acquired object pose data and the bone orientation data obtained from motion-driven processing can both be transformed to the global coordinate system for comparison. Specifically, the object pose data corresponding to the target image frame can be transformed to obtain the target bone orientation data of the target object in the global coordinate system. Then, based on the target bone orientation data and the predicted bone orientation data, bone orientation error information can be generated.

[0193] The pose data of an object collected by an inertial measurement unit can be used to obtain the corresponding skeletal orientation data. The skeletal orientation data can represent the action of the target object from the perspective of skeletal motion. By using the skeletal orientation error information between the target skeletal orientation data and the predicted skeletal orientation data, error constraint processing can be performed on the predicted action data. This can achieve constraint on the predicted action data from the perspective of skeletal motion, further improving the effectiveness and accuracy of the constraint on the predicted action data.

[0194] In this embodiment, there is a corresponding relationship between the keypoint information obtained from keypoint detection of the target object and the joints of each bone of the target object. For example, each bone has two ends, each end can correspond to a joint, and each joint can correspond to at least one keypoint. Detecting keypoints can be understood as detecting joints. Since the detection confidence of each keypoint is different, that is, the detection confidence of each joint may be different, it is necessary to determine the joint confidence based on the joint detection confidence, and then determine the weight of each bone. Please refer to [link / reference] for details. Figure 7 It illustrates a method for constraint processing based on bone weights, which may include:

[0195] S710. Based on the target image frame, perform key point detection to obtain the key point information of the target object; the key point information includes the detection confidence of each of the multiple target key points; the multiple skeletons of the target object in the target image frame correspond to the multiple target key points.

[0196] S720. If the detection confidence of the target key point corresponding to any bone is less than the target confidence threshold, determine the second weight corresponding to the bone and the third weight corresponding to the target key point corresponding to the bone; the second weight is greater than the third weight.

[0197] S730. Based on the second weight corresponding to any one of the bones, the bone orientation error information is weighted to generate second weighted error information.

[0198] S740. Based on the third weight corresponding to the target key point of any skeleton, the error information between the detection information of the target key point and the prediction information of the target key point is weighted to obtain the third weighted error information; the prediction information of the target key point is obtained based on the action-driven processing of the prediction action data.

[0199] S750. Based on the second weighted error information and the third weighted error information, the predicted action data is subjected to error constraint processing to obtain the target action data.

[0200] The keypoint detection method here can be the same as the one described above. Figure 4 as well as Figure 5 The methods shown are the same, so they will not be repeated here.

[0201] For keypoint detection of the target image frame, if the detection confidence of the target keypoint corresponding to any bone is less than the target confidence threshold, refer to equation (2). This corresponds to the two ranges c<τ1 and τ1≤c≤τ2, indicating that keypoints within these two ranges may have inaccurate detection. For inaccurately detected keypoints, the dependence on keypoint detection information can be reduced, while the dependence on the orientation data of the bone corresponding to the inaccurately detected keypoint can be increased, so that the dependence on keypoint detection information is less than the dependence on the orientation data of the bone corresponding to the inaccurately detected keypoint. That is, if the detection confidence of the target keypoint corresponding to any bone is less than the target confidence threshold, the weight of the bone orientation error information corresponding to that bone is increased, and the weight of the keypoint detection error information corresponding to that bone is decreased. That is, the second weight is increased and the third weight is decreased, so that the data constraints based on that bone are more dependent on the bone orientation data and the dependence on keypoint detection information is reduced.

[0202] Furthermore, bone orientation constraints can be achieved through equation (4):

[0203]

[0204] Where E ori For the second weighted error information, w i Here, I represents the second weight corresponding to the skeleton, and R represents the identity matrix. rel We can obtain it through equation (5):

[0205]

[0206] in, These include the transformation between the IMU inertial coordinate system and the global coordinate system, the bone orientation data of the IMU data, the transformation between the IMU local coordinate system and the corresponding bone local coordinate system, and the predicted bone orientation data in the global coordinate system. This refers to the target skeleton orientation data in the global coordinate system.

[0207] Based on the key point detection results, the confidence information of each target detection point is determined, thereby determining the degree of dependence on the image information of the target image frame and the degree of dependence on the object pose data when performing data constraint processing. Furthermore, when the detection points are inaccurate when performing key point detection based on the image information of the target image frame, the dependence on the corresponding object pose data is increased, thereby improving the accuracy of error information determination, further improving the optimization results of target action data, and improving the optimization efficiency of target action data.

[0208] The predicted action data corresponding to the target image frame may include the predicted shape data of the target object. Correspondingly, the predicted object shape data corresponding to the target image frame can be compared with the object shapes of historical image frames that have already undergone error constraint processing to constrain the predicted action data from the shape dimension of the target object; please refer to [link to details]. Figure 8 It illustrates a data constraint method based on object shape, which may include:

[0209] S810. Based on the object shape prediction data of the target object in the target image frame and the shape smoothing data corresponding to the historical image frame, determine the shape error information; the historical image frame is the image frame in the image frame sequence that is temporally preceding the target image frame; the shape smoothing data is obtained by temporally smoothing the object shape data of the target object in the historical image frame.

[0210] S820. Based on the shape error information, the target image frame, and the object pose data corresponding to the target image frame, error constraint processing is performed on the predicted motion data to obtain the target motion data.

[0211] The object shape data may include shape data and bone length data. Correspondingly, the target object's shape prediction data includes the target object's shape prediction data and the length prediction data of multiple bones of the target object. Further, when determining shape error information, it may specifically include:

[0212] Based on the predicted shape data of the target object in the target image frame and the shape smoothing data corresponding to the historical image frames, shape error information is determined; the historical image frames are image frames in the image frame sequence that are temporally preceding the target image frame; the shape smoothing data is obtained by temporally smoothing the shape data of the target object in the historical image frames.

[0213] Based on the predicted length data of multiple bones of the target object in the target image frame and the length smoothing data corresponding to the historical image frame, bone length error information is determined; the length smoothing data is obtained by performing temporal smoothing on the length data of multiple bones of the target object in the historical image frame.

[0214] Shape error information is obtained based on shape error information and bone length error information.

[0215] The determination of shape error information can be based on equation (6):

[0216]

[0217] Among them, E shape This is information on shape error. The shape prediction data corresponding to the target image frame. mv This is the shape smoothing data corresponding to the historical image frames.

[0218] The determination of bone length error information can be based on equation (7):

[0219]

[0220] Among them, E bone This is information on bone length error. For the bone length prediction data corresponding to the target image frame, bone mv This is the smoothed bone length data corresponding to the historical image frames.

[0221] Using exponential moving average as an example to illustrate shape mv and bone mv For details of the update process, please refer to Figures (8) and (9):

[0222]

[0223]

[0224] Where, α s As preset parameters, shape mv1 For the shape smoothing data corresponding to the target image frame, bone mv1 This is the smoothed data of the bone length corresponding to the target image frame.

[0225] By applying temporal constraints to the predicted shape data of the target image frame using the smoothed shape data corresponding to historical image frames, the temporal consistency of the target object's shape can be ensured, thereby improving the temporal stability of the target object's shape.

[0226] In this embodiment, the prediction driving data obtained by data-driven analysis of the predicted motion data may also include joint prediction data of the target object; after data constraint processing of each image frame, there will generally be a corresponding error between the target motion data and the actual motion data, and this error information can be determined as loss information; accordingly, when performing error constraint processing on the current target image frame, the dependence of the current target image frame on the historical image frames can be determined based on the loss information corresponding to the historical image frames; please refer to [link / reference] for details. Figure 9 It illustrates a method for error constraint based on loss information from historical image frames, which may include:

[0227] S910. Based on the joint prediction data corresponding to the target image frame and the joint constraint data corresponding to the previous image frame of the target image frame, joint error information is obtained; the joint constraint data corresponding to the previous image frame is obtained by performing data error constraint on the joint prediction data of the target object in the previous image frame; the previous image frame is an image frame that is temporally located before the target image frame in the image frame sequence and is adjacent to the target image frame.

[0228] S920. Determine the fourth weight based on the motion data loss information corresponding to the previous image frame; the motion data loss information corresponding to the previous image frame is determined based on the object pose data corresponding to the previous image frame and the constraint driving data corresponding to the previous image frame; the constraint driving data corresponding to the previous image frame is obtained by performing data error constraints on the predicted motion data corresponding to the previous image frame.

[0229] S930. The joint error information is weighted based on the fourth weight to generate fourth weighted error information.

[0230] S940. The joint prediction data is iteratively updated based on the fourth weighted error information to obtain the joint constraint data corresponding to the target image frame.

[0231] S950. Based on the target image frame, the object pose data corresponding to the target image frame, and the joint constraint data corresponding to the target image frame, error constraint processing is performed on the predicted motion data to obtain the target motion data.

[0232] The loss between the target motion data and the actual motion data can be characterized by the loss between the object pose data and the constraint-driven data after error constraint processing. This loss can be considered motion data loss information. The existence of an error between the object pose data and the constraint-driven data indicates an error between the target motion data and the actual motion data. Therefore, motion data loss information can characterize the error information between the target motion data and the actual motion data. For ease of calculation, the motion data loss information in this embodiment can be obtained by calculating the error based on the object pose data and the constraint-driven data after error constraint processing. If the motion data loss information of the previous image frame is greater than or equal to the average loss information, then when calculating the error information corresponding to the joint prediction data and the joint constraint data, the dependence on the previous image frame is reduced, i.e., the fourth weight is decreased. Conversely, if the motion data loss information of the previous image frame is less than the average loss information, then when calculating the error information corresponding to the joint prediction data and the joint constraint data, the dependence on the previous image frame is increased, i.e., the fourth weight is increased.

[0233] When performing error constraint processing on the joint prediction data corresponding to the target image frame, the joint constraint data corresponding to the target image frame can be used as a basis for error constraint processing, so that the joint data of the target object remains stable between adjacent image frames, avoiding the problem of sudden changes in joint data. In addition, the corresponding fourth weight is determined based on the motion data loss information corresponding to the previous image frame, which can determine the corresponding degree of dependence based on the actual optimization situation of the previous image frame. This can effectively break away from the trend of poor prediction results, and maintain the temporal stability of motion capture results when the optimization results were relatively accurate in the past.

[0234] For details on the calculation method of the fourth weight, please refer to [link / reference needed]. Figure 10 The method may include:

[0235] S1010. Determine the constraint driving data corresponding to the target motion data of the previous image frame.

[0236] S1020. Based on the object pose data corresponding to the previous image frame and the constraint driving data, determine the motion data loss information corresponding to the previous image frame.

[0237] S1030. Based on the motion data loss information corresponding to the previous image frame and the loss smoothing information corresponding to the previous image frame, determine the fourth weight; the loss smoothing information corresponding to the previous image frame is obtained by performing temporal smoothing on the motion data loss information of historical image frames in the image frame sequence that are temporally preceding the previous image frame.

[0238] Furthermore, the fourth weight can be determined based on the ratio of motion data loss information of the previous image frame to motion data loss information of historical image frames, wherein the motion data loss information of historical image frames can be obtained by temporal smoothing based on the motion data loss information of each historical image frame before the previous image frame.

[0239] The joint prediction data may further include joint pose prediction data and joint position prediction data. The joint pose prediction data may be one piece of data in the predicted motion data corresponding to the target image frame, and the joint position prediction data may be data obtained by motion driving based on the predicted motion data.

[0240] The error information corresponding to the joint position data can be obtained based on equation (10):

[0241]

[0242] Among them, E pos,tc For error information corresponding to joint position data, w tAs the fourth weight, J is the three-dimensional joint position prediction data obtained through motion-driven processing. 3d,last This is the joint position constraint data corresponding to the previous image frame.

[0243] The error information corresponding to the joint posture data can be obtained based on equation (11):

[0244]

[0245] Among them, E ang,tc This refers to the error information corresponding to the joint posture data. For the joint pose prediction data corresponding to the target image frame, pose last This is the joint pose constraint data corresponding to the previous image frame.

[0246] The error information corresponding to the joint prediction data is as follows:

[0247] E tc =E pos,tc +E ang,tc (12)

[0248] Furthermore, the fourth weight w t The following conditions must be met:

[0249] w t ∝L imu,last / L imu,mv -1 (13)

[0250] Among them, L imu,last L represents the motion data loss information corresponding to the previous image frame. imu,mv This is the loss smoothing information corresponding to the previous image frame.

[0251] Therefore, when determining the fourth weight, it is based on the motion data loss information and loss smoothing information corresponding to the previous image frame. The loss smoothing information can smoothly represent the average loss level of historical image frames, thereby improving the adaptability of the fourth weight to the image frame sequence. Determining the corresponding fourth weight based on the motion data loss information corresponding to the previous image frame can determine the corresponding dependence based on the actual optimization situation of the previous image frame. This can effectively escape the trend of poor prediction results, and maintain the temporal stability of motion capture results when the optimization results were relatively accurate in the past.

[0252] In this embodiment, when performing error constraint processing on predicted action data based on object pose data, the object pose data can be transformed into data that is easier to calculate. For example, through the coordinate system transformation process described above, the object pose data can be converted into target bone orientation data in the global coordinate system. Then, the target bone orientation data is compared with the predicted bone orientation data obtained after action-driven processing, thereby achieving error constraint processing on the predicted action data. Coordinate system transformation parameters are required during coordinate system transformation and bone orientation loss calculation. To facilitate the subsequent direct use of coordinate system transformation parameters to perform coordinate system transformation on the object pose data, the coordinate system transformation parameters can be iteratively optimized in advance to obtain optimized coordinate system transformation parameters. Please refer to [link / reference] for details. Figure 11 It illustrates a method for optimizing coordinate system transformation parameters, which may include:

[0253] S1110. Obtain a calibration image frame containing the calibration object, and object pose data corresponding to the calibration image frame; the pose measurement module on the calibration object is configured in the same way as the pose measurement module on the target object.

[0254] S1120. Perform motion data prediction on the calibration object in the calibration image frame to obtain predicted motion data corresponding to the calibration image frame.

[0255] It should be noted that when predicting the action of the calibration object in the calibration image frame, the prediction is performed directly based on the image data in the calibration image frame. The predicted action data corresponding to the calibration image frame is not combined with the object pose data of the calibration object because the target coordinate system transformation parameters for transforming the object pose data have not yet been determined. In this embodiment, the data constraint is applied to the object pose data of the calibration object based on the predicted action data corresponding to the calibration image frame and the initial coordinate system transformation parameters, so as to optimize the initial coordinate system transformation parameters.

[0256] S1130. Perform motion-driven processing based on the predicted motion data corresponding to the calibration image frame to obtain the predicted bone orientation data of the calibration object in the global coordinate system.

[0257] S1140. Based on the initial coordinate system transformation parameters, perform coordinate system transformation processing on the object pose data corresponding to the calibration image frame to obtain the calibration skeleton orientation data of the calibration object in the global coordinate system.

[0258] S1150. Based on the predicted bone orientation data of the calibration object in the global coordinate system and the calibration bone orientation data, obtain the calibration orientation error information.

[0259] S1160. The initial coordinate system transformation parameters are iteratively updated based on the calibration orientation error information to obtain the target coordinate system transformation parameters; the target coordinate system transformation parameters are used to perform coordinate system transformation processing on the object pose data corresponding to the target image frame.

[0260] When performing error constraint processing on the predicted action data, the object pose data corresponding to the target image frame can be transformed based on the target coordinate system transformation parameters to obtain the target skeleton orientation data of the target object in the global coordinate system; based on the target image frame and the target skeleton orientation data of the target object in the global coordinate system, the predicted action data corresponding to the target image frame is subjected to error constraint processing to obtain the target action data corresponding to the target image frame.

[0261] The calibration image frame can be an image frame containing a silent pose, which can be a simple pose such as Tpose or Apose. By selecting an image frame containing a silent pose, the complexity of data processing can be simplified and the efficiency of data processing can be improved.

[0262] The coordinate system transformation parameters to be optimized in this embodiment include R. ig ,R ib , where R ig R represents the transformation parameters between the IMU inertial coordinate system and the global coordinate system. ib These are the transformation parameters between the IMU's local coordinate system and the corresponding skeleton's local coordinate system; for the steps to solve for the coordinate system transformation parameters, please refer to [link to documentation]. Figure 12 ,include:

[0263] 1. Calibrate the internal and external parameters of the camera system.

[0264] 2. Based on the correspondence between the peak acceleration of the IMU and the landing of the foot in the video, time synchronization between the camera system and the IMU system is performed to obtain IMU data corresponding to each frame of the image.

[0265] 3. Detect key points in each frame and smooth them using Kalman filtering.

[0266] 4. Use the monocular SMPL prediction method to estimate the initial SMPL parameters.

[0267] 5. Based on the results of (4), select several T / A attitude frames with better prediction quality for calibration.

[0268] 6. Using the SMPL prediction results, IMU data, and keypoint detection results of T / A frames as input, optimize the solution for R. ig ,R ib .

[0269] In this embodiment, a nonlinear least squares method is used to optimize R.ig ,R ib The formula for calculating the residual is as follows:

[0270]

[0271] Where R ig Transformation between the IMU inertial coordinate system and the global coordinate system, R i For IMU data, skeletal orientation data, R ib Transformation between the IMU local coordinate system and the corresponding skeleton local coordinate system This provides predicted bone orientation data in the global coordinate system. This refers to the target skeleton orientation data in the global coordinate system. Specifically, it requires that the skeleton orientation measured by the IMU be consistent with the skeleton orientation predicted by the monocular SMPL, where Φ represents the extracted axis-angle vector. Each IMU sensor and its attached skeleton requires optimization to solve for R. ib Compared to traditional methods that optimize R individually for each IMU sensor, this approach offers a more efficient alternative. ig In this disclosure, the same R is jointly optimized for each IMU sensor. ig ,R ig This represents the transformation between the inertial coordinate system and the global world coordinate system, which should remain invariant for each IMU sensor. Joint optimization guarantees R... ig Maintaining consistency across all IMU sensors yields more accurate results.

[0272] Therefore, this disclosure proposes a method for error constraint processing of the predicted action data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame, as well as an optimization calibration method for R. ig ,R ib This method is of great significance for creating multimodal motion capture datasets.

[0273] Correspondingly, the constraint process for predicting action data can also be found in [reference needed]. Figure 13 It mainly includes:

[0274] 1. Acquire image data and IMU data through the data acquisition and processing module;

[0275] 2. Based on the collected data, keypoint detection and Kalman filtering smoothing are performed to obtain smoothed keypoints;

[0276] 3. Perform monocular SMPL prediction based on image data to obtain initial SMPL parameters;

[0277] 4. Based on the acquired image data, IMU data, and smoothing key points, the initial SMPL parameters are optimized to obtain the optimized SMPL parameters.

[0278] This disclosure, under a monocular camera setup, adds sparse IMU sensor data and designs various constraints to effectively improve motion capture performance under self-occlusion conditions, thereby enhancing overall motion capture accuracy and temporal stability. A multi-level thresholding scheme for 2D keypoint confidence can adaptively handle motion sequences of varying complexity, improving the capture performance of complex motion sequences. Through iterative optimization of R... ig ,R ib It can obtain accurate calibration results, which is of great significance for the production of multimodal datasets.

[0279] Figure 14 This is an object motion data processing apparatus illustrated according to an exemplary embodiment. (Refer to...) Figure 14 The device includes:

[0280] The target data acquisition unit 1410 is configured to acquire an image frame sequence containing a target object, and object pose data corresponding to each image frame in the image frame sequence; the object pose data is obtained based on a pose measurement module set on the target object.

[0281] The first data prediction unit 1420 is configured to perform motion data prediction on the target object in the target image frame to obtain predicted motion data corresponding to the target image frame; the target image frame is an image frame in the image frame sequence that has not undergone motion data prediction.

[0282] Error constraint unit 1430 is configured to perform error constraint processing on the predicted motion data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame, so as to obtain target motion data corresponding to the target image frame.

[0283] In one exemplary embodiment, the error constraint unit includes:

[0284] The first driving unit is configured to perform action-driven processing based on the predicted action data to obtain predicted driving data; the predicted driving data is observation data corresponding to the predicted action data.

[0285] The first constraint processing unit is configured to perform error constraint processing on the prediction driving data based on the target image frame and the object pose data corresponding to the target image frame, to obtain target driving data.

[0286] The target action data determination unit is configured to determine the target action data based on the target driving data.

[0287] In one exemplary embodiment, the error constraint unit includes:

[0288] The first detection unit is configured to perform key point detection based on the target image frame to obtain the key point information of the target object.

[0289] The second driving unit is configured to perform action-driven processing based on the predicted action data to obtain the predicted key point information of the target object.

[0290] The second constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the error information between the detected key point information and the predicted key point information, and the object pose data corresponding to the target image frame, to obtain the target motion data.

[0291] In one exemplary embodiment, the detection key point information includes the detection confidence scores corresponding to each of the multiple target key points;

[0292] The second constraint processing unit includes:

[0293] The target confidence threshold determination unit is configured to perform a process based on the detection confidence of each of the plurality of target key points in the target image frame, the detection confidence of each of the plurality of target key points in historical image frames, and the smoothing confidence threshold corresponding to the historical image frame, to obtain a target confidence threshold corresponding to the target image frame; the historical image frame is an image frame in the image frame sequence that is temporally preceding the target image frame; the smoothing confidence threshold is determined based on the average confidence of the plurality of target key points in the historical image frames;

[0294] The first weight determination unit is configured to perform an operation based on a preset confidence threshold, a target confidence threshold corresponding to the target image frame, and the detection confidence of each of the plurality of target key points in the target image frame, to determine a first weight corresponding to each of the plurality of target key points in the target image frame; wherein the preset confidence threshold is less than the target confidence threshold corresponding to the target image frame;

[0295] The first weighting unit is configured to perform weighted processing on the error information of the detected key point information and the predicted key point information based on the first weights corresponding to the plurality of target key points, and generate first weighted error information.

[0296] The third constraint processing unit is configured to perform error constraint processing on the predicted action data based on the first weighted error information and the object pose data corresponding to the target image frame, so as to obtain the target action data.

[0297] In one exemplary embodiment, the error constraint unit includes:

[0298] The first coordinate system transformation unit is configured to transform the object pose data corresponding to the target image frame from the inertial coordinate system to the global coordinate system to obtain the target skeleton orientation data of the target object in the global coordinate system.

[0299] The third driving unit is configured to perform motion-driven processing based on the predicted motion data to obtain the predicted bone orientation data of the target object in the global coordinate system.

[0300] The first error information determination unit is configured to generate bone orientation error information based on the target bone orientation data and the predicted bone orientation data.

[0301] The fourth constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the target image frame and the bone orientation error information to obtain the target motion data.

[0302] In one exemplary embodiment, the fourth constraint processing unit includes:

[0303] The second detection unit is configured to perform key point detection based on the target image frame to obtain key point information of the target object; the key point information includes the detection confidence of each of the multiple target key points; the multiple skeletons of the target object in the target image frame correspond to the multiple target key points.

[0304] The second weight determination unit is configured to determine a second weight corresponding to any bone and a third weight corresponding to the target key point of any bone when the detection confidence of the target key point corresponding to any bone is less than a target confidence threshold; the second weight is greater than the third weight.

[0305] The second weighting unit is configured to perform weighted processing on the bone orientation error information based on the second weight corresponding to each bone, and generate second weighted error information.

[0306] The third weighting unit is configured to perform weighted processing on the error information between the detection information and the prediction information of the target key point based on the third weight corresponding to the target key point of any skeleton, to obtain the third weighted error information; the prediction information of the target key point is obtained based on the action-driven processing of the predicted action data.

[0307] The fifth constraint processing unit is configured to perform error constraint processing on the predicted action data based on the second weighted error information and the third weighted error information to obtain the target action data.

[0308] In one exemplary embodiment, the predicted motion data includes object shape prediction data of the target object;

[0309] The error constraint unit includes:

[0310] The shape error information determination unit is configured to perform shape error information determination based on object shape prediction data of the target object in the target image frame and shape smoothing data corresponding to historical image frames; the historical image frames are image frames in the image frame sequence that are temporally preceding the target image frame; the shape smoothing data are obtained by temporally smoothing the object shape data of the target object in the historical image frames.

[0311] The sixth constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the shape error information, the target image frame, and the object pose data corresponding to the target image frame, to obtain the target motion data.

[0312] In one exemplary embodiment, the predicted motion data includes joint prediction data of the target object;

[0313] The error constraint unit includes:

[0314] The joint error information determination unit is configured to perform joint error information based on joint prediction data corresponding to the target image frame and joint constraint data corresponding to the previous image frame of the target image frame; the joint constraint data corresponding to the previous image frame is obtained by performing data error constraints on the joint prediction data of the target object in the previous image frame; the previous image frame is an image frame that is temporally preceding the target image frame in the image frame sequence and is adjacent to the target image frame;

[0315] The third weight determination unit is configured to determine the fourth weight based on the motion data loss information corresponding to the previous image frame; the motion data loss information corresponding to the previous image frame is determined based on the object pose data corresponding to the previous image frame and the constraint driving data corresponding to the previous image frame; the constraint driving data corresponding to the previous image frame is obtained by applying data error constraints to the predicted motion data corresponding to the previous image frame.

[0316] The fourth weighting unit is configured to perform weighting processing on the joint error information based on the fourth weight to generate fourth weighted error information;

[0317] The iterative update unit is configured to perform iterative updates on the joint prediction data based on the fourth weighted error information to obtain the joint constraint data corresponding to the target image frame.

[0318] The seventh constraint processing unit is configured to perform error constraint processing on the predicted motion data based on the target image frame, the object pose data corresponding to the target image frame, and the joint constraint data corresponding to the target image frame, to obtain the target motion data.

[0319] In one exemplary embodiment, the apparatus further includes:

[0320] The fourth driving unit is configured to execute constraint driving data corresponding to the target action data of the previous image frame;

[0321] The loss information unit is configured to perform an action data loss based on the object pose data corresponding to the previous image frame and the constraint driving data;

[0322] The third weight determination unit includes:

[0323] The weight calculation unit is configured to determine the fourth weight based on the motion data loss information corresponding to the previous image frame and the loss smoothing information corresponding to the previous image frame; the loss smoothing information corresponding to the previous image frame is obtained by temporally smoothing the motion data loss information of historical image frames in the image frame sequence that are temporally preceding the previous image frame.

[0324] In one exemplary embodiment, the apparatus further includes:

[0325] The calibration data acquisition unit is configured to acquire a calibration image frame containing the calibration object and the object pose data corresponding to the calibration image frame; the pose measurement module on the calibration object is configured in the same way as the pose measurement module on the target object.

[0326] The second data prediction unit is configured to perform motion data prediction on the calibration object in the calibration image frame to obtain predicted motion data corresponding to the calibration image frame.

[0327] The fifth driving unit is configured to perform motion-driven processing based on the predicted motion data corresponding to the calibration image frame to obtain the predicted bone orientation data of the calibration object in the global coordinate system.

[0328] The second coordinate system transformation unit is configured to perform coordinate system transformation processing on the object pose data corresponding to the calibration image frame based on the initial coordinate system transformation parameters, so as to obtain the calibration bone orientation data of the calibration object in the global coordinate system.

[0329] The second error information determination unit is configured to execute the predicted bone orientation data and the calibrated bone orientation data based on the calibrated object in the global coordinate system to obtain calibration orientation error information.

[0330] The third iterative update unit is configured to perform iterative updates on the initial coordinate system transformation parameters based on the calibration orientation error information to obtain the target coordinate system transformation parameters; the target coordinate system transformation parameters are used to perform coordinate system transformation processing on the object pose data corresponding to the target image frame.

[0331] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0332] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform any of the methods described above.

[0333] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform any of the methods described above.

[0334] Figure 15 This is a block diagram illustrating an electronic device for processing object motion data according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 15 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an object motion data processing method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0335] Figure 16 This is a block diagram illustrating an electronic device for processing object motion data according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 16 As shown, this electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an object action data processing method.

[0336] Those skilled in the art will understand that Figure 15 and Figure 16 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0337] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0338] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An object action data processing method characterized by, The method comprises the following steps: obtaining an image frame sequence containing a target object and object pose data corresponding to each image frame in the image frame sequence; the object pose data is obtained based on a pose measurement module arranged on the target object; action data prediction is performed on the target object in a target image frame to obtain predicted action data corresponding to the target image frame; the target image frame is an image frame in the image frame sequence that has not been subjected to action data prediction; based on the target image frame and the object pose data corresponding to the target image frame, error constraint processing is performed on the predicted action data corresponding to the target image frame to obtain target action data corresponding to the target image frame; the error constraint processing based on the target image frame and the object pose data corresponding to the target image frame on the predicted action data corresponding to the target image frame to obtain target action data corresponding to the target image frame comprises the following steps: based on the target image frame, key point detection is performed to obtain detection key point information of the target object; the detection key point information comprises detection confidence of each target key point; based on the predicted action data, action driving processing is performed to obtain predicted key point information of the target object; based on the detection confidence of each target key point in the target image frame, the detection confidence of each target key point in a historical image frame, and a smoothing confidence threshold corresponding to the historical image frame, a target confidence threshold corresponding to the target image frame is obtained; the historical image frame is an image frame in the image frame sequence that is located in time sequence before the target image frame; the smoothing confidence threshold is determined based on the average confidence of the target key points in the historical image frame; based on a pre-set confidence threshold, the target confidence threshold corresponding to the target image frame, and the detection confidence of each target key point in the target image frame, a first weight corresponding to each target key point in the target image frame is determined; the pre-set confidence threshold is smaller than the target confidence threshold corresponding to the target image frame; based on the first weight corresponding to each target key point, weighted processing is performed on the error information of the detection key point information and the predicted key point information to generate first weighted error information; based on the first weighted error information and the object pose data corresponding to the target image frame, error constraint processing is performed on the predicted action data to obtain the target action data.

2. The method of claim 1, wherein, the error constraint processing based on the target image frame and the object pose data corresponding to the target image frame on the predicted action data corresponding to the target image frame to obtain target action data corresponding to the target image frame comprises the following steps: based on the predicted action data, action driving processing is performed to obtain predicted driving data; the predicted driving data is observation data corresponding to the predicted action data; based on the target image frame, the object pose data corresponding to the target image frame, error constraint processing is performed on the predicted driving data to obtain target driving data; determine the target action data based on the target driving data.

3. The method according to claim 1 or 2, characterized in that, The error constraint processing on the predicted action data based on the target image frame and the skeleton orientation error information comprises: transforming the object pose data corresponding to the target image frame from an inertial coordinate system to a global coordinate system to obtain target skeleton orientation data of the target object in the global coordinate system; performing action driving processing based on the predicted action data to obtain predicted skeleton orientation data of the target object in the global coordinate system; generating skeleton orientation error information based on the target skeleton orientation data and the predicted skeleton orientation data; performing error constraint processing on the predicted action data based on the target image frame and the skeleton orientation error information to obtain the target action data.

4. The method of claim 3, wherein, The error constraint processing on the predicted action data based on the target image frame and the skeleton orientation error information to obtain the target action data comprises: performing key point detection based on the target image frame to obtain detection key point information of the target object; the detection key point information comprises detection confidence of each target key point; a plurality of skeletons of the target object in the target image frame have a corresponding relationship with the plurality of target key points; in a case where the detection confidence of any skeleton corresponding target key point is less than a target confidence threshold, determining a second weight corresponding to the any skeleton and a third weight corresponding to the any skeleton corresponding target key point; the second weight is greater than the third weight; performing weighted processing on the skeleton orientation error information based on the second weight corresponding to the any skeleton to generate second weighted error information; performing weighted processing on error information between the detection information of the target key point and predicted information of the target key point based on the third weight corresponding to the target key point to obtain third weighted error information; the predicted information of the target key point is obtained based on action driving processing on the predicted action data; performing error constraint processing on the predicted action data based on the second weighted error information and the third weighted error information to obtain the target action data.

5. The method of claim 1, wherein, The predicted action data comprises object shape prediction data of the target object; The error constraint processing on the predicted action data based on the target image frame and the object pose data corresponding to the target image frame to obtain the target action data corresponding to the target image frame comprises: determining shape error information based on the object shape prediction data of the target object in the target image frame and shape smoothing data corresponding to a historical image frame; the historical image frame is an image frame in the image frame sequence that is located in time sequence before the target image frame; the shape smoothing data is obtained based on time sequence smoothing of object shape data of the target object in the historical image frame; The target action data is obtained by performing error constraint processing on the predicted action data based on the shape error information, the target image frame, and object pose data corresponding to the target image frame.

6. The method of claim 1 or 2, wherein, The predicted action data includes joint prediction data of the target object. The target action data corresponding to the target image frame is obtained by performing error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and object pose data corresponding to the target image frame, and the target action data corresponding to the target image frame is obtained by performing error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and object pose data corresponding to the target image frame. Joint error information is obtained based on joint prediction data corresponding to the target image frame and joint constraint data corresponding to a previous image frame of the target image frame; the joint constraint data corresponding to the previous image frame is obtained by performing data error constraint on joint prediction data of the target object in the previous image frame; the previous image frame is an image frame located in time sequence before the target image frame and adjacent to the target image frame in the image frame sequence; A fourth weight is determined based on action data loss information corresponding to the previous image frame; the action data loss information corresponding to the previous image frame is determined based on object pose data corresponding to the previous image frame and constraint driving data corresponding to the previous image frame; the constraint driving data corresponding to the previous image frame is obtained by performing data error constraint on predicted action data corresponding to the previous image frame; The joint error information is weighted based on the fourth weight to generate fourth weighted error information; The joint prediction data is iteratively updated based on the fourth weighted error information to obtain joint constraint data corresponding to the target image frame; The target action data is obtained by performing error constraint processing on the predicted action data based on the target image frame, object pose data corresponding to the target image frame, and joint constraint data corresponding to the target image frame.

7. The method of claim 6, wherein, Before the fourth weight is determined based on the action data loss information corresponding to the previous image frame, the method further includes: Constraint driving data corresponding to target action data of the previous image frame is determined; Action data loss information corresponding to the previous image frame is determined based on object pose data corresponding to the previous image frame and the constraint driving data; The fourth weight is determined based on the action data loss information corresponding to the previous image frame and loss smoothing information corresponding to the previous image frame; the loss smoothing information corresponding to the previous image frame is obtained by performing time sequence smoothing on action data loss information of historical image frames located in time sequence before the previous image frame in the image frame sequence. Before the target action data corresponding to the target image frame is obtained by performing error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and object pose data corresponding to the target image frame, the method further includes:

8. The method of claim 1, wherein, ​ Obtaining a calibration image frame containing a calibration object, and object pose data corresponding to the calibration image frame; the pose measurement module on the calibration object is consistent with the setting mode of the pose measurement module on the target object; Performing action data prediction on the calibration object in the calibration image frame to obtain predicted action data corresponding to the calibration image frame; Performing action driving processing based on the predicted action data corresponding to the calibration image frame to obtain predicted skeletal orientation data of the calibration object in the global coordinate system; Performing coordinate system transformation processing on the object pose data corresponding to the calibration image frame based on initial coordinate system transformation parameters to obtain calibration skeletal orientation data of the calibration object in the global coordinate system; Obtaining calibration orientation error information based on the predicted skeletal orientation data of the calibration object in the global coordinate system and the calibration skeletal orientation data; Performing iterative updating on the initial coordinate system transformation parameters based on the calibration orientation error information to obtain target coordinate system transformation parameters; the target coordinate system transformation parameters are used for coordinate system transformation processing on the object pose data corresponding to the target image frame.

9. An object action data processing apparatus, characterized by comprising: Comprise: An acquisition unit configured to perform obtaining an image frame sequence containing a target object, and object pose data corresponding to each image frame in the image frame sequence; The object pose data is obtained based on a pose measurement module arranged on the target object; A data prediction unit configured to perform action data prediction on the target object in a target image frame to obtain predicted action data corresponding to the target image frame; The target image frame is an image frame in the image frame sequence that has not been subjected to action data prediction; An error constraint unit configured to perform error constraint processing on the predicted action data corresponding to the target image frame based on the target image frame and the object pose data corresponding to the target image frame to obtain target action data corresponding to the target image frame; The error constraint unit comprises: A first detection unit configured to perform key point detection based on the target image frame to obtain detection key point information of the target object; the detection key point information comprises detection confidence of each target key point; A second driving unit configured to perform action driving processing based on the predicted action data to obtain predicted key point information of the target object; A target confidence threshold determination unit configured to perform obtaining a target confidence threshold corresponding to the target image frame based on detection confidence of each target key point in the target image frame, detection confidence of each target key point in a historical image frame, and a smoothing confidence threshold corresponding to the historical image frame; the historical image frame is an image frame in the image frame sequence that is located before the target image frame in time sequence; the smoothing confidence threshold is determined based on average confidence of the target key points in the historical image frame; The first weight determination unit is configured to determine a first weight corresponding to each of the plurality of target key points in the target image frame based on a preset confidence threshold, a target confidence threshold corresponding to the target image frame, and a detection confidence corresponding to each of the plurality of target key points in the target image frame; and the preset confidence threshold is less than the target confidence threshold corresponding to the target image frame. The first weighting unit is configured to perform weighting processing on the error information of the detection key point information and the predicted key point information based on the first weight corresponding to each of the plurality of target key points, to generate first weighted error information. The third constraint processing unit is configured to perform error constraint processing on the predicted action data based on the first weighted error information and object pose data corresponding to the target image frame, to obtain the target action data.

10. An electronic device, comprising: comprise: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the object action data processing method of any one of claims 1 to 8.

11. A computer readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can perform the object action data processing method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • System and method for tracking hand motion using strong coupling fusion of image sensor and inertial sensor

    KR102456872B1