Pose correction and its model training methods, electronic devices, computer storage media and software products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-14
AI Technical Summary
然而,现有的动作捕捉技术在处理大尺度场地内的刚体定位任务时,普遍存在定位距离与精度的矛盾,即:随着定位距离的增加,系统稳定性大幅下降,导致输出的位姿轨迹在连续帧上产生剧烈抖动,从而形成数据层面的噪声
[0010]根据本申请实施例提供的方案,以设置于目标对象上的多个标记点为基础,对包含该目标对象的视频帧序列中的目标对象的位姿观测数据序列,进行针对多个标记点的时空特征提取,从而从空间维度和时间维度上对多个标记点进行空间关系和时序关系的建模,获得的时空特征既能够反映多个标记点之间的空间位置关系,又能够反映动态变化关系。在此基础上,确定多个标记点对应的位姿校正数据,并基于此对目标对象进行位姿校正,可以使得校正后的目标对象的位姿更加准确、平滑,有效抑制动作捕捉技术在大尺度场景下因定位距离增加而引入的轨迹抖动噪声,提升视频帧序列中目标对象的渲染稳定性与虚实合成画面的连贯性,进而改善虚拟拍摄所得视频的整体质量。
Smart Images

Figure CN122574076A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video technology, and in particular to a pose correction method, a pose correction model training method, an electronic device, a computer storage medium, and a computer program product. Background Technology
[0002] Virtual filming is a technique that uses a virtual scene as the background and a real object as the foreground for filming. Taking the presentation of a virtual scene on an LED screen as an example, when actors perform in front of the LED screen, the pre-made virtual scene is played on the LED screen. At the same time, the director's monitor (rendering workstation) displays the effect of the actors filming in the virtual scene in real time.
[0003] Motion capture technology is a crucial tool in the process of generating videos based on virtual shooting. Motion capture utilizes optical, inertial, or mechanical sensors to record and digitize the position coordinates and rotational posture of objects or human bodies in a real-world scene in three-dimensional space in real time. It is a fundamental technology for mapping physical motion to a virtual digital environment. However, existing motion capture technologies generally suffer from a trade-off between positioning distance and accuracy when handling rigid body positioning tasks in large-scale environments. Specifically, as the positioning distance increases, system stability decreases significantly, causing severe jitter in the output pose trajectory across consecutive frames, thus creating noise at the data level. This noise at the data level is directly transmitted to the rendering end, causing visible jitter in the video frame sequence and severely affecting the continuity of the composite image. Summary of the Invention
[0004] In view of this, embodiments of this application provide a pose correction and model training scheme to at least partially solve the above problems.
[0005] According to a first aspect of the embodiments of this application, a pose correction method is provided, comprising: acquiring a video frame sequence containing a target object and a pose observation data sequence of the target object, wherein the target object is provided with a plurality of marker points for pose estimation; extracting spatiotemporal features of the plurality of marker points based on the pose observation data sequence to obtain corresponding spatiotemporal features; determining pose correction data corresponding to the plurality of marker points based on the spatiotemporal features; and performing pose correction on the target object in the video frame sequence based on the pose correction data.
[0006] According to a second aspect of the embodiments of this application, a pose correction model training method is provided, comprising: acquiring video frame sequence samples and pose observation data sequence samples of sample objects contained in the video frame sequence samples, wherein the sample objects are provided with multiple marker points for pose estimation; using a pose correction model to be trained, performing spatiotemporal feature extraction on the multiple marker points based on the pose observation data sequence samples to obtain corresponding spatiotemporal feature samples; determining pose correction data samples corresponding to the multiple marker points based on the spatiotemporal feature samples; and training the pose correction model based on the pose correction data samples and a preset loss function.
[0007] According to a third aspect of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method described in the first or second aspect.
[0008] According to a fourth aspect of the embodiments of this application, a computer storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the method as described in the first or second aspect.
[0009] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that instruct a computing device to perform an operation corresponding to the method described in the first or second aspect.
[0010] According to the scheme provided in the embodiments of this application, based on multiple marker points set on the target object, the spatiotemporal features of the target object's pose observation data sequence in the video frame sequence containing the target object are extracted for multiple marker points. This allows for the modeling of spatial and temporal relationships between the multiple marker points from both spatial and temporal dimensions. The obtained spatiotemporal features reflect both the spatial positional relationships and dynamic changes between the multiple marker points. Based on this, pose correction data corresponding to the multiple marker points is determined, and pose correction is performed on the target object accordingly. This results in a more accurate and smoother pose of the corrected target object, effectively suppressing trajectory jitter noise introduced by motion capture technology in large-scale scenes due to increased positioning distance. This improves the rendering stability of the target object in the video frame sequence and the coherence of the virtual-real composite image, thereby improving the overall quality of the video obtained from virtual shooting. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0012] Figure 1 This is a schematic diagram of an exemplary virtual shooting system applicable to embodiments of this application; Figure 2A This is a flowchart illustrating the steps of a pose correction model training method according to an embodiment of this application. Figure 2B for Figure 2A A schematic diagram of an exemplary model structure in the illustrated embodiment; Figure 2C for Figure 2A A schematic diagram illustrating an example of a training process in the illustrated embodiment; Figure 3A This is a flowchart illustrating the steps of a pose correction method according to an embodiment of this application. Figure 3B for Figure 3A A schematic diagram of a scenario example in the illustrated embodiment; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0013] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.
[0014] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.
[0015] To facilitate understanding of the solutions provided in the embodiments of this application, the following is combined with... Figure 1 First, we will provide an example of the virtual shooting system to which it is applicable.
[0016] As mentioned earlier, virtual filming technology involves both "real" and "virtual." The "real" refers to the physical screens, cameras, and workstations in the actual filming environment. The "virtual" refers to the virtual scenes constructed based on filming needs. These virtual scenes can be presented through physical screens, or projected onto surfaces such as screens or walls using a physical screen as an intermediary. During virtual filming, actors in the physical environment can perform in front of physical screens, screens, or walls displaying virtual backgrounds, while being filmed by physical cameras, thus achieving a combination of "virtual" and "real."
[0017] For example, refer to Figure 1 This illustrates an exemplary virtual shooting system applicable to embodiments of this application. For example... Figure 1 As shown, the virtual shooting system 100 may include: a physical camera 102, a screen 104 for presenting a virtual scene, and a rendering workstation 106. The image presented on the screen 104 is a view of a pre-constructed virtual scene from the shooting perspective of the virtual camera. Those skilled in the art should understand that in this embodiment, the positional relationship between the physical camera 102 and the screen 104 is not limited; the physical camera 102 only needs to be able to capture a good image of the actual object and a complete or partial area of the virtual scene presented on the screen 104. When an actual object, such as an actor, performs actions in front of the screen 104, with the virtual scene presented on the screen 104 as the background, the physical camera 102 captures the action, achieving effective "virtual" and "real" fusion to create an effect similar to an actual object moving in a real scene. For example, the screen 104 in this example can be an LED screen.
[0018] The physical camera 102 captures video based on the virtual scene presented on the screen 104 and the actors' performance in this scene, and transmits the capture results to the rendering workstation 106 in the form of a video stream for real-time display on the monitor of the rendering workstation 106.
[0019] In the aforementioned system, the virtual shooting system 100 further includes a motion capture camera array 108, which may include multiple optical cameras deployed around the actual shooting environment to capture video frames of multiple marker points on the target object (a rigid body in the actual shooting environment) in real time. The marker points are identifiable points attached to the surface of the target object for assisting pose estimation. In the motion capture system (which includes at least the motion capture camera array 108 and the motion capture data processing software, such as a motion capture workstation 110), the marker points are typically implemented in the form of LED reflective spheres or reflective patches. Multiple optical cameras simultaneously capture images from different angles, allowing the calculation of the three-dimensional spatial coordinates of each marker point. In this embodiment, the marker points are implemented in the form of reflective patches. For example, eight reflective patches can be set on the rigid body, and these eight reflective patches are not collinear. In the video frames captured by the motion capture camera array, these eight reflective patches are reflected as eight reflective points.
[0020] The virtual shooting system 100 can perform real-time calculations based on video frames captured by the motion capture camera array 108 to obtain the three-dimensional spatial coordinates of each marker point. For example, the motion capture camera array 108 can perform real-time calculations on the captured video frames to obtain the three-dimensional spatial coordinates of each marker point; or, the virtual shooting system 100 can use the motion capture workstation 110 to calculate the three-dimensional spatial coordinates of each marker point based on the video frames captured by the motion capture camera array 108. Furthermore, the motion capture workstation 110 can calculate the six-degree-of-freedom pose of the target object (pose is a spatial state description of the target object in three-dimensional space, covering its position components in the three coordinate axes and its rotational components around the three coordinate axes) based on the three-dimensional spatial coordinates of each marker point, that is, a six-dimensional state quantity including three-dimensional position coordinates and three-dimensional rotational attitude.
[0021] For example, in the virtual shooting system 100 described above, the physical camera 102 can be specifically used for video shooting. It is aimed at the actor and the virtual background screen to capture high-quality video frame images for post-production, and its output is directly presented on the director's monitor of the rendering workstation 106. The motion capture camera array 108 can be an optical sensor system independently deployed around the shooting location and dedicated to pose estimation. Its working object is multiple marker points attached to the target object (rigid body or actor), and it captures multi-view images of the marker points to calculate the three-dimensional spatial coordinates of the marker points.
[0022] The two camera systems described above operate in parallel without interfering with each other during virtual filming: the motion capture camera array 108 calculates the pose data of the target object in real time, and after pose correction by the motion capture workstation 110, outputs it to the rendering workstation 106 to drive the synchronous movement of the 3D model in the virtual scene, forming a virtual background image that matches the real site conditions; the physical camera 102 records the actor's performance in front of this virtual background synchronously, achieving virtual-real compositing. Thus, the two systems each perform their respective functions, jointly supporting virtual filming.
[0023] However, in related technologies, when the distance between the target object and the motion capture camera array 108 is far, the pose trajectory calculated by the motion capture system often exhibits severe jitter between consecutive video frames. This jitter noise is directly transmitted to the rendering workstation 106, causing the corresponding 3D model in the virtual scene to produce visible tremors, which seriously affects the continuity of the virtual-real composite image.
[0024] To address this, this application introduces a pose correction stage between the motion capture stage and the video frame rendering stage. The pose observation data sequence is obtained by taking the three-dimensional spatial coordinates of the target object from the video frame sequence formed by multiple video frames acquired by the motion capture camera array 108. Spatiotemporal features are extracted based on the pose observation data sequence, and then pose correction data is determined. The pose of the target object is corrected accordingly. The rendering stage then drives the three-dimensional model to be rendered based on the pose correction results, thereby eliminating the visual jitter present in related technologies.
[0025] In one alternative approach, pose correction can be achieved through a pose correction model. To facilitate understanding of the embodiments of this application, the training process of the pose correction model will be described first, followed by the description of the pose correction process.
[0026] Reference Figure 2A The diagram illustrates a flowchart of a pose correction model training method according to an embodiment of this application. The pose correction model training method includes the following steps: Step S202: Obtain video frame sequence samples and pose observation data sequence samples of sample objects contained in the video frame sequence samples.
[0027] The video frame sequence sample contains multiple consecutive video frames, and at least some of the video frames contain a sample object. In this embodiment, the sample object is set as a rigid body. When subjected to external forces, the shape and size of the rigid body do not change; that is, the distance between any two points within the rigid body remains constant. In practical applications, the rigid body can be implemented in any suitable form, such as a rigid sphere, a rigid cube, or a rigid cylinder, and this embodiment does not impose any limitations on this.
[0028] Furthermore, the sample object is equipped with multiple marker points for pose estimation. In this embodiment, eight LED reflective patches are set on the sample object. These eight LED reflective patches are not collinear and appear as eight LED reflective points in the video frame. The eight LED reflective points provide more redundant constraints for calculating the rigid body pose, improving the robustness of the calculation; moreover, they can cover a larger rigid body surface, reducing the risk of failure due to partial occlusion or detachment of reflective points. In dynamic environments (such as during motion), the eight reflective points can better handle rapid movement or changes in viewpoint. However, those skilled in the art should understand that the eight LED reflective points are merely illustrative; in practical applications, more or fewer reflective points can be used to calculate the rigid body pose.
[0029] In one example, the motion capture camera array can perform real-time calculations based on captured video frames to obtain the three-dimensional spatial coordinates P={p1,p2,...,p8} of eight LED reflective points on a rigid body in each video frame. These three-dimensional spatial coordinates {p1,p2,...,p8} are the raw output directly calculated by the motion capture camera array and serve as the fundamental data describing the position of each marker point in the world coordinate system.
[0030] Furthermore, to capture the dynamic trend of the sample object's motion, in this embodiment, instead of processing individual video frames in isolation, a sliding time window of length M is constructed. The pose observation data samples of the sample object in each video frame of the corresponding video frame window sequence form a pose observation data sequence sample, denoted as St={C t M+1 ,...,C t}, where C t This represents the 3D spatial coordinates of the 8 LED reflectors in frame t and the 6-DoF pose (6 Degrees of Freedom) of the rigid body determined by these coordinates, i.e., the pose observation data sample of the rigid body determined by the 8 LED reflectors in frame t. The 6-DoF pose of the rigid body in frame t is calculated based on the 3D spatial coordinates of the 8 LED reflectors. For example, the position components of each reflector along the three coordinate axes in the 6-DoF pose originate from the 3D spatial coordinates of each reflector, while the rotational components around the three coordinate axes can be calculated by the motion capture workstation based on the 3D spatial coordinates. This time-window-based data organization provides cross-frame contextual information for subsequent temporal modeling, laying the foundation for capturing dynamic motion trends. Since the 6-DoF pose is calculated based on 3D spatial coordinates, in another optional approach, the above St can also be expressed as St={P} t M+1 ,...,Pt}, where P t The three-dimensional spatial coordinates of the 8 LED reflective points in frame t are represented. Here, "frame t" can correspond to the "frame t" image captured from different angles by multiple optical cameras in the motion capture camera array 108 at the same shooting moment.
[0031] Optionally, to eliminate the influence of dimensions and accelerate model convergence, the aforementioned data St can be normalized to transform it into a uniform numerical range. Furthermore, the video frame sequence samples also contain pose ground truth labels for each marker point, corresponding to the two forms of St mentioned above. These pose ground truth labels can be either 6-DoF pose ground truth labels or three-dimensional spatial coordinate ground truth labels, serving as supervision conditions for model training. These pose ground truth labels indicate the six-DOF pose ground truth or three-dimensional spatial coordinate ground truth values that accurately reflect the motion state of the sample object, obtained through high-precision measurement under noise-free or low-noise conditions.
[0032] For each video frame sequence sample, the pose correction model takes its corresponding normalized St as input and the corresponding pose ground truth label as supervision condition for subsequent model training.
[0033] This step prepares the sample data for training the pose correction model.
[0034] Step S204: Using the pose correction model to be trained, perform spatiotemporal feature extraction on multiple marker points based on pose observation data sequence samples to obtain corresponding spatiotemporal feature samples; based on the spatiotemporal feature samples, determine the pose correction data samples corresponding to multiple marker points.
[0035] This step is implemented by a pose correction model, which first obtains spatiotemporal feature samples for multiple marker points. Spatiotemporal features are feature vectors or feature maps extracted by jointly analyzing multiple consecutive frames of video frame sequence samples (or video frame sequences) that describe the correlation between objects or scenes in spatial dimensions (such as shape, texture, and position) and temporal dimensions (such as motion trajectory, velocity, and acceleration). In this embodiment, the multiple consecutive frames of video frame sequence samples (or video frame sequences) are represented by St after the aforementioned normalization process. Spatiotemporal features integrate the static spatial features and dynamic temporal features of the sample object, forming a high-order abstract representation of the sample object (specifically, the reflective points of the sample object in this embodiment). To distinguish it from the inference stage, the spatiotemporal features obtained in the training stage are referred to as spatiotemporal feature samples in this embodiment. Based on the obtained spatiotemporal feature samples, the pose correction model performs further processing to determine the pose correction data samples corresponding to multiple marker points, the specific implementation of which will be described in detail below.
[0036] In one alternative approach, the pose correction model can be a Transformer-based model. An exemplary pose correction model structure is as follows: Figure 2B As shown, it includes an input encoding layer, a spatial feature extraction layer, a temporal feature extraction layer, a feature fusion layer, and a residual regression layer connected in sequence.
[0037] in: The input coding layer receives the input, normalized pose observation data sequence sample St of the sample object, and maps the pose observation data samples of multiple marker points in each frame into a fixed-dimensional embedding vector for subsequent processing.
[0038] The spatial feature extraction layer is used to extract the spatial features of multiple marker points based on pose observation data sequence samples, so as to model the relative geometric structure relationship between multiple marker points. When some marker points are disturbed by noise or temporarily occluded, the spatial information of other marker points can be used to help infer the reasonable position of the disturbed point, thereby improving the robustness of the system to local noise.
[0039] In one alternative approach, the spatial feature extraction layer can use self-attention processing to extract features based on the relative geometric relationships of multiple marker points from pose observation data sequence samples, thereby obtaining corresponding spatial feature samples. The self-attention mechanism can automatically learn the association weights between elements within the pose observation data sequence samples. For the spatial dimension, multiple marker points within the same frame are treated as a sequence, and the self-attention mechanism can model the relative geometric relationships between any two marker points without needing to manually define the topological structure between them beforehand. This allows the pose correction model to adaptively learn rigid body structural constraints in a data-driven manner, resulting in stronger generalization ability. For example, the spatial feature extraction layer can be implemented as a Transformer encoder structure.
[0040] The temporal feature extraction layer is used to extract temporal features of multiple marker points between different video frames based on pose observation data sequence samples, and to model the dependency relationship between multiple marker points between adjacent frames. It can distinguish between motion trends that conform to physical laws and noise jumps that violate motion laws from the temporal context.
[0041] In one alternative approach, the temporal feature extraction layer can utilize multi-head attention processing to extract features based on the adjacent frame dependencies of multiple marker points from the pose observation data sequence samples, thereby obtaining corresponding temporal feature samples. In the Transformer architecture, the multi-head attention mechanism is used to establish global dependencies between any positions within a sequence. In the temporal dimension, the pose observation data sequence samples corresponding to consecutive video frames within a time window are considered as a single sequence. The multi-head attention mechanism can then establish global dependencies between the current frame and the sample objects in the preceding and following frames. This allows the pose correction model to perceive the motion inertia of sample objects in past frames and the potential trends of sample objects in future frames. Thus, it can judge whether the pose of the current frame sample object conforms to physical laws from a holistic rather than a local perspective, accurately identifying and eliminating instantaneous jumps (noise) that do not conform to motion laws, while retaining legitimate sudden stops or turns. For example, the temporal feature extraction layer can also be implemented as a Transformer encoder structure.
[0042] It should be noted that those skilled in the art can flexibly set the specific network structure, number of attention heads, embedding dimension, and other hyperparameters of the spatial feature extraction layer and the temporal feature extraction layer, and the embodiments of this application do not impose any restrictions on this.
[0043] The feature fusion layer is used to fuse spatial feature samples and temporal feature samples to generate spatiotemporal feature samples rich in spatiotemporal context information. In one optional approach, spatial feature samples and temporal feature samples can be concatenated. However, this is not a limitation; other fusion methods, such as element-wise weighted summation, are also applicable to the schemes of this application embodiment. The spatiotemporal feature samples contain the relative geometric relationships and cross-frame temporal dependencies between multiple marker points. In one example, it can be represented as:
[0044] in, A sequence of pose observation data representing a sample object; This represents the shared model parameters between the spatial feature extraction layer and the temporal feature extraction layer. Represents spatiotemporal feature samples; This refers to spatiotemporal feature extraction methods. Within a unified Transformer encoder framework, this represents the overall underlying feature mapping parameters shared by both spatial and temporal feature extraction. Both feature extraction processes rely on the same set of encoding weights for data processing, effectively capturing spatial and temporal relationships. By sharing these underlying encoding parameters, the model can simultaneously optimize spatial and temporal awareness within a unified feature space during training, avoiding feature space fragmentation caused by parameter separation, reducing the total number of model parameters, and mitigating the risk of overfitting. This spatiotemporal feature sample integrates information from multiple marker points in video frame sequences across both spatial and temporal dimensions, providing a rich information foundation for subsequent processing.
[0045] The residual regression layer is used to transform spatiotemporal features into pose correction data. Unlike end-to-end solutions that directly predict the absolute pose of the target object, the pose observation data predicted by the residual regression layer in this embodiment deviates from the corresponding residual, i.e., noise. Therefore, on the one hand, the residual values have a smaller range relative to the absolute coordinates, resulting in a flatter loss function landscape, which helps alleviate gradient explosion or vanishing problems during training and accelerates model convergence; on the other hand, the model's attention is focused on compensating for noise-sensitive regions, avoiding excessive smoothing that damages details of legitimate sudden stops, turns, and other high-frequency actions.
[0046] In one optional approach, when performing residual regression correction, the residual regression layer obtains residual values based on spatiotemporal feature samples of multiple marker points and reference feature samples corresponding to the multiple marker points at their initial reference positions. Then, based on these residual values, the pose observation data samples of the multiple marker points are corrected. Here, the reference features represent the feature representations corresponding to each marker point when the sample object is in a certain reference static state or a preset initial pose. By comparing the current spatiotemporal features with the reference features, the deviation relative to the reference state can be more accurately located, thus making the representation of the residuals more physically meaningful and further improving the correction accuracy. Exemplarily, those skilled in the art can flexibly set a static state or a specific reference pose, such as during the calibration phase, as a reference to determine the reference features; this application embodiment does not impose any limitations on this. The calibration stage refers to the stage before pose correction. In this stage, the sample object or target object (such as a rigid body) is placed at a preset reference position (i.e., the initial reference position) to keep it in a static state. The motion capture camera array collects and calculates the three-dimensional spatial coordinates of each marker point at this time, thereby obtaining the corresponding pose data. The feature representation obtained after feature extraction of the above data is recorded as the reference feature of each marker point.
[0047] In one example, the residual regression layer can be implemented as a fully connected network, and its computation process can be described as follows:
[0048]
[0049] in, These are the model parameters for the residual regression layer; The residual value is the value predicted by the model, which is the difference between the spatiotemporal feature samples corresponding to the pose observation data sequence samples and the baseline features. The residual function is determined based on the baseline characteristics; This is a sequence sample of pose observation data; The pose data is the residual regression-corrected data, which is the final output pose-corrected data sample. Thus, the corrected pose data is obtained by superimposing the original pose observation data sample with the predicted residual value. The model only needs to focus on learning the distribution characteristics of the noise, without having to reconstruct the entire pose from scratch.
[0050] It should be noted that in pose observation data sequence samples, such as the St sequence sample, the coordinate sequence sample is St={P}. t M+1 ,...,P t When}, the residual value corresponds to the coordinate residual value; while the pose observation data sequence sample, such as the St sequence sample, is a six-degree-of-freedom pose sequence sample, i.e., St={Ct} When M+1,...,Ct}, the residual values correspond to the six-degree-of-freedom pose residual values. Furthermore, in actual pose correction, only the above-mentioned values can be considered. The pose observation data samples of the target, such as the pose observation data samples corresponding to the target video frame (e.g., the last frame in the window), are superimposed with the prediction residual values to determine the pose correction data samples; or, for the above... The pose observation data samples corresponding to some key video frames are superimposed with the predicted residual values to determine the pose correction data samples. Furthermore, the structure of the above pose correction model is merely an example. Those skilled in the art will understand that the specific implementation methods of each layer are not unique. Any structure capable of achieving the corresponding spatiotemporal feature extraction and residual regression functions is applicable to the scheme of this application embodiment, and this application embodiment does not impose any limitations on it.
[0051] Based on this, in one optional approach, performing spatiotemporal feature extraction on multiple marker points based on pose observation data sequence samples to obtain corresponding spatiotemporal feature samples may include: extracting features of the relative geometric structure relationships of multiple marker points based on pose observation data sequence samples to obtain corresponding spatial feature samples (e.g., through the spatial feature extraction layer of the pose correction model, extracting features of the relative geometric structure relationships of multiple marker points based on pose observation data sequence samples to obtain corresponding spatial feature samples); and extracting features of the adjacent frame dependencies of multiple marker points based on pose observation data sequence samples to obtain corresponding temporal feature samples (e.g., through the temporal feature extraction layer of the pose correction model, extracting features of the adjacent frame dependencies of multiple marker points based on pose observation data sequence samples to obtain corresponding temporal feature samples); and obtaining spatiotemporal feature samples based on spatial feature samples and temporal feature samples (e.g., through the feature fusion layer of the pose correction model, obtaining spatiotemporal feature samples based on spatial feature samples and temporal feature samples).
[0052] Based on this, optionally, determining the pose correction data samples corresponding to multiple marker points based on spatiotemporal feature samples may include: performing residual regression correction on multiple marker points based on spatiotemporal feature samples to obtain pose correction data samples corresponding to multiple marker points. Further optionally, performing residual regression correction on multiple marker points based on spatiotemporal feature samples to obtain pose correction data samples corresponding to multiple marker points can be implemented as follows: obtaining residual values based on the spatiotemporal feature samples of multiple marker points and the reference feature samples corresponding to multiple marker points at their initial reference positions; correcting the pose observation data samples of multiple marker points based on the residual values to obtain pose correction data samples corresponding to multiple marker points. For example, the pose correction data samples corresponding to multiple marker points can be obtained by correcting the pose observation data samples of multiple marker points based on spatiotemporal feature samples through the residual regression layer of the pose correction model.
[0053] Step S206: Train the pose correction model based on pose correction data samples and a preset loss function.
[0054] In this step, the predicted output of the pose correction model, i.e., the pose correction data sample, is compared with the true pose value to calculate the value of the loss function. This value is then used as a signal to backpropagate and update the parameters of the pose correction model, so that the pose correction model gradually converges to a state that can accurately correct pose noise until the training termination condition is met.
[0055] In one alternative approach, this step can be implemented as follows: based on pose correction data samples of multiple marker points and the ground truth pose values corresponding to multiple marker points, the pose correction model is trained using a preset loss function. The preset loss function includes: a first sub-loss function for evaluating the positional accuracy loss of multiple marker points, a second sub-loss function for evaluating the velocity smoothing loss of multiple marker points, and a third sub-loss function for evaluating the acceleration smoothing loss of multiple marker points. By explicitly embedding rigid body kinematic constraints such as velocity continuity and acceleration smoothness through this joint loss function, the model's output can conform to Newtonian mechanics. Compared to simple optimization methods targeting positional error, while this simple method can reduce the average positioning error of a single video frame, it cannot constrain the temporal continuity of the model's output. Even if the prediction error of each video frame is small, if there are irregular jumps in the prediction results between adjacent video frames, the final rendered 3D model will still exhibit high-frequency jitter. However, through the joint optimization of this embodiment, by explicitly introducing rigid body kinematics prior constraints into the loss function, the model output can achieve a smoother trajectory that better conforms to physical laws.
[0056] The first sub-loss function measures the loss of positional accuracy. In one alternative approach, the L1 norm positional loss can be used to measure the deviation between the corrected 3D feature point coordinates and their corresponding ground truth values, thus dominating the model's ability to correct for original observation noise. In one example, this first sub-loss function can be expressed as:
[0057] in, The number of video frames in a training batch; The number of marker points in each video frame; The model predicts the first The first frame Pose correction data samples of 1 marker point; This represents the corresponding truth value. The L1 norm is more robust to penalizing outlier noise points than the L2 norm, which helps improve the model's performance under extreme jitter conditions.
[0058] The second sub-loss function measures velocity smoothing loss. It applies a velocity continuity constraint by minimizing the first-order difference (velocity consistency) of pose predictions between adjacent video frames, thus suppressing high-frequency jitter in the model output and improving trajectory smoothness and controller robustness. In one example, it can be expressed as:
[0059] in, and The predicted numbers are respectively the first two. Frame and the The first frame The pose correction data samples for each marker point are the same as those in the first sub-loss function. This loss function term keeps the displacement between adjacent video frames smooth, thus constraining the change in motion speed of the sample object to prevent it from becoming too drastic.
[0060] The third sub-loss function measures the acceleration smoothing loss. It imposes an acceleration rationality constraint by minimizing the second-order difference (acceleration consistency) of pose predictions between adjacent video frames, thus suppressing non-physical acceleration abrupt changes and improving trajectory smoothness and controller robustness. In one example, it can be expressed as:
[0061] in, , and The predicted numbers are respectively the first two. Frame, First In the frame and the first Frame number The pose correction data samples for each marker point are the same as those in the first sub-loss function. This loss function term can constrain the rate of change of motion acceleration of the sample object, preventing sudden acceleration changes in the corrected trajectory, thereby further eliminating residual jitter at the rendering end.
[0062] In actual training, the three sub-loss functions are combined in a weighted manner to form the total loss. The weight coefficients can be flexibly adjusted by those skilled in the art according to the actual needs of the scenario, and this application does not limit this.
[0063] The following, combined with Figure 2C The above training process will be illustrated with an example.
[0064] Figure 2C The training process described above is divided into four stages: data preprocessing and time series construction, spatiotemporal feature encoding, residual regression correction, and multi-objective joint optimization. These stages will be explained in the following order.
[0065] Phase 1: Data preprocessing and time series construction phase.
[0066] In this stage, the raw data calculated in real time by the motion capture camera array is received, namely the three-dimensional spatial coordinates {p1,p2,...,p8} of the eight LED reflective points on the rigid body surface in each video frame. Then, a sliding time window of length M is constructed. Based on the three-dimensional spatial coordinates {p1,p2,...,p8} of the reflective points in a single video frame, an input sequence with temporal context is formed, namely, the pose observation data sequence samples corresponding to the sample objects in M frames, denoted as St={C t M+1,...,C t} or St={P t M+1 ,...,P t}, where C t P represents the three-dimensional spatial coordinates of the eight reflective points based on frame t and the initial 6-DoF pose of the rigid body calculated from them. t This represents the three-dimensional spatial coordinates of the eight LED reflective points in frame t. Then, the above data is normalized to eliminate dimensional differences and accelerate subsequent model convergence. After this stage of processing, the output is a standardized input sequence St that can be directly used by subsequent models.
[0067] Phase Two: Spatiotemporal Feature Encoding Phase.
[0068] In this example, the pose correction model is built on the Transformer structure, which receives the input sequence St and performs joint feature encoding (i.e., feature extraction) on the input sequence St in both spatial and temporal dimensions.
[0069] The spatial feature encoding part utilizes a self-attention mechanism to automatically learn the relative geometric relationships between eight marker points on a rigid body. Therefore, even if some marker points are briefly occluded, the model can infer the overall rigid body shape from the spatial distribution of the remaining marker points, enhancing its robustness to local disturbances. The temporal feature encoding part treats the pose observation data sequence samples corresponding to consecutive video frames within a time window as a sequence. It uses a multi-head attention mechanism to establish global dependencies between the current frame and previous / next frames, enabling the model to perceive the motion inertia of historical frames and the motion trend of future frames. This allows for accurate differentiation between legal actions that conform to physical laws and noisy transitions that violate motion laws, accurately identifying and deleting high-frequency transitions (noise), and enhancing global dependencies. The spatial and temporal feature encoding parts share parameters and output a spatiotemporal feature vector. , can be represented as ,in This is a sequence of pose observation data for the sample object; These are shared model parameters for the spatial feature extraction layer and the temporal feature extraction layer.
[0070] Phase 3: Residual Regression Correction Phase.
[0071] During this phase, through a fully connected network ( Figure 2C The diagram in the image is an "ERROR correction network" based on spatiotemporal feature vectors. Perform residual regression correction. In this example, the model does not directly predict the absolute 6-DoF pose or spatial coordinates of the rigid body, but focuses on learning the amount of difference (residual) between the pose observation data sequence samples and the baseline data corresponding to the baseline feature samples. Specifically, first, the residual values are predicted through a fully connected network. ,in These are the model parameters for the residual regression layer, i.e., the parameters of the fully connected network; subsequently, the predicted residuals are superimposed onto the pose observation data sequence samples. The corrected data is the pose correction data sample. This stage outputs a set of corrected pose correction data samples for rigid body markers. This serves as the predicted result for subsequent loss calculations.
[0072] Phase 4: Multi-objective joint optimization phase.
[0073] This stage utilizes a constructed 3D joint loss function that integrates geometric accuracy and kinematic constraints to perform end-to-end optimization of the model. (Position accuracy loss) The L1 norm is used to measure and correct the deviation between the six-DOF pose or 3D spatial coordinates and the corresponding ground truth, which dominates the model's denoising ability; velocity smoothing loss. Minimize the first-order difference of predicted coordinates between adjacent frames, and apply a velocity continuity constraint; use acceleration smoothing loss. Minimize the second-order difference between the predicted coordinates of adjacent frames and apply an acceleration rationality constraint. The three sub-loss functions are combined by weights to form the total loss, which is used as the optimization signal to backpropagate and train the pose correction model until the training termination condition is met.
[0074] Through close collaboration in the above four stages, the rigid body 6-DoF pose of 8 marker points was enhanced across the entire chain, fundamentally improving the rendering stability and realism in large-scale virtual shooting scenes.
[0075] The pose correction model training method in this embodiment trains the model based on pose correction data sequence samples from multiple marker points and the true pose value. It employs a multi-task joint loss function that simultaneously includes position accuracy loss, velocity smoothing loss, and acceleration smoothing loss. Position accuracy loss primarily eliminates the influence of observation noise on position coordinates. Velocity smoothing loss and acceleration smoothing loss, respectively, explicitly embed the laws of rigid body motion into the optimization objective from the first-order and second-order kinematic levels, in the form of mathematical constraints, ensuring that the model's output trajectory meets the requirements of velocity continuity and acceleration rationality. This synergistic optimization of the three losses ensures that the trained model not only approximates the true value in single-frame position accuracy but also outputs a smooth and continuous trajectory conforming to Newtonian mechanics at the temporal level. This mathematical optimization eliminates the root cause of high-frequency jitter at the rendering end, meeting the requirements of visually stable trajectory for film-level rendering.
[0076] After training the pose correction model, the trained model can be applied to the inference stage to perform real-time correction on the raw pose data output by the motion capture system during actual virtual shooting.
[0077] Reference Figure 3A The diagram illustrates a pose correction method according to an embodiment of this application, the method comprising the following steps: Step S302: Obtain the video frame sequence containing the target object and the pose observation data sequence of the target object.
[0078] The target object is marked with multiple marker points for pose estimation. For example, eight marker points.
[0079] For example, the target object can be a rigid body with multiple non-collinear marker points, such as eight LED reflective patches. The motion capture camera array captures video and outputs a series of ordered video frames at a certain frame rate, forming a video frame sequence. Each video frame (which may include video frames captured simultaneously from different angles by multiple optical cameras in the motion capture camera array 108, referred to here as "one frame") corresponds to the three-dimensional spatial coordinate observation values of multiple marker points on the target object at the time of that video frame's capture, as well as the six-degree-of-freedom pose estimation values calculated from them.
[0080] Similar to the training phase, the inference phase still uses a sliding time window to organize the six-degree-of-freedom pose estimation values or three-dimensional spatial coordinates of the target object in each video frame of the video frame sequence, forming a pose observation data sequence of the target object (three-dimensional spatial coordinate sequence or six-degree-of-freedom pose data sequence), and then performs normalization processing so that cross-frame context information can be obtained through spatiotemporal feature extraction.
[0081] This step transforms discrete single-frame pose observation data into a sequence that reflects temporal continuity by acquiring a video frame sequence containing the target object, providing an informational foundation for subsequent noise analysis based on spatiotemporal features. Using multi-frame sequences rather than isolated single-frame data as the processing unit allows subsequent steps to distinguish noise from real motion within the temporal context, providing a data foundation for the entire pose correction process.
[0082] Step S304: Based on the pose observation data sequence, perform spatiotemporal feature extraction for multiple marker points to obtain the corresponding spatiotemporal features.
[0083] In this step, on the one hand, in the spatial dimension, features are extracted based on the relative geometric structural relationships between multiple marker points within the same video frame; on the other hand, in the temporal dimension, features are extracted based on the motion dependencies between multiple marker points in adjacent video frames within a video frame sequence. These two aspects of features together constitute the spatiotemporal features of multiple marker points, thereby characterizing the motion state of the target object.
[0084] In one alternative approach, extracting spatiotemporal features from multiple marker points based on the pose observation data sequence to obtain the corresponding spatiotemporal features may include: extracting features based on the relative geometric structural relationships of multiple marker points from the pose observation data sequence to obtain corresponding spatial features; and extracting features based on the adjacent frame dependencies of multiple marker points from the pose observation data sequence to obtain corresponding temporal features; and obtaining spatiotemporal features based on spatial and temporal features. This allows for targeted modeling of features in both spatial and temporal dimensions, which are then fused to obtain a comprehensive representation, taking into account the informational characteristics of both dimensions.
[0085] Optionally, feature extraction of the relative geometric relationships of multiple marker points based on the pose observation data sequence to obtain the corresponding spatial features can include: using self-attention processing to extract features of the relative geometric relationships of multiple marker points based on the pose observation data sequence to obtain the corresponding spatial features. The self-attention mechanism, by calculating the attention weight matrix between elements within the sequence, can adaptively aggregate marker point information with high reference value for the current pose estimation. In the spatial dimension, when the observed coordinates of a marker point deviate significantly due to occlusion, abnormal reflection, or signal loss, the model can use the self-attention mechanism to refer to the spatial distribution of other marker points with better observation quality in the current frame, inferring the reasonable location of the disturbed marker point, thereby enhancing the system's resistance to local interference.
[0086] In another alternative approach, feature extraction of adjacent-frame dependencies for multiple marker points based on the pose observation data sequence to obtain corresponding temporal features can be achieved through multi-head attention processing. The multi-head attention mechanism uses multiple parallel attention heads to simultaneously establish dependencies between the current frame and each historical frame within the time window from different feature subspace perspectives, essentially examining the motion state of the target object in the current frame from multiple different temporal viewpoints. This allows the model to not only perceive short-term motion inertia but also grasp long-term motion trends, thereby accurately identifying temporal abrupt changes. If the position of a marker point in a frame exhibits a jump that does not conform to the motion pattern compared to the preceding and following frames, the temporal features extracted by the multi-head attention mechanism will form a significant signal response, providing a basis for judgment in subsequent residual regression.
[0087] This step extracts spatiotemporal features based on the pose observation data sequence. The extraction of spatial features enables the model to perceive the inherent geometric constraints of the rigid body structure, thus maintaining the consistency of overall pose estimation even when some marker points are disturbed. The extraction of temporal features enables the model to perceive the continuity of motion, allowing for identification and correction through contextual reference when abnormal frames occur. The fusion of these two approaches provides comprehensive features for subsequent pose correction data determination, incorporating both spatial geometric priors and temporal dynamic information. This improves the accuracy and robustness of the pose correction data, effectively resisting the data quality degradation caused by the decrease in signal-to-noise ratio due to increased positioning distance.
[0088] Step S306: Based on spatiotemporal features, determine the pose correction data corresponding to multiple marker points.
[0089] In this step, based on the obtained spatiotemporal features, data that can be directly used to correct the pose of the target object is generated to correct the original pose observation data output by the motion capture system. The data used to correct the target object's pose can be a residual relative to the original pose observation data. This embodiment uses a residual approach, which reduces the computational load of the model and improves the model's processing speed and efficiency.
[0090] In one alternative approach, this step can be implemented as follows: based on spatiotemporal features, residual regression correction is performed on multiple marker points to obtain pose correction data corresponding to the multiple marker points. Through residual regression correction, the amount of deviation from the original pose observation data (i.e., the residual) can be determined, and the predicted residual is superimposed on the original pose observation data as the correction result. This method replaces the method of predicting absolute coordinates in related techniques with the method of predicting the deviation, reducing the model's data processing load and allowing the model to focus on noise compensation.
[0091] Optionally, based on spatiotemporal features, residual regression correction is performed on multiple marker points to obtain pose correction data corresponding to these marker points. This can be achieved by: obtaining residual values based on the spatiotemporal features of multiple marker points and the reference features corresponding to these marker points at their initial reference positions; and then performing residual regression correction on the pose observation data of these marker points based on the residual values to obtain the pose correction data corresponding to these marker points. By introducing reference features and comparing the spatiotemporal features of the current state with the reference reference corresponding to these features, a more physically meaningful relative deviation can be used as the basis for representing the residuals. This gives the residual regression correction a clear reference anchor point, further improving the correction accuracy and stability.
[0092] This step, based on spatiotemporal features that integrate spatial and temporal information, determines pose correction data through residual regression, forming a comprehensive description of the target object's motion state at the current moment. Further, based on this description, the residual values of the pose observation data for each marker point in the current frame are predicted. These residual values comprehensively reflect the combined effects of various factors such as measurement noise, transmission delay, and signal occlusion. Superimposing these residual values onto the original pose observation data yields the pose correction data. Because the numerical amplitude of the residual regression target is relatively small, and the model establishes a systematic understanding of noise patterns through spatiotemporal features, the determined pose correction data can effectively eliminate non-physical noise jumps while preserving realistic motion details, thus obtaining a correction output that combines high accuracy and high smoothness.
[0093] Step S308: Based on the pose correction data, perform pose correction on the target object in the video frame sequence.
[0094] In this step, the pose correction data determined in the previous step is applied to update the actual pose of the target object in the video frame sequence, completing the transformation from the original noisy pose to the corrected smooth pose. The corrected pose data not only meets the positional accuracy requirements but also possesses good temporal smoothness and physical plausibility. The corrected pose data will be output to the rendering workstation as control commands to drive the motion of the 3D model in the virtual scene.
[0095] It should be noted that if the pose observation data sequence of the target object is a three-dimensional spatial coordinate sequence, then the pose correction data can be the corrected three-dimensional spatial coordinates. Based on the corrected three-dimensional spatial coordinates, multiple rotation angles can be calculated to obtain the corresponding pose data, and then the pose correction of the target object can be performed based on this. If the pose observation data sequence of the target object is a six-degree-of-freedom pose data sequence, then the pose correction data can be the corrected six-degree-of-freedom pose, which can be directly used for the pose correction of the target object.
[0096] Based on this corrected pose data, the 3D model in the rendering workstation will move along a smooth and continuous trajectory, eliminating inter-frame jitter caused by measurement noise. At the same time, because the residual regression mechanism removes noise while preserving as much high-frequency detail as possible in real motion, natural movement features such as sudden stops and turns of actors or props can also be accurately presented, avoiding the motion lag caused by over-smoothing.
[0097] The following, combined with Figure 3B The above pose correction process will be illustrated by an example scenario.
[0098] This example sets up a virtual shooting task to capture a virtual scene taking place in a fantasy virtual space. In this virtual scene, a 3D model of a "magic box" is suspended in mid-air and needs to move along with a table. This 3D model is synchronized in real-time by a physical prop—a solid wooden cube placed on the tabletop—to achieve the visual effect of the "magic box" moving and rotating in the virtual space as the table moves. This solid wooden cube is the target object in this example. It has eight LED reflective patches on its six faces as markers, numbered p1 to p8. These eight points are not collinear, providing sufficient geometric constraints for the motion capture system to support six-degree-of-freedom pose calculation.
[0099] An array of motion capture cameras was deployed around the filming location. Initially, the prop table and the solid wood cube were stationary. Initial filming was performed using the motion capture camera array, and the reference data for eight LED marker points, including reference coordinates and the reference pose data calculated from them, were obtained based on the video frames obtained from the initial filming.
[0100] During the actual filming, the prop table was pushed to a certain distance from the motion capture camera array. Due to the decrease in signal-to-noise ratio, the 3D spatial coordinates of each reflective patch calculated by the motion capture system showed obvious random fluctuations between adjacent frames. Taking marker point p1 as an example, when the cube was stationary, the difference in its X-axis coordinates between adjacent frames reached approximately 'a' millimeters. When rendered on the rendering workstation, this manifested as a visible tremor in the 3D model of the "magic box".
[0101] In response to the above problems, such as Figure 3BAs shown, the data captured by the motion capture camera array is first processed to obtain a video frame sequence containing a rigid wooden cube. Based on the 3D spatial coordinates of the eight LED markers of the rigid wooden cube in each frame of the video frame sequence, the corresponding pose observation data is calculated. Then, the data is organized using a sliding time window of length M=16, with each frame containing pose observation data for eight markers. The 16 frames of pose observation data form a pose observation data sequence. This pose observation data sequence is normalized to form the final input sequence. Subsequently, this input sequence is fed into a pose correction model for spatiotemporal feature extraction: In the spatial dimension, self-attention processing is used to model the relative geometric relationship between the eight markers. When a marker shifts its coordinates due to brief occlusion, the model can infer its reasonable position using the spatial distribution of the remaining markers. In the temporal dimension, multi-head attention processing is used to establish the dependency relationship between the current frame and the previous 15 frames, identifying abnormal jump frames that do not conform to the uniform translational motion law of the rigid wooden cube.
[0102] Taking pose correction only for the current frame as an example, after obtaining the spatiotemporal features, the coordinate residuals of each marker point are first output through residual regression processing. Taking p1 as an example, its predicted residuals are approximately (-6.8mm, -4.9mm, -4.2mm). These residuals are then superimposed onto the three-dimensional spatial coordinates of the original pose observation data of the solid wood cube rigid body in the current frame. When the actor quickly pushes the solid wood cube rigid body, because this motion conforms to the laws of mechanical inertia, the predicted residuals are small, and the true rapid displacement details of the solid wood cube rigid body are completely preserved, avoiding the sense of motion lag caused by excessive smoothing.
[0103] Furthermore, based on the corrected 3D spatial coordinates of the eight marker points, the 6-DoF pose of the rigid wooden cube is recalculated to obtain the corrected pose correction data. In this example, the pose correction model is set in the motion capture workstation. After obtaining the pose correction data, the motion capture workstation passes it to the rendering workstation. The rendering workstation uses the pose correction data to update the position and orientation of the "Magic Box" 3D model, driving the "Magic Box" 3D model in the virtual scene to move synchronously, thus presenting the effect of the "Magic Box" moving with the table in the final synthesized video frame image. As a result, in the static segments of the rigid wooden cube, the "Magic Box" 3D model no longer exhibits trembling, and in the moving segments, the trajectory of the "Magic Box" 3D model is smooth and continuous, effectively improving the coherence of the virtual-real composite image.
[0104] This embodiment uses multiple marker points set on a target object as a basis to extract spatiotemporal features from the pose observation data sequence of the target object in a video frame sequence containing the target object. This allows for the modeling of spatial and temporal relationships between the multiple marker points, reflecting both their spatial positional relationships and dynamic changes. Based on this, pose correction data corresponding to the multiple marker points is determined, and the target object's pose is corrected accordingly. This results in a more accurate and smoother pose for the corrected target object, effectively suppressing trajectory jitter noise introduced by motion capture technology in large-scale scenes due to increased positioning distance. This improves the rendering stability of the target object in the video frame sequence and the coherence of the virtual-real composite image, thereby enhancing the overall quality of the video obtained from virtual shooting.
[0105] Furthermore, embodiments of this application also provide an electronic device. (Refer to...) Figure 4 The diagram shows a structural schematic of an electronic device according to Embodiment 5 of this application. The electronic device can be implemented as a workstation of a virtual shooting system or as other devices capable of executing the methods in any of the foregoing method embodiments. The specific embodiments of this application do not limit the specific implementation of the electronic device.
[0106] like Figure 4 As shown, the electronic device may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.
[0107] in: The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408.
[0108] Communication interface 404 is used to communicate with other electronic devices or servers.
[0109] The processor 402 is used to execute program 410, specifically to perform the relevant steps in any of the above method embodiments.
[0110] Specifically, program 410 may include program code that includes computer operation instructions.
[0111] Processor 402 may be a CPU, a GPU (Graphics Processing Unit), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The electronic device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0112] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0113] Program 410 may include multiple computer instructions. Specifically, program 410 may use multiple computer instructions to cause processor 402 to perform the operation corresponding to any of the methods described in the foregoing multiple method embodiments.
[0114] The specific implementation of each step in procedure 410 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0115] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in any of the foregoing method embodiments. The computer storage medium includes, but is not limited to, compact disc read-only memory (CD-ROM), random access memory (RAM), floppy disk, hard disk, or magneto-optical disk.
[0116] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the methods in the above-described multiple method embodiments.
[0117] Furthermore, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used for training the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0118] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.
[0119] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., Random Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0120] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0121] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.
Claims
1. A pose correction method, comprising: Acquire a video frame sequence containing a target object and a pose observation data sequence of the target object, wherein multiple marker points for pose estimation are set on the target object; Based on the pose observation data sequence, spatiotemporal features are extracted for the multiple marker points to obtain the corresponding spatiotemporal features; Based on the spatiotemporal characteristics, the pose correction data corresponding to the multiple marker points are determined; Based on the pose correction data, pose correction is performed on the target object in the video frame sequence.
2. The method according to claim 1, wherein, The step of extracting spatiotemporal features from the multiple marker points based on the pose observation data sequence to obtain the corresponding spatiotemporal features includes: Based on the pose observation data sequence, feature extraction is performed on the relative geometric structure relationship of the multiple marker points to obtain the corresponding spatial features; Furthermore, based on the pose observation data sequence, feature extraction is performed on the adjacent frame dependencies of the multiple marker points to obtain the corresponding temporal features; The spatiotemporal features are obtained based on the spatial features and the temporal features.
3. The method according to claim 2, wherein, The step of extracting features of the relative geometric structure relationship of the multiple marker points based on the pose observation data sequence to obtain corresponding spatial features includes: extracting features of the relative geometric structure relationship of the multiple marker points based on the pose observation data sequence through self-attention processing to obtain corresponding spatial features. And / or, The step of extracting features of the adjacent frame dependencies of the multiple marker points based on the pose observation data sequence to obtain corresponding temporal features includes: extracting features of the adjacent frame dependencies of the multiple marker points based on the pose observation data sequence through multi-head attention processing to obtain corresponding temporal features.
4. The method according to any one of claims 1-3, wherein, The step of determining the pose correction data corresponding to the plurality of marker points based on the spatiotemporal features includes: Based on the spatiotemporal characteristics, residual regression correction is performed on the multiple marker points to obtain pose correction data corresponding to the multiple marker points.
5. The method according to claim 4, wherein, The step of performing residual regression correction on the multiple marker points based on the spatiotemporal features to obtain pose correction data corresponding to the multiple marker points includes: Based on the spatiotemporal characteristics of the multiple marker points and the reference characteristics corresponding to the multiple marker points at the initial reference position, residual values are obtained; Based on the residual values, the pose observation data of the multiple marker points are corrected to obtain the pose correction data corresponding to the multiple marker points.
6. A method for training a pose correction model, comprising: Acquire video frame sequence samples and pose observation data sequence samples of sample objects contained in the video frame sequence samples, wherein multiple marker points for pose estimation are set on the sample objects; Using the pose correction model to be trained, spatiotemporal features are extracted for the multiple marker points based on the pose observation data sequence samples to obtain corresponding spatiotemporal feature samples; based on the spatiotemporal feature samples, pose correction data samples corresponding to the multiple marker points are determined. The pose correction model is trained based on the pose correction data samples and the preset loss function.
7. The method according to claim 6, wherein, The step of extracting spatiotemporal features from the multiple marker points based on the pose observation data sequence samples to obtain corresponding spatiotemporal feature samples includes: Based on the pose observation data sequence samples, feature extraction is performed on the relative geometric structure relationship of the multiple marker points to obtain the corresponding spatial feature samples; Furthermore, based on the pose observation data sequence samples, feature extraction is performed on the adjacent frame dependencies of the multiple marker points to obtain the corresponding temporal feature samples; Based on the spatial feature samples and the temporal feature samples, the spatiotemporal feature samples are obtained.
8. The method according to claim 6 or 7, wherein, The step of determining the pose correction data samples corresponding to the plurality of marker points based on the spatiotemporal feature samples includes: Based on the spatiotemporal feature samples, residual regression correction is performed on the multiple marker points to obtain pose correction data samples corresponding to the multiple marker points.
9. The method according to claim 8, wherein, The step of performing residual regression correction on the plurality of marker points based on the spatiotemporal feature samples to obtain pose correction data samples corresponding to the plurality of marker points includes: Based on the spatiotemporal feature samples of the multiple marker points and the reference feature samples corresponding to the multiple marker points at the initial reference position, residual values are obtained; Based on the residual values, the pose observation data samples of the multiple marker points are corrected to obtain the pose correction data samples corresponding to the multiple marker points.
10. The method according to claim 6, wherein, The step of training the pose correction model based on the pose correction data samples and a preset loss function includes: Based on the pose correction data samples of the multiple marker points and the true pose values corresponding to the multiple marker points, the pose correction model is trained using a preset loss function. The preset loss function includes: a first sub-loss function for evaluating the positional accuracy loss of the multiple marker points, a second sub-loss function for evaluating the velocity smoothing loss of the multiple marker points, and a third sub-loss function for evaluating the acceleration smoothing loss of the multiple marker points.
11. An electronic device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform an operation corresponding to the method as described in any one of claims 1-5; or to perform an operation corresponding to the method as described in any one of claims 6-10.
12. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-5; or implements the method as described in any one of claims 6-10.
13. A computer program product comprising computer instructions that instruct a computing device to perform an operation corresponding to any one of the methods described in claims 1-5; or, to perform an operation corresponding to any one of the methods described in claims 6-10.