Object attitude smoothing method and storage medium

By cascading spatiotemporal feature extraction and smoothing of object pose information in target video frames and their reference frames, the problem of poor pose consistency in object pose estimation is solved, and smoother and more consistent object pose information is achieved.

CN119941542APending Publication Date: 2025-05-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311456155.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The object pose estimation obtains poor coherence of the object pose, and it is necessary to provide an object pose smoothing scheme to improve the coherence of the pose.

Method used

By obtaining object pose information in the target video frame and its forward and backward reference frames, cascading spatiotemporal feature extraction with a preset number of times is obtained to obtain space-time fusion features, and the pose sequence is smoothed based on this.

Benefits of technology

The smooth processing of object pose information in the target video frame is realized, and the coherence and real-timeness of the pose are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941542A_ABST
    Figure CN119941542A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an object attitude smoothing method and a storage medium, and relates to the technical field of data processing, the method comprises the following steps: obtaining attitude information of an object in a target video frame, forward reference frames of the target video frame and backward reference frames of the target video frame, the number of the forward reference frames being greater than the number of the backward reference frames; a preset number of cascaded spatial-temporal feature extraction is carried out on the attitude sequence including the obtained attitude information to obtain spatial-temporal fusion features of the attitude sequence, and the spatial-temporal feature extraction comprises the steps of carrying out feature extraction in a spatial dimension and carrying out feature extraction in a time dimension according to a feature extraction result of the spatial dimension; carrying out smoothing processing on the attitude sequence based on the space-time fusion features; and on the basis of the attitude sequence after smoothing processing, attitude information obtained after smoothing processing is carried out on the attitude information of the object in the target video frame is obtained. By applying the scheme provided by the embodiment of the invention, the coherence of the estimated object attitude can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an object posture smoothing method and a storage medium. Background Art

[0002] Object pose estimation is an important topic in the field of computer vision and is also the basis for many specific downstream applications. Taking the application of object pose estimation in VR (Virtual Reality) game scenes as an example, electronic devices can estimate the player's human body pose, and drive the virtual character to make corresponding actions based on the estimated human body pose, thereby realizing various interactions in the game.

[0003] In the related art, electronic devices generally collect object videos and estimate the object posture for each video frame one by one.

[0004] However, the consistency of the object pose estimated is poor.

[0005] In view of the above situation, it is necessary to provide an object posture smoothing solution to smooth the estimated object posture and improve the consistency of the object posture. Summary of the invention

[0006] The purpose of the embodiment of the present application is to provide an object posture smoothing method and a storage medium to improve the consistency of the estimated object posture. The specific technical solution is as follows:

[0007] In a first aspect, an embodiment of the present application provides an object posture smoothing method, comprising:

[0008] Obtaining posture information of an object in a target video frame, a forward reference frame of the target video frame, and a backward reference frame of the target video frame, wherein the number of the forward reference frames is greater than the number of the backward reference frames;

[0009] Performing a preset number of cascaded spatiotemporal feature extractions on a posture sequence including the obtained posture information to obtain a spatiotemporal fusion feature of the posture sequence, wherein the spatiotemporal feature extraction includes: performing feature extraction in a spatial dimension, and performing feature extraction in a temporal dimension based on the feature extraction result of the spatial dimension;

[0010] Based on the spatiotemporal fusion features, smoothing the posture sequence;

[0011] Based on the smoothed posture sequence, the posture information of the object in the target video frame is obtained after smoothing the posture information.

[0012] In a second aspect, an embodiment of the present application provides an object posture smoothing device, comprising:

[0013] A first posture information obtaining module, used to obtain posture information of an object in a target video frame, a forward reference frame of the target video frame, and a backward reference frame of the target video frame, wherein the number of the forward reference frames is greater than the number of the backward reference frames;

[0014] A spatiotemporal feature extraction module is used to perform a preset number of cascade spatiotemporal feature extractions on a posture sequence including the obtained posture information to obtain a spatiotemporal fusion feature of the posture sequence, wherein the spatiotemporal feature extraction includes: performing feature extraction in a spatial dimension, and performing feature extraction in a temporal dimension based on the feature extraction result of the spatial dimension;

[0015] A smoothing processing module, used for smoothing the posture sequence based on the spatiotemporal fusion features;

[0016] The second posture information obtaining module is used to obtain the posture information of the object in the target video frame after smoothing the posture information based on the smoothed posture sequence.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0018] Memory, used to store computer programs;

[0019] The processor is used to implement the method described in the first aspect when executing the program stored in the memory.

[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0021] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect.

[0022] As can be seen from the above, when the scheme provided in the embodiment of the present application is applied to perform object posture smoothing, the posture information of the object in the target video frame, the forward reference frame of the target video frame and the backward reference frame of the target video frame is first obtained, and a preset number of cascaded spatiotemporal feature extractions are performed on the posture sequence including the obtained posture information to obtain the spatiotemporal fusion features of the posture sequence, thereby smoothing the posture sequence based on the spatiotemporal fusion features, and finally, based on the smoothed posture sequence, the posture information of the object in the target video frame after smoothing is obtained, thereby achieving object posture smoothing.

[0023] Among them, for the target video frame, the forward reference frame is a historical video frame, and the backward reference frame is a future video frame. In addition, since both the previous reference frame and the backward reference frame are temporally associated with the target video frame, the previous reference frame, the backward reference frame and the target video frame are also mutually associated in content. In this way, when the target video frame is smoothed based on the posture sequence containing the forward video frame and the backward video frame, the historical video frames and the future video frames associated with the content of the target video frame are comprehensively considered, which is conducive to eliminating the jitter of the posture information of the target video frame based on the association of the content of each video frame, so that the posture information of the target video frame is smoother and more coherent.

[0024] Of course, implementing any product or method of the present application does not necessarily require achieving all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0026] Figure 1 A schematic diagram of a flow chart of a first object posture smoothing method provided in an embodiment of the present application;

[0027] Figure 2 A schematic diagram of an object posture smoothing process is provided for an embodiment of the present application;

[0028] Figure 3 A schematic diagram of a flow chart of a second object posture smoothing method provided in an embodiment of the present application;

[0029] Figure 4 A schematic diagram of the structure of a posture smoothing model provided in an embodiment of the present application;

[0030] Figure 5 A schematic diagram of a data processing flow provided in an embodiment of the present application;

[0031] Figure 6 A schematic diagram of a difference obtaining process provided in an embodiment of the present application;

[0032] Figure 7 A schematic diagram for comparing smoothing effects is provided for an embodiment of the present application;

[0033] Figure 8 A schematic diagram of the structure of an object posture smoothing device provided in an embodiment of the present application;

[0034] Fig. 9A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field based on the present application belong to the scope of protection of the present application.

[0036] First, the execution subject of the solution provided in the embodiment of the present application is described.

[0037] The executor of the solution provided in the embodiment of the present application is: any electronic device with data processing, storage and other functions.

[0038] The object posture smoothing solution provided in the embodiment of the present application is described in detail below.

[0039] See also Figure 1 , is a flow chart of a first object posture smoothing method provided in an embodiment of the present application, the method comprising the following steps S101-S104.

[0040] Step S101: obtaining posture information of an object in a target video frame, a forward reference frame of the target video frame, and a backward reference frame of the target video frame.

[0041] The target video frame can be understood as a video frame for which object posture smoothing is to be performed.

[0042] It should be noted that the electronic device can continuously collect video frames, and each collected video frame can be used as a target video frame. In this way, by executing this solution multiple times, the posture of the object in each video frame collected by the electronic device can be smoothed.

[0043] The forward reference frame is a video frame captured before the target video frame, and the backward reference frame is a video frame captured after the target video frame. In addition, the number of forward reference frames is greater than the number of backward reference frames.

[0044] In one case, the forward reference frame may be a plurality of continuous video frames including video frames adjacent to the target video frame; similarly, the backward reference frame may also be a plurality of continuous video frames including video frames adjacent to the target video frame.

[0045] The embodiment of the present application does not limit the specific number of forward reference frames and the specific number of backward reference frames. It only needs to satisfy that the number of forward reference frames is greater than the number of backward reference frames.

[0046] For example, the number of forward reference frames may be 7, 8, 9, etc., and the number of backward reference frames may be 2, 3, etc.

[0047] In one case, the number of backward reference frames may be 1.

[0048] The backward reference frame is a video frame used for reference when smoothing the object posture of the target video frame, and the backward reference frame is collected after the target video frame. For the target video frame, the backward reference frame is future information. When the number of backward reference frames is 1, the future information referenced when smoothing the object posture of the target video frame is minimized, which greatly reduces the future information required for reference when smoothing the object posture of the target video frame, thereby greatly improving the real-time performance of the object posture smoothing solution.

[0049] The objects in the target video frame, the forward reference frame, and the backward reference frame may be people, animals, robots, and the like.

[0050] The embodiment of the present application does not limit the specific method of obtaining the posture information of the object in the video frame, which is introduced below by way of example.

[0051] Specifically, object detection may be performed on the video frame first, and then an object pose estimation algorithm may be used to estimate the pose of the detected object to obtain pose information.

[0052] For example, when the object is a human body, object detection algorithms such as the YOLOv5 algorithm and the Fast R-CNN (Fast Region-Convolutional Neural Network) algorithm can be used to perform object detection on the video frame, and then human posture estimation algorithms such as the DeepPose algorithm and the PARE algorithm can be used to perform posture estimation to obtain posture information.

[0053] Depending on the adopted posture estimation algorithm, the above-mentioned posture information may include posture parameters and / or shape parameters.

[0054] In one embodiment of the present application, the above-mentioned posture information may be posture information that satisfies a preset format.

[0055] The above-mentioned preset format may be a parameter format required by a specific human body model.

[0056] For example, for the SMPL (Skinned Multi-Person Linear) model, the required parameter format is: posture parameter θ∈R 24×3 , and the shape parameter β∈R 10 , R represents vector space.

[0057] The SMPL model is constructed by using the posture parameter θ∈R 24×3 Represents human posture through shape parameters β∈R 10 To represent the shape of the human body, based on the above posture parameters and shape parameters, the SMPL model can output a three-dimensional representation mesh M(θ,β)∈R^(6890×3) with posture, and then combined with the position of the joint points, it can render a visual complete object posture.

[0058] The above joint point position J 3D It can be obtained by the following expression:

[0059] J 3D =WM∈R J×3

[0060] Among them, the above W represents the pre-trained linear regressor, M represents the above three-dimensional representation grid, and J represents the number of joint points.

[0061] Step S102: performing a preset number of cascaded spatiotemporal feature extractions on the posture sequence including the obtained posture information to obtain spatiotemporal fusion features of the posture sequence.

[0062] The above-mentioned spatiotemporal feature extraction includes: performing feature extraction in the spatial dimension, and performing feature extraction in the temporal dimension based on the feature extraction result of the spatial dimension.

[0063] Specifically, the following method can be used to perform cascaded spatiotemporal feature extraction to obtain spatiotemporal fusion features of the posture sequence.

[0064] In one implementation, spatial features may be first extracted from each posture information in the target sequence in the spatial dimension, the obtained spatial features may be fused to obtain spatial fusion features, and then feature extraction may be performed on the spatial fusion features in the time dimension to obtain spatiotemporal fusion features, and the target sequence may be updated to the spatiotemporal fusion features, and the step of spatial feature extraction may be returned until the preset number of spatiotemporal feature extractions is reached. For details on the specific implementation, please refer to the following Figure 3 Steps S302 to S305 in the illustrated embodiment are not described in detail here.

[0065] In another implementation, the above posture sequence can be input into a pre-trained posture smoothing model to obtain the above spatiotemporal fusion features. The specific implementation will be given in the subsequent detailed introduction of the posture smoothing model, which will not be described in detail here.

[0066] Step S103: Based on the spatiotemporal fusion features, the posture sequence is smoothed.

[0067] The above-mentioned spatiotemporal fusion features are: features of a posture sequence including posture information of objects in target video frames, forward reference frames and backward reference frames, and the above-mentioned spatiotemporal fusion features include both features of the posture sequence in the time dimension and features of the posture sequence in the space dimension.

[0068] In this way, based on the above-mentioned spatiotemporal fusion features, the posture sequence can be smoothed in a manner of reducing the jitter of the posture information.

[0069] Specifically, the posture sequence may be smoothed in the following manner.

[0070] In one implementation, the spatiotemporal fusion features can be input into the regression head (RH) layer in the posture smoothing model, so that the regression head layer performs smoothing processing on the posture sequence. The specific implementation will be given in the subsequent detailed introduction of the posture smoothing model, which will not be described in detail here.

[0071] In another embodiment, the adjustment parameters corresponding to each posture information in the posture sequence can be determined based on the spatiotemporal fusion features, and the posture sequence can be smoothed using the adjustment parameters. The above adjustment parameters can be obtained based on the correspondence between the preset spatiotemporal fusion features and the adjustment parameters, which will not be described in detail here.

[0072] Step S104: based on the smoothed posture sequence, obtaining the posture information of the object in the target video frame after smoothing.

[0073] After the above step S103, a smoothed posture sequence is obtained. At this time, the smoothed posture information corresponding to the target video frame can be determined from the smoothed posture sequence.

[0074] Specifically, the penultimate target number of gesture information may be determined from the smoothed gesture sequence, and the determined gesture information is the gesture information after the gesture information of the object in the target video frame is smoothed.

[0075] The above target number is: the number of backward reference frames plus 1.

[0076] For example, if the number of backward reference frames is 1, the second to last posture information can be determined from the smoothed posture sequence as the smoothed posture information corresponding to the target video frame.

[0077] As can be seen from the above, when the scheme provided in the embodiment of the present application is applied to perform object posture smoothing, the posture information of the object in the target video frame, the forward reference frame of the target video frame and the backward reference frame of the target video frame is first obtained, and a preset number of cascaded spatiotemporal feature extractions are performed on the posture sequence including the obtained posture information to obtain the spatiotemporal fusion features of the posture sequence, thereby smoothing the posture sequence based on the spatiotemporal fusion features, and finally, based on the smoothed posture sequence, the posture information of the object in the target video frame after smoothing is obtained, thereby achieving object posture smoothing.

[0078] Among them, for the target video frame, the forward reference frame is a historical video frame, and the backward reference frame is a future video frame. In addition, since both the previous reference frame and the backward reference frame are temporally associated with the target video frame, the previous reference frame, the backward reference frame and the target video frame are also mutually associated in content. In this way, when the target video frame is smoothed based on the posture sequence containing the forward video frame and the backward video frame, the historical video frames and the future video frames associated with the content of the target video frame are comprehensively considered, which is conducive to eliminating the jitter of the posture information of the target video frame based on the association of the content of each video frame, so that the posture information of the target video frame is smoother and more coherent.

[0079] In addition, the number of backward reference frames is smaller than the number of forward reference frames, that is, when smoothing the object posture of the target video frame, the future information used is less than the historical information. This ensures that less future information is used to smooth the object posture of the target video frame, thereby improving the real-time performance of the object posture smoothing scheme.

[0080] Furthermore, when extracting features from the posture sequence, multiple cascade feature extractions are performed in the time dimension and the space dimension, so that deep features of the posture sequence can be extracted in the time dimension and the space dimension, making the extracted features richer and more comprehensive, and improving the smoothing effect when smoothing the object posture of the target video frame based on the extracted features.

[0081] Combine the following Figure 2 , a more intuitive explanation of the object posture smoothing solution provided in the embodiment of the present application is provided.

[0082] See also Figure 2 , which is a schematic diagram of an object posture smoothing process provided in an embodiment of the present application.

[0083] Depend on Figure 2 It can be seen that the number of forward reference frames of the target video frame is 5, and the number of backward reference frames is 1.

[0084] After smoothing the posture sequence including the posture information of the target reference frame, the posture information of the forward reference frame and the posture information of the backward reference frame, a smoothed posture sequence is obtained. In this way, the smoothed posture information corresponding to the target video frame can be determined from the smoothed posture sequence as the posture information obtained after the object posture smoothing is performed on the target video frame.

[0085] exist Figure 1 On the basis of the embodiment shown, when performing spatiotemporal feature extraction, spatial feature extraction can be performed on each posture information in the target sequence in the spatial dimension, and the obtained spatial features can be fused to obtain spatial fusion features, and then feature extraction can be performed on the spatial fusion features in the time dimension to obtain spatiotemporal fusion features, and the target sequence can be updated to the spatiotemporal fusion features, and the step of performing spatial feature extraction can be returned until the preset number of spatiotemporal feature extractions is reached. In view of the above situation, the embodiment of the present application provides a second object posture smoothing method.

[0086] See also Figure 3 , is a flow chart of a second object posture smoothing method provided in an embodiment of the present application, the method includes the following steps S301-S307.

[0087] Step S301: obtaining posture information of an object in a target video frame, a forward reference frame of the target video frame, and a backward reference frame of the target video frame.

[0088] The above step S301 is the same as the above Figure 1 Step S101 in the illustrated embodiment is the same and will not be described again here.

[0089] Step S302: extracting spatial features of each posture information in the target sequence in the spatial dimension to obtain spatial features of each posture information.

[0090] The initial value of the target sequence is a posture sequence including the acquired posture information.

[0091] Each posture information in the target sequence is the posture information of the object in a video frame. Therefore, performing spatial feature extraction on each posture information can be understood as: capturing the characteristics of each individual posture information itself.

[0092] Specifically, any spatial feature extraction method can be used to extract spatial features from each posture information, such as using an MLP (Multi-Layer Perceptron) algorithm to extract features from posture information, using a deep connection network to extract features from posture information, etc., and the extracted features are used as the spatial features of each posture information.

[0093] In one case, spatial features of each posture information may be extracted based on a spatial encoder (Spatial Transformer Encoder) in a spatiotemporal transformer module (STM) layer of an object posture smoothing model, as described in subsequent embodiments.

[0094] In one embodiment of the present application, the obtained posture information includes: joint point position information of the object.

[0095] In this case, different methods can be used to extract the spatial features of the posture information according to different times of spatial feature extraction, which are explained below respectively.

[0096] In the first case, if feature extraction is performed in the spatial dimension for the first time, the following steps A to C can be used to perform spatial feature extraction.

[0097] Step A: Based on the joint point position information included in each posture information in the target sequence, the first position code of each joint point corresponding to each posture information is obtained.

[0098] Each piece of posture information may include multiple joint position information, for example, may include 24 joint position information.

[0099] Specifically, the position information of each joint point corresponding to each posture information may be individually position-coded, and the coding result may be used as the first position code of each joint point corresponding to each posture information.

[0100] Step B: superimposing the first position code of each joint point corresponding to each posture information on each posture information.

[0101] After the information is superimposed in this way, the position code of the joint points is integrated into the posture information, so that when the posture information is subsequently spatially extracted, the spatial features representing the position information of the joint points can be extracted.

[0102] Step C: extracting spatial features of the posture information after superimposing the first position code in the spatial dimension to obtain spatial features of each posture information.

[0103] The specific method of performing spatial feature extraction in step C may be similar to the spatial feature extraction method described above, with the only difference being that the posture information is replaced by the posture information after superimposing the first position encoding, which will not be described in detail here.

[0104] In the second case, if it is not the first time to extract features in the spatial dimension, spatial features can be directly extracted for each posture information in the target sequence in the spatial dimension to obtain the spatial features of each posture information.

[0105] The specific implementation method of this situation can be the same as the spatial feature extraction method introduced above, and will not be repeated here.

[0106] As can be seen from the above, if the posture information includes the position information of the joint points of the object, then when the feature extraction is performed in the spatial dimension for the first time, the first position code of each joint point corresponding to each posture information can be obtained based on the position information of the joint points included in each posture information, and the first position code of each joint point corresponding to the posture information can be superimposed on each posture information, and then the spatial feature extraction of the posture information after the superimposed first position code is performed in the spatial dimension to obtain the spatial features of each posture information. Among them, the first position code is obtained based on the joint point position information, which can reflect the posture of the object in a more fine-grained manner. In this way, when obtaining the spatial features of each posture information, not only the overall features of the posture information are considered, but also the first position code that can reflect the posture of the object in a more fine-grained manner is considered, thereby improving the accuracy of the spatial features finally obtained.

[0107] Step S303: Fuse the spatial features of each posture information to obtain the spatial fusion features of the target sequence.

[0108] After obtaining the spatial features of each posture information, the obtained spatial features can be fused using a variety of feature fusion methods such as feature connection and feature weighted fusion, which is not limited in this application.

[0109] At this time, the obtained spatial fusion features reflect the overall spatial characteristics of each posture information obtained.

[0110] Step S304: extracting time features from the spatial fusion features in the time dimension to obtain spatiotemporal fusion features.

[0111] The above-mentioned spatial fusion features reflect the overall spatial features of each posture information obtained. In this step, time features are extracted from the spatial fusion features in the time dimension, which can be understood as: capturing the correlation features between each posture information.

[0112] Specifically, any temporal feature extraction method can be used to extract temporal features from spatial fusion features, such as using an MLP algorithm to extract spatial fusion features, using a deep connection network to extract spatial fusion features, etc., and the extracted features are used as spatiotemporal fusion features.

[0113] Among them, when the spatial fusion feature is obtained based on the above-mentioned first position coding, this step extracts the time feature of the spatial fusion feature in the time dimension, which is equivalent to extracting the time feature of the spatial feature with the position coding added.

[0114] In one case, temporal features may be extracted from spatial fusion features based on a temporal encoder (Temporal Transformer Encoder) in a spatiotemporal cyclic coding layer of an object posture smoothing model, as described in subsequent embodiments.

[0115] If the obtained posture information includes the position information of the joint points of the object, then, according to the different times of temporal feature extraction, different methods can be used to extract spatial fusion features for temporal features, which are explained below respectively.

[0116] In the first case, if feature extraction is performed in the time dimension for the first time, the following steps D-F can be used to extract spatial features.

[0117] Step D: Based on the joint point position information included in each posture information, obtain the second position codes of all joint points corresponding to each posture information.

[0118] Specifically, the position information of all joint points corresponding to each posture information may be coded as a whole, and the coding result may be used as the second position code of all joint points corresponding to each posture information.

[0119] Step E: Superimpose the obtained second position code on the spatial fusion feature.

[0120] After the position codes are superimposed in this way, the position codes of all joint points are integrated into the spatial fusion features, so that when the superimposed spatial fusion features are subsequently spatially extracted, the spatiotemporal fusion features representing the position information of all joint points can be extracted.

[0121] Step F: extracting the temporal features of the spatial fusion features after superimposing the second position code in the temporal dimension to obtain the spatiotemporal fusion features.

[0122] The feature extraction method of step D can be similar to the temporal feature extraction method described above. The only difference is that the spatial fusion feature is replaced by the spatial fusion feature after superimposing the second position encoding, which will not be repeated here.

[0123] In the second case, if it is not the first time to extract features in the time dimension, then time features can be directly extracted from the spatial fusion features in the time dimension to obtain spatiotemporal fusion features.

[0124] The specific implementation method of this situation can be the same as the time feature extraction method introduced above, and will not be repeated here.

[0125] As can be seen from the above, if the posture information includes the position information of the joint points of the object, then when the feature extraction is performed in the time dimension for the first time, the second position code of all the joint points corresponding to each posture information can be obtained based on the position information of the joint points included in each posture information, and the obtained second position code is superimposed on the spatial fusion feature, and the spatial fusion feature after superimposing the second position code is subjected to time feature extraction in the time dimension to obtain the spatiotemporal fusion feature. Among them, the second position code is obtained based on the joint point position information, which can reflect the posture of the object in a more fine-grained manner. In this way, when obtaining the spatiotemporal fusion feature, not only the overall spatial fusion feature of the posture information is considered, but also the second position code that can reflect the posture of the object in a more fine-grained manner is considered, thereby improving the accuracy of the finally obtained spatiotemporal fusion feature.

[0126] Step S305: Determine whether the number of spatiotemporal feature extractions is less than a preset number. If yes, update the target sequence to the spatiotemporal fusion feature and return to step S302. If no, execute step S306.

[0127] When the number of spatiotemporal feature extractions is less than the preset number, the process can return to step S302 and perform the next round of spatiotemporal feature extraction again. In this way, after the preset number of cascaded spatiotemporal feature extractions, the deep spatiotemporal fusion features of the posture sequence can be obtained.

[0128] Step S306: Based on the spatiotemporal fusion features, the posture sequence is smoothed.

[0129] Step S307: based on the smoothed posture sequence, obtaining the posture information of the object in the target video frame after smoothing.

[0130] The above steps S306 and S307 are the same as the above steps. Figure 1 In the illustrated embodiment, step S103 is the same as step S104 and will not be described in detail herein.

[0131] As can be seen from the above, in this embodiment, spatial features are first extracted for each posture information in the target sequence in the spatial dimension, and the obtained spatial features are fused to obtain spatial fusion features, which reflect the overall spatial features of each posture information obtained; then, feature extraction is performed on the above spatial fusion features in the time dimension to obtain spatiotemporal fusion features, which reflect both the characteristics of the posture sequence in the spatial dimension and the characteristics of the posture sequence in the time dimension; and the target sequence is updated to the spatiotemporal fusion features, and the step of spatial feature extraction is returned until the preset number of spatiotemporal feature extractions is reached. It can be seen that through the above-mentioned hierarchical feature extraction method, the deep spatiotemporal fusion features of the posture sequence can be finally obtained, which is conducive to improving the smoothing effect of the subsequent object posture smoothing based on the spatiotemporal fusion features.

[0132] The following is a detailed introduction to the posture smoothing model mentioned above.

[0133] The above-mentioned posture smoothing model may include N layers of spatiotemporal cyclic coding layers and one regression head layer, and the spatiotemporal cyclic coding layer includes a spatial decoder and a temporal decoder.

[0134] In one embodiment of the present application, the above-mentioned posture smoothing model may further include a linear embedding (LE) layer in addition to the spatiotemporal cyclic coding layer and the regression head layer.

[0135] See also Figure 4 , is a structural diagram of a posture smoothing model provided in an embodiment of the present application.

[0136] It can be seen that the above-mentioned posture smoothing model includes a linear embedding layer, N spatiotemporal cyclic coding layers of spatiotemporal cyclic coding layer 1-spatial cyclic coding layer N, and a regression head layer.

[0137] The above LE layer can also be called a fully connected layer, which is used to map the posture sequence to a high-dimensional space.

[0138] Specifically, for the posture sequence X, the LE layer can use the following expression to perform high-dimensional mapping on X to obtain the initial feature Z 0 :

[0139] Z 0 =LE(X)

[0140] Among them, LE() represents a high-dimensional mapping operation, Z 0 ∈R F×J×C , R represents the vector space, F represents the total number of video frames, J represents the number of joint points corresponding to the posture information, and C represents the feature dimension after high-dimensional mapping, such as 32 dimensions.

[0141] The above-mentioned spatial encoder and temporal encoder are introduced below.

[0142] The encoder includes a multi-head self-attention (MSA) module, a multi-layer perceptron (MLP) module, a layer normalization module, and a residual connection module.

[0143] Each head in the above multi-head self-attention (MSA) module is used to linearly map the input pose sequence to obtain the query matrix Q, key matrix K and value matrix V, and then calculate the scaled dot product attention head of the i-th head by the following formula i :

[0144]

[0145] in, denote the query matrix, key matrix, and value matrix obtained by the i-th head, respectively. N denotes the number of elements in the pose sequence. m Represents the dimension of each element, Attention() represents the operation of scaling and clicking attention, and Softmax() represents the normalized exponential operation function.

[0146] Finally, the scaled dot product attention obtained from the h heads is connected according to the following expression to obtain the connected attention MSA:

[0147] MSA=Concat(head 1 ,…,head h )W o

[0148] Among them, Concat() represents the connection processing function, head 1 ,…,head h represents the scaled dot product attention obtained by h heads, Represents the linear mapping weight.

[0149] The above MLP module consists of two linear layers, which are used to transform the features represented by the following expression:

[0150] MLP(x)=σ(xW 1 +b 1 )W 2 +b 2 ,

[0151] Among them, x represents the input feature, σ represents the GELU activation function, and W 1 and W 2 Represent the weights of the two linear layers, b 1 and b 2 Represents the bias term.

[0152] It can be seen that based on the above modules, the encoder can use the self-attention mechanism to perform deep feature extraction on the input feature sequence.

[0153] In one embodiment of the present application, when cascading spatiotemporal features of a posture sequence to obtain spatiotemporal fusion features, the above posture sequence can be input into the first spatiotemporal cyclic coding layer in the posture smoothing model, and each cascaded spatiotemporal cyclic coding layer performs spatiotemporal feature extraction in turn to obtain the spatiotemporal fusion features of the posture sequence output by the last spatiotemporal cyclic coding layer.

[0154] Specifically, the above-mentioned spatiotemporal cyclic coding layer includes a spatial encoder and a temporal encoder. The spatial encoder is used to extract features in the spatial dimension, and the temporal encoder is used to extract features in the temporal dimension for the output result of the spatial encoder.

[0155] The processing of the posture sequence by the spatiotemporal cyclic coding layer is described in detail below.

[0156] If the posture sequence input to the first layer STM is called the initial feature Z 0 , then, for the l-th layer STM, its input feature can be called Z l-1 , the output feature can be called Z l .

[0157] For example, the input feature of the first layer STM is Z 0 , the output feature is Z 1 , the input feature of the second layer STM is Z 1 , the output feature is Z 2 .

[0158] The data processing processes of the first layer STM and the lth layer STM other than the first layer are described below respectively.

[0159] For Tier 1 STM:

[0160] First, E SPos and Fusion is performed to obtain the updated position encoding containing the joint points The updated Input the spatial encoder in the first layer STM to extract spatial features and obtain spatial features

[0161] Among them, E SPos ∈R J×C , represents the initial feature Z 0 The position encoding of each joint point of the object described by the posture feature corresponding to each video frame; f∈F, represents the initial feature Z 0 The pose features corresponding to each video frame.

[0162] Then, the spatial features corresponding to each video frame are fused Get the spatial fusion features of each video frame extracted by the first layer STM ζZ 1 ∈R F×(J×C) , and for E TPos With ζZ 1 Fusion is performed to obtain the updated ζZ containing the position encoding of the joint points 1 , where R represents the vector space, J represents the number of joint points corresponding to the posture information, C represents the feature dimension, and F represents the total number of video frames. They respectively represent the spatial features corresponding to F video frames extracted by the first layer of STM.

[0163] Among them, l TPos ∈E F×(J×C) , represents the initial feature Z 0 The position encoding of all joint points of the object described by the posture features corresponding to each video frame.

[0164] Finally, the updated ζZ 1 Input the time encoder in the first layer STM to extract the time feature and obtain the output feature Z of the first layer STM 1 It can be seen that the output feature Z 1 It is a feature of space-time fusion.

[0165] For layer l STM except layer 1:

[0166] First, directly Input into the spatial encoder in the lth layer STM to extract spatial features and obtain spatial features

[0167] in, f∈F, represents the input feature Z l-1 The pose features corresponding to each video frame.

[0168] Then, the spatial features corresponding to each video frame are fused Get the spatial fusion features of each video frame extracted by the lth layer STM ζZ l ∈R F×(J×C) Where R represents the vector space, J represents the number of joint points corresponding to the posture information, C represents the feature dimension, and F represents the total number of video frames. They respectively represent the spatial features corresponding to F video frames extracted by the l-th layer STM.

[0169] Finally, directly convert ζZ l Input the time encoder in the l-th layer STM to extract the time feature and obtain the output feature Z of the l-th layer STM l It can be seen that the output feature Z l It is a feature of space-time fusion.

[0170] In this way, L STMs can fully extract the spatiotemporal features of the posture sequence, and the output Z of the last layer of STM is L This is the final spatiotemporal fusion feature.

[0171] In one embodiment of the present application, when smoothing the posture sequence based on the spatiotemporal fusion features, the spatiotemporal fusion features can be input into the regression head layer to smooth the posture sequence to obtain a smooth posture sequence.

[0172] Specifically, the regression head layer can simultaneously predict the smooth posture information of all video frames based on the spatiotemporal fusion features and the posture sequence in a sequence-to-sequence (seq2seq) manner to obtain a smooth posture sequence.

[0173] It can be seen that through the cascade feature extraction of multiple spatiotemporal cyclic coding layers in the posture smoothing model, the spatiotemporal fusion features of the posture sequence can be quickly and comprehensively extracted. Then, the above-mentioned spatiotemporal fusion features are input into the regression head layer, which can quickly and accurately smooth the posture sequence to obtain a smooth posture sequence.

[0174] Next, combine Figure 5 , which provides a more intuitive explanation of the data processing flow of the first layer of STM.

[0175] See also Figure 5 , which is a schematic diagram of a data processing flow provided in an embodiment of the present application.

[0176] Figure 5 In the above, C represents the posture feature corresponding to each video frame in the initial feature, and pe1-pe24 represent the position encoding of each joint point of the object described by the posture feature corresponding to each video frame in the initial feature, corresponding to the aforementioned E SPos ; J×C represents the spatial feature obtained by the spatial encoder after extracting the spatial feature from the initial feature, PE1-PE9 represent the position encoding of all joint points of the object described by the posture feature of each video frame in the initial feature, corresponding to the aforementioned E TPos ; F×J×C represents the input features of the first layer of STM, that is, the spatiotemporal fusion features.

[0177] It can be seen that for the first layer of STM, when the spatial encoder (Spatial Transformer Encoder) performs feature extraction, it also considers the position encoding of each joint point of the object described by the posture features of each video frame, and when the temporal encoder (Temporal Transformer Encoder) performs feature extraction, it also considers the position encoding of all joint points of the object described by the posture features of each video frame. This can obtain the spatiotemporal fusion features of the posture sequence in a more fine-grained and comprehensive manner.

[0178] The training method of the above-mentioned posture smoothing model is introduced below through steps G to J.

[0179] Step G: Obtain a sample pose sequence and annotated pose sequence corresponding to a sample video frame.

[0180] The sample posture sequence includes posture information obtained by performing posture estimation on objects in the sample video frame, the forward sample reference frame of the sample video frame, and the backward sample reference frame of the sample video frame.

[0181] The above posture information can be combined with the above Figure 1 The manner of obtaining the posture information described in step S101 of the illustrated embodiment is the same and will not be repeated here.

[0182] The annotated posture sequence includes: real posture information of the object in the sample video frame, the forward sample reference frame and the backward sample reference frame, and the number of the forward sample reference frames is greater than the number of the backward sample reference frames.

[0183] The above-mentioned real posture information can be obtained through on-site actual measurement and other methods, which will not be elaborated here.

[0184] Step H: Input the sample posture sequence into a preset model to obtain a smooth posture sequence output by the model after smoothing the sample posture sequence.

[0185] The network architecture of the above preset model can be as mentioned above Figure 4 As shown, no further description is given here.

[0186] Step I: Based on the labeled pose sequence and the smoothed pose sequence, determine the loss value of the model for smoothing.

[0187] In one embodiment of the present application, the loss value can be obtained based on at least one of the following differences:

[0188] The first one is to mark the first posture difference between the object posture represented by the posture sequence and the object posture represented by the smooth posture sequence.

[0189] The first posture difference L 1 It can be expressed as follows:

[0190]

[0191] Among them, λ rot represents the preset rotation loss parameter, λ joint Represents the preset joint point loss parameters, and Y T They represent the above smooth posture sequence and the labeled posture sequence respectively, and They respectively represent the joint point position sequence included in the above-mentioned smooth posture sequence and the joint point position sequence included in the annotated posture sequence.

[0192] In this way, the loss value during model training is obtained based on the first posture difference between the object posture represented by the labeled posture sequence and the object posture represented by the smooth posture sequence, which can constrain the smooth posture sequence output by the model to be close to the labeled posture sequence, which is beneficial to improving the object posture smoothing effect of the model.

[0193] The second type is the speed difference between the change speed of the object posture represented by the labeled posture sequence and the change speed of the object posture represented by the smooth posture sequence.

[0194] The above speed difference L vel It can be expressed as follows:

[0195]

[0196] Among them, λ rot represents the preset rotation loss parameter, λ joint Represents the preset joint point loss parameters, and V(Y T ) represent the rotation change speed of the object in the adjacent video frames described by the smooth posture sequence and the annotated posture sequence, respectively. and They respectively represent the changing speed of the joint point positions of the objects in the adjacent video frames of the above smooth posture sequence and the annotated posture sequence.

[0197] In this way, the loss value during model training is obtained based on the speed difference between the changing speed of the object posture represented by the labeled posture sequence and the changing speed of the object posture represented by the smooth posture sequence. This can constrain the changing speed of the object posture represented by the smooth posture sequence output by the model to be close to the changing speed of the object posture represented by the labeled posture sequence, which is beneficial to improving the object posture smoothing effect of the model.

[0198] The third type, the difference between the second posture difference and the third posture difference, can also be called the queue difference.

[0199] The second posture difference is: the difference between the object postures represented by the annotated posture sequences of adjacent sample video frames, and the third posture difference is: the difference between the object postures represented by the smooth posture sequences of adjacent sample video frames.

[0200] The difference L between the second posture difference and the third posture difference queue It can be expressed as follows:

[0201]

[0202] Among them, λ rot represents the preset rotation loss parameter, λ joint Represents the preset joint point loss parameters, and represents the smoothed pose sequence of two adjacent sample video frames, by the unsmoothed pose sequence X of two adjacent sample video frames T and X T-1 After smoothing, Y T and Y T-1 represents the annotated pose sequence of two adjacent sample video frames, and The smooth posture sequence of two adjacent sample video frames includes the joint point position sequence, and The joint point position sequence included in the labeled posture sequence of two adjacent sample video frames.

[0203] Combine the following Figure 6 , a more intuitive explanation of the above adjacent posture sequence is given.

[0204] Figure 6 Each numbered rectangle in the pose sequence represents a video frame.

[0205] Depend on Figure 6 It can be seen that the annotated posture sequence corresponding to sample video frame 8 contains the posture information corresponding to sample video frames 1-9, corresponding to the aforementioned Y T-1 , the annotated posture sequence corresponding to sample video frame 9 contains the posture information corresponding to sample video frames 2-10, corresponding to the aforementioned Y T , based on Y T and Y T-1 The second posture difference can be obtained.

[0206] The unsmoothed posture sequence corresponding to sample video frame 8 includes posture information corresponding to sample video frames 1-9, corresponding to the aforementioned X T-1 , the unsmoothed posture sequence corresponding to sample video frame 9 contains the posture information corresponding to sample video frames 2-10, corresponding to the aforementioned X T .

[0207] After the model smoothes the unsmoothed posture sequence, the smoothed posture sequence corresponding to sample video frame 8 still contains the posture information corresponding to sample video frames 1-9, corresponding to the above expression The smoothed posture sequence corresponding to sample video frame 9 still contains the posture information corresponding to sample video frames 2-10, corresponding to the above expression based on and The third posture difference can be obtained.

[0208] In this way, the difference between the second posture difference and the third posture difference can be determined.

[0209] It can be seen that the second posture difference actually reflects the difference between the annotated posture sequences corresponding to adjacent sample video frames; the third posture difference actually reflects the queue difference between the smooth posture sequences corresponding to adjacent sample video frames. In this way, the loss value obtained during model training based on the difference between the second posture difference and the third posture difference can constrain the queue difference between the smooth posture sequences obtained by the model for two adjacent sample video frames to be close to the queue difference between the annotated posture sequences, so that in the final reasoning, the posture information of the sample video frames after smoothing can be smoothly connected, reducing the probability of jitter.

[0210] Step J: Based on the loss value, adjust the network parameters of the model to obtain a posture smoothing model.

[0211] It can be seen that in this way, based on the sample posture sequence and the labeled posture sequence corresponding to the sample video frames, a posture smoothing model can be quickly and reasonably trained.

[0212] In one embodiment of the present application, the input and output of the above-mentioned posture smoothing model may be standardized.

[0213] Specifically, the posture sequence input to the model is: a 9D (Dimension) rotation matrix that can continuously represent the rotation semantics, and the smooth posture sequence output by the model is: a 6D rotation representation that can continuously represent the rotation semantics.

[0214] Specifically, rotating Euler angles, rotating axis angle coordinates, rotating quaternions, rotating matrices, and 6-dimensional rotation representations can all represent the posture of human joints. However, the inventors have found in practice that rotating Euler angles, rotating axis angle coordinates, and rotating quaternions cannot continuously represent rotation semantics, and as input or output of the model, they will cause data ambiguity, increase the difficulty of model training, and make it difficult to converge.

[0215] First, the above 9D rotation matrix is ​​introduced.

[0216] The 9D rotation matrix represents the posture rotation angle of the joint and can continuously represent the rotation semantics. The rotation matrix is ​​an orthogonal matrix and the determinant is always equal to 1. As the input of the network, it can accelerate the convergence of network training and help find a standardized input data distribution, which is beneficial to the training and reasoning of the network.

[0217] The above-mentioned 6D rotation representation is further introduced.

[0218] In one case, the above 6D rotation representation A can be expressed by the following expression:

[0219] A=[a 1 ,a 2 ]∈R 3×2

[0220] Where R represents the vector space, a 1 ,a 2 Represents the rotation parameters included in the 6D rotation representation.

[0221] make

[0222] b 1 =N(a 1 ), b 2 =N(a 2 -(b 1 ·a 2 )b 1 ), b 3 =b 1 ×b 2

[0223] Among them, b 1 、b 2 、b 3 represents the rotation parameters included in the 9D rotation matrix, and N() represents the normalization process. Then, the 9D rotation matrix B composed of b = [b 1 ,b 2 ,b 3 ]∈R 3×3 That is, it is a standard orthogonal matrix with determinant equal to 1.

[0224] The above A can be called a 6D rotation representation. A can not only continuously represent the rotation semantics, but also derive the above 9D rotation matrix B, which is beneficial to the training and convergence of the model.

[0225] It can be seen that using the 9D rotation matrix that can continuously represent the rotation semantics as the model input and using the 6D rotation representation that can continuously represent the rotation semantics as the model output standardizes the input and output of the model, is conducive to eliminating data ambiguity, and can speed up the reasoning and training of the model, which is conducive to the real-time deployment of this solution in the industrial field and improves the efficiency and stability of the object posture smoothing solution.

[0226] The following compares the posture smoothing effects of the object posture smoothing solution provided in the embodiment of the present application with those of the object posture smoothing solution in the related art.

[0227] See also Figure 7 , which provides a schematic diagram for comparing smoothing effects in an embodiment of the present application.

[0228] Figure 7 A comparison chart of smoothing effects of smoothing an AMASS dataset using the solution provided in the embodiment of the present application and the solution of related technologies.

[0229] Among them, MPJPE represents the average position error of each joint, MPJTE represents the average velocity error of each joint between adjacent frames, PVE represents the error of each vertex, Accel represents the acceleration of the joint in the coordinate system of the image acquisition device, and Noise represents the original noise contained in the unsmoothed data set.

[0230] It can be seen that compared with One-Euro, Sav-Gol, GausslD, SmoothNet and other methods, the solution provided in the embodiment of the present application has achieved the best results in indicators such as MPJPE, MPJTE, and PVE.

[0231] In terms of time consumption, the solution provided in the embodiment of the present application can be run within a few milliseconds, achieving real-time smoothing of the object's posture.

[0232] Corresponding to the above-mentioned object posture smoothing method, an embodiment of the present application provides an object posture smoothing device.

[0233] See also Figure 8 , is a schematic diagram of the structure of an object posture smoothing device provided in an embodiment of the present application, the device includes the following modules:

[0234] A first posture information obtaining module 801 is used to obtain posture information of an object in a target video frame, a forward reference frame of the target video frame, and a backward reference frame of the target video frame, wherein the number of the forward reference frames is greater than the number of the backward reference frames;

[0235] The spatiotemporal feature extraction module 802 is used to perform a preset number of cascade spatiotemporal feature extractions on a posture sequence including the obtained posture information to obtain a spatiotemporal fusion feature of the posture sequence, wherein the spatiotemporal feature extraction includes: performing feature extraction in the spatial dimension, and performing feature extraction in the time dimension based on the feature extraction result of the spatial dimension;

[0236] A smoothing processing module 803, used for smoothing the posture sequence based on the spatiotemporal fusion feature;

[0237] The second posture information obtaining module 804 is used to obtain the posture information of the object in the target video frame after smoothing the posture information based on the smoothed posture sequence.

[0238] As can be seen from the above, when the scheme provided in the embodiment of the present application is applied to perform object posture smoothing, the posture information of the object in the target video frame, the forward reference frame of the target video frame and the backward reference frame of the target video frame is first obtained, and a preset number of cascaded spatiotemporal feature extractions are performed on the posture sequence including the obtained posture information to obtain the spatiotemporal fusion features of the posture sequence, thereby smoothing the posture sequence based on the spatiotemporal fusion features, and finally, based on the smoothed posture sequence, the posture information of the object in the target video frame after smoothing is obtained, thereby achieving object posture smoothing.

[0239] Among them, for the target video frame, the forward reference frame is a historical video frame, and the backward reference frame is a future video frame. In addition, since both the previous reference frame and the backward reference frame are temporally associated with the target video frame, the previous reference frame, the backward reference frame and the target video frame are also mutually associated in content. In this way, when the target video frame is smoothed based on the posture sequence containing the forward video frame and the backward video frame, the historical video frames and the future video frames associated with the content of the target video frame are comprehensively considered, which is conducive to eliminating the jitter of the posture information of the target video frame based on the association of the content of each video frame, so that the posture information of the target video frame is smoother and more coherent.

[0240] In addition, the number of backward reference frames is smaller than the number of forward reference frames, that is, when smoothing the object posture of the target video frame, the future information used is less than the historical information. This ensures that less future information is used to smooth the object posture of the target video frame, thereby improving the real-time performance of the object posture smoothing scheme.

[0241] Furthermore, when extracting features from the posture sequence, multiple cascade feature extractions are performed in the time dimension and the space dimension, so that deep features of the posture sequence can be extracted in the time dimension and the space dimension, making the extracted features richer and more comprehensive, and improving the smoothing effect when smoothing the object posture of the target video frame based on the extracted features.

[0242] In one embodiment of the present application, the spatiotemporal feature extraction module 802 includes:

[0243] A spatial feature extraction submodule is used to extract spatial features of each posture information in the target sequence in the spatial dimension to obtain spatial features of each posture information, wherein the initial value of the target sequence is a posture sequence including the obtained posture information;

[0244] A spatial feature fusion submodule is used to fuse the spatial features of each posture information to obtain the spatial fusion features of the target sequence;

[0245] The time feature extraction submodule is used to extract the time feature of the spatial fusion feature in the time dimension to obtain the spatiotemporal fusion feature; if the number of spatiotemporal feature extractions is less than a preset number, the target sequence is updated to the spatiotemporal fusion feature to trigger the spatial feature extraction submodule.

[0246] As can be seen from the above, in this embodiment, spatial features are first extracted for each posture information in the target sequence in the spatial dimension, and the obtained spatial features are fused to obtain spatial fusion features, which reflect the overall spatial features of each posture information obtained; then, feature extraction is performed on the above spatial fusion features in the time dimension to obtain spatiotemporal fusion features, which reflect both the characteristics of the posture sequence in the spatial dimension and the characteristics of the posture sequence in the time dimension; and the target sequence is updated to the spatiotemporal fusion features, and the step of spatial feature extraction is returned until the preset number of spatiotemporal feature extractions is reached. It can be seen that through the above-mentioned hierarchical feature extraction method, the deep spatiotemporal fusion features of the posture sequence can be finally obtained, which is conducive to improving the smoothing effect of the subsequent object posture smoothing based on the spatiotemporal fusion features.

[0247] In one embodiment of the present application, the posture information includes: joint point position information of the object;

[0248] The spatial feature extraction submodule is specifically used for obtaining the first position code of each joint point corresponding to each posture information based on the joint point position information included in each posture information in the target sequence when performing feature extraction in the spatial dimension for the first time, superimposing the first position code of each joint point corresponding to the posture information on each posture information, performing spatial feature extraction on the posture information after superimposing the first position code in the spatial dimension, and obtaining the spatial features of each posture information; when performing feature extraction in the spatial dimension not for the first time, directly performing spatial feature extraction on each posture information in the target sequence in the spatial dimension to obtain the spatial features of each posture information.

[0249] As can be seen from the above, if the posture information includes the position information of the joint points of the object, then when the feature extraction is performed in the spatial dimension for the first time, the first position code of each joint point corresponding to each posture information can be obtained based on the position information of the joint points included in each posture information, and the first position code of each joint point corresponding to the posture information can be superimposed on each posture information, and then the spatial feature extraction of the posture information after the superimposed first position code is performed in the spatial dimension to obtain the spatial features of each posture information. Among them, the first position code is obtained based on the joint point position information, which can reflect the posture of the object in a more fine-grained manner. In this way, when obtaining the spatial features of each posture information, not only the overall features of the posture information are considered, but also the first position code that can reflect the posture of the object in a more fine-grained manner is considered, thereby improving the accuracy of the spatial features finally obtained.

[0250] In one embodiment of the present application, the time feature extraction submodule is specifically used to obtain the second position code of all joint points corresponding to each posture information based on the joint point position information included in each posture information when performing feature extraction in the time dimension for the first time, superimpose the obtained second position code on the spatial fusion feature, and perform time feature extraction on the spatial fusion feature after superimposing the second position code in the time dimension to obtain a time-space fusion feature; when feature extraction is not performed in the time dimension for the first time, time feature extraction is directly performed on the spatial fusion feature in the time dimension to obtain a time-space fusion feature. If the number of time-space feature extractions is less than a preset number, the target sequence is updated to the time-space fusion feature to trigger the spatial feature extraction submodule.

[0251] As can be seen from the above, if the posture information includes the position information of the joint points of the object, then when the feature extraction is performed in the time dimension for the first time, the second position code of all the joint points corresponding to each posture information can be obtained based on the position information of the joint points included in each posture information, and the obtained second position code is superimposed on the spatial fusion feature, and the spatial fusion feature after superimposing the second position code is subjected to time feature extraction in the time dimension to obtain the spatiotemporal fusion feature. Among them, the second position code is obtained based on the joint point position information, which can reflect the posture of the object in a more fine-grained manner. In this way, when obtaining the spatiotemporal fusion feature, not only the overall spatial fusion feature of the posture information is considered, but also the second position code that can reflect the posture of the object in a more fine-grained manner is considered, thereby improving the accuracy of the finally obtained spatiotemporal fusion feature.

[0252] In one embodiment of the present application, the number of the backward reference frames is 1.

[0253] The backward reference frame is a video frame used for reference when smoothing the object posture of the target video frame, and the backward reference frame is collected after the target video frame. For the target video frame, the backward reference frame is future information. When the number of backward reference frames is 1, the future information referenced when smoothing the object posture of the target video frame is minimized, which greatly reduces the future information required for reference when smoothing the object posture of the target video frame, thereby greatly improving the real-time performance of the object posture smoothing solution.

[0254] In one embodiment of the present application, the object includes at least one of a human, an animal, and a robot.

[0255] In one embodiment of the present application, the spatiotemporal feature extraction module 802 is specifically used to input the posture sequence including the obtained posture information into the first spatiotemporal cyclic coding layer in the pre-trained posture smoothing model, and the spatiotemporal feature extraction is performed in sequence on the spatiotemporal cyclic coding layers connected at each level, and the spatiotemporal fusion features of the posture sequence output by the last spatiotemporal cyclic coding layer are obtained, wherein the posture smoothing model includes: a preset number of spatiotemporal cyclic coding layers and a regression head layer, the spatiotemporal cyclic coding layer includes a spatial encoder and a time encoder, the spatial encoder is used to extract features in the spatial dimension, and the time encoder is used to extract features in the time dimension for the output result of the spatial encoder;

[0256] The smoothing processing module 803 is specifically used to input the spatiotemporal fusion features into the regression head layer, and perform smoothing on the posture sequence to obtain a smoothed posture sequence.

[0257] It can be seen that through the cascade feature extraction of multiple spatiotemporal cyclic coding layers in the posture smoothing model, the spatiotemporal fusion features of the posture sequence can be quickly and comprehensively extracted. Then, the above-mentioned spatiotemporal fusion features are input into the regression head layer, which can quickly and accurately smooth the posture sequence to obtain a smooth posture sequence.

[0258] In one embodiment of the present application, the posture smoothing model is trained in the following manner:

[0259] A sample posture sequence and annotated posture sequence corresponding to a sample video frame are obtained, wherein the sample posture sequence includes: posture information obtained by performing posture estimation on the sample video frame, the forward sample reference frame of the sample video frame and the backward sample reference frame of the sample video frame; the annotated posture sequence includes: real posture information of the object in the sample video frame, the forward sample reference frame and the backward sample reference frame, and the number of forward sample reference frames is greater than the number of backward sample reference frames; the sample posture sequence is input into a preset model to obtain a smoothed posture sequence output by the model after smoothing the sample posture sequence; based on the annotated posture sequence and the smoothed posture sequence, a loss value of the model for smoothing is determined; based on the loss value, the network parameters of the model are adjusted to obtain a posture smoothing model.

[0260] It can be seen that in this way, based on the sample posture sequence and the labeled posture sequence corresponding to the sample video frames, a posture smoothing model can be quickly and reasonably trained.

[0261] In one embodiment of the present application, the loss value is obtained based on at least one of the following differences:

[0262] A first posture difference between the object posture represented by the labeled posture sequence and the object posture represented by the smooth posture sequence; a speed difference between the change speed of the object posture represented by the labeled posture sequence and the change speed of the object posture represented by the smooth posture sequence; a difference between the second posture difference and the third posture difference, wherein the second posture difference is: the difference between the object postures represented by the labeled posture sequences of adjacent sample video frames, and the third posture difference is: the difference between the object postures represented by the smooth posture sequences of adjacent sample video frames.

[0263] The loss value during model training is obtained based on the first posture difference between the object posture represented by the labeled posture sequence and the object posture represented by the smooth posture sequence, which can constrain the smooth posture sequence output by the model to be close to the labeled posture sequence, which is beneficial to improving the object posture smoothing effect of the model; the loss value during model training is obtained based on the speed difference between the speed of change of the object posture represented by the labeled posture sequence and the speed of change of the object posture represented by the smooth posture sequence, which can constrain the object posture change speed represented by the smooth posture sequence output by the model to be close to the object posture change speed represented by the labeled posture sequence, which is beneficial to improving the object posture smoothing effect of the model; the second posture difference actually reflects the difference between the labeled posture sequences corresponding to adjacent sample video frames; the third posture difference actually reflects the queue difference between the smooth posture sequences corresponding to adjacent sample video frames. In this way, the loss value during model training is obtained based on the difference between the second posture difference and the third posture difference, which can constrain the queue difference between the smooth posture sequences obtained by the model for two adjacent sample video frames to be close to the queue difference between the labeled posture sequences, so that in the final reasoning, the posture information of the sample video frames after smoothing can be smoothly connected, reducing the probability of jitter.

[0264] In one embodiment of the present application, the posture sequence is: a 9D rotation matrix capable of continuously representing rotation semantics;

[0265] The smooth posture sequence is: a 6D rotation representation that can continuously represent rotation semantics.

[0266] It can be seen that using the 9D rotation matrix that can continuously represent the rotation semantics as the model input and using the 6D rotation representation that can continuously represent the rotation semantics as the model output standardizes the input and output of the model, is conducive to eliminating data ambiguity, and can speed up the reasoning and training of the model, which is conducive to the real-time deployment of this solution in the industrial field and improves the efficiency and stability of the object posture smoothing solution.

[0267] In the technical solution of this application, the operations involved in obtaining, storing, using, processing, transmitting, providing and disclosing user personal information are all carried out with the user's authorization.

[0268] It should be noted that the two-dimensional face images, two-dimensional body images, etc. in this embodiment are from public data sets.

[0269] The present application also provides an electronic device, such as Fig. 9 As shown, including:

[0270] Memory 901, used for storing computer programs;

[0271] The processor 902 is used to implement the aforementioned object posture smoothing method when executing the program stored in the memory 901.

[0272] Furthermore, the electronic device may further include a communication bus and / or a communication interface, and the processor 902, the communication interface, and the memory 901 communicate with each other via the communication bus.

[0273] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0274] The communication interface is used for communication between the above electronic device and other devices.

[0275] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0276] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0277] In another embodiment provided in the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the aforementioned object posture smoothing methods are implemented.

[0278] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any of the object posture smoothing methods in the above embodiments.

[0279] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a solid-state hard disk (SSD), etc.

[0280] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0281] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, electronic device and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0282] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A method for smoothing an object posture, characterized in that: include: Obtaining posture information of an object in a target video frame, a forward reference frame of the target video frame, and a backward reference frame of the target video frame, wherein the number of the forward reference frames is greater than the number of the backward reference frames; Performing a preset number of cascaded spatiotemporal feature extractions on a posture sequence including the obtained posture information to obtain a spatiotemporal fusion feature of the posture sequence, wherein the spatiotemporal feature extraction includes: performing feature extraction in a spatial dimension, and performing feature extraction in a temporal dimension based on the feature extraction result of the spatial dimension; Based on the spatiotemporal fusion features, smoothing the posture sequence; Based on the smoothed posture sequence, the posture information of the object in the target video frame is obtained after smoothing the posture information.

2. The method according to claim 1, characterized in that The step of performing a preset number of cascaded spatiotemporal feature extractions on a posture sequence including the acquired posture information to obtain spatiotemporal fusion features of the posture sequence includes: Extracting spatial features of each posture information in the target sequence in the spatial dimension to obtain spatial features of each posture information, wherein the initial value of the target sequence is a posture sequence including the obtained posture information; Fusing the spatial features of each posture information to obtain the spatial fusion features of the target sequence; Performing time feature extraction on the spatial fusion feature in the time dimension to obtain a spatiotemporal fusion feature; If the number of spatiotemporal feature extractions is less than a preset number, the target sequence is updated to the spatiotemporal fusion features, and the process returns to the step of extracting spatial features for each posture information in the target sequence in the spatial dimension.

3. The method according to claim 2, characterized in that The posture information includes: joint point position information of the object; The step of extracting spatial features of each posture information in the target sequence in the spatial dimension to obtain the spatial features of each posture information includes: When performing feature extraction in the spatial dimension for the first time, based on the joint point position information included in each posture information in the target sequence, the first position code of each joint point corresponding to each posture information is obtained, the first position code of each joint point corresponding to the posture information is superimposed on each posture information, and spatial feature extraction is performed on the posture information after superimposing the first position code in the spatial dimension to obtain the spatial feature of each posture information; When feature extraction is performed in the spatial dimension for the first time, spatial feature extraction is performed directly on each posture information in the target sequence in the spatial dimension to obtain the spatial features of each posture information.

4. The method according to claim 2, characterized in that: The extracting time features of the spatial fusion features in the time dimension to obtain the spatiotemporal fusion features includes: When performing feature extraction in the time dimension for the first time, based on the joint point position information included in each posture information, a second position code of all joint points corresponding to each posture information is obtained, the obtained second position code is superimposed on the spatial fusion feature, and time feature extraction is performed on the spatial fusion feature after superimposing the second position code in the time dimension to obtain a spatiotemporal fusion feature; When the feature extraction is performed in the time dimension for the first time, the time feature extraction is performed directly on the spatial fusion feature in the time dimension to obtain the spatiotemporal fusion feature.

5. The method according to any one of claims 1 to 4, characterized in that The number of the backward reference frame is 1; and / or The object includes at least one of a human, an animal, and a robot.

6. The method according to any one of claims 1 to 4, characterized in that The step of performing a preset number of cascaded spatiotemporal feature extractions on a posture sequence including the acquired posture information to obtain spatiotemporal fusion features of the posture sequence includes: Input the posture sequence including the obtained posture information into the first spatiotemporal cyclic coding layer in the pre-trained posture smoothing model, and sequentially perform spatiotemporal feature extraction on the spatiotemporal cyclic coding layers of each level of cascade connection to obtain the spatiotemporal fusion features of the posture sequence output by the last spatiotemporal cyclic coding layer, wherein the posture smoothing model comprises: a preset number of spatiotemporal cyclic coding layers and a regression head layer, the spatiotemporal cyclic coding layer comprises a spatial encoder and a temporal encoder, the spatial encoder is used to perform feature extraction in the spatial dimension, and the temporal encoder is used to perform feature extraction in the temporal dimension for the output result of the spatial encoder; The step of smoothing the posture sequence based on the spatiotemporal fusion feature includes: The spatiotemporal fusion features are input into the regression head layer, and the posture sequence is smoothed to obtain a smoothed posture sequence.

7. The method according to claim 6, characterized in that The posture smoothing model is trained in the following manner: Obtaining a sample pose sequence and an annotated pose sequence corresponding to a sample video frame, wherein the sample pose sequence includes: pose information obtained by performing pose estimation on the sample video frame, a forward sample reference frame of the sample video frame, and a backward sample reference frame of the sample video frame; the annotated pose sequence includes: real pose information of the object in the sample video frame, the forward sample reference frame, and the backward sample reference frame, and the number of the forward sample reference frames is greater than the number of the backward sample reference frames; Inputting the sample posture sequence into a preset model to obtain a smooth posture sequence output by the model after smoothing the sample posture sequence; Determining a loss value of smoothing of the model based on the labeled posture sequence and the smoothed posture sequence; Based on the loss value, the network parameters of the model are adjusted to obtain a posture smoothing model.

8. The method according to claim 7, characterized in that The loss value is obtained based on at least one of the following differences: a first posture difference between a posture of the object represented by the labeled posture sequence and a posture of the object represented by the smoothed posture sequence; a speed difference between a speed of change of the object posture represented by the annotated posture sequence and a speed of change of the object posture represented by the smoothed posture sequence; The difference between the second posture difference and the third posture difference, wherein the second posture difference is: the difference between the object postures represented by the annotated posture sequence of adjacent sample video frames, and the third posture difference is: the difference between the object postures represented by the smooth posture sequence of adjacent sample video frames.

9. The method according to claim 6, characterized in that The posture sequence is: a 9D rotation matrix capable of continuously representing rotation semantics; The smooth posture sequence is: a 6D rotation representation that can continuously represent rotation semantics.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.