Skeletal motion prediction processing method, device and limb motion prediction processing method

Through machine learning models, the bone motion state vector is characterized by encoding and decoding and extracting motion intentions, which solves the problem that traditional motion prediction technology cannot accurately predict the target object's own motion details, and realizes the detailed prediction of bone motion.

CN114495283BActive Publication Date: 2025-08-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210143192.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-01-19
Publication Date
2025-08-22
Estimated Expiration
2038-01-19

AI Technical Summary

Technical Problem

Traditional motion prediction techniques cannot achieve accurate prediction of the target object's own motion details, and usually can only predict based on the position changes of fixed observation points.

Method used

By obtaining multiple continuous historical bone motion state vectors, using machine learning models for feature encoding and decoding, extracting motion intentions, and predicting the bone motion state at the next moment.

Benefits of technology

The detailed prediction of the bone movement of the target object is achieved, and the movement state can be accurately predicted at the next moment, improving the accuracy and detailed description of the motion prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495283B_ABST
    Figure CN114495283B_ABST
Patent Text Reader

Abstract

The present invention relates to a skeletal motion prediction processing method, device, and limb motion prediction processing method, which includes: inputting multiple continuous skeletal motion state vectors of historical observations into a pre-trained machine learning model to perform feature encoding to obtain corresponding skeletal motion feature vectors; determining the skeletal motion implicit state vector at the previous moment before the current moment; the skeletal motion implicit state vector at the previous moment is obtained by extracting the motion intention of the skeletal motion feature vector at the previous moment; obtaining the skeletal motion state vector at the current moment; decoding the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment to calculate the skeletal motion state vector at the next moment. The solution of the present application realizes motion prediction processing of the details of predicting the skeletal motion of the target object itself.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application submitted to the China Patent Office on January 19, 2018, with application number 2018100552132, and the invention name is "Skeletal motion prediction and processing method, device and limb motion prediction and processing method", all of which are incorporated by reference into this application. Technical Field

[0002] The present invention relates to the field of computer technology, and in particular to a skeletal motion prediction processing method, device, and limb motion prediction processing method. Background Art

[0003] With the rapid development of science and technology, various cutting-edge technologies are gaining more and more attention. For example, motion prediction refers to predicting the next possible motion based on the motion that has already occurred.

[0004] However, traditional motion prediction technology is not yet mature. It typically selects a fixed observation point within the target object (for example, the center of the target object) and analyzes its movement to predict the target object's next possible location. Obviously, traditional motion prediction technology can only predict the position change of the target object as a whole, and cannot predict the specific movement details of the target object itself. Summary of the Invention

[0005] Based on this, it is necessary to provide a skeletal motion prediction processing method, device, computer equipment and storage medium to address the problem that traditional methods cannot predict the motion details of the target object itself, and also provide a limb motion prediction processing method, device, computer equipment and storage medium.

[0006] A skeletal motion prediction and processing method, the method comprising:

[0007] Obtain multiple continuous historical skeletal motion state vectors;

[0008] Performing feature code encoding on each of the skeletal motion state vectors using a machine learning model to generate skeletal motion feature vectors corresponding to each of the skeletal motion state vectors;

[0009] Determine the skeletal motion implicit state vector at the previous moment before the current moment; the skeletal motion implicit state vector at the previous moment is obtained by extracting motion intention from the skeletal motion feature vector at the previous moment;

[0010] Get the current skeletal motion state vector;

[0011] Decoding is performed based on the skeleton motion implicit state vector at the previous moment and the skeleton motion state vector at the current moment to calculate the skeleton motion state vector at the next moment.

[0012] A skeletal motion prediction and processing device, comprising:

[0013] An encoding module is used to obtain a plurality of continuous historical skeletal motion state vectors; perform feature code encoding on each of the skeletal motion state vectors through a machine learning model to generate skeletal motion feature vectors corresponding to each of the skeletal motion state vectors;

[0014] An implicit state vector determination module is used to determine the implicit state vector of the skeletal motion at the previous moment before the current moment; the implicit state vector of the skeletal motion at the previous moment is obtained by extracting the motion intention of the skeletal motion feature vector at the previous moment;

[0015] A skeleton motion state vector acquisition module is used to obtain the skeleton motion state vector at the current moment;

[0016] The decoding prediction module is used to decode according to the skeleton movement implicit state vector at the previous moment and the skeleton movement state vector at the current moment to calculate the skeleton movement state vector at the next moment.

[0017] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0018] Obtain multiple continuous historical skeletal motion state vectors;

[0019] Performing feature code encoding on each of the skeletal motion state vectors using a machine learning model to generate skeletal motion feature vectors corresponding to each of the skeletal motion state vectors;

[0020] Determine the skeletal motion implicit state vector at the previous moment before the current moment; the skeletal motion implicit state vector at the previous moment is obtained by extracting motion intention from the skeletal motion feature vector at the previous moment;

[0021] Get the current skeletal motion state vector;

[0022] Decoding is performed based on the skeleton motion implicit state vector at the previous moment and the skeleton motion state vector at the current moment to calculate the skeleton motion state vector at the next moment.

[0023] A storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0024] Obtain multiple continuous historical skeletal motion state vectors;

[0025] Performing feature code encoding on each of the skeletal motion state vectors using a machine learning model to generate skeletal motion feature vectors corresponding to each of the skeletal motion state vectors;

[0026] Determine the skeletal motion implicit state vector at the previous moment before the current moment; the skeletal motion implicit state vector at the previous moment is obtained by extracting motion intention from the skeletal motion feature vector at the previous moment;

[0027] Get the current skeletal motion state vector;

[0028] Decoding is performed based on the skeleton motion implicit state vector at the previous moment and the skeleton motion state vector at the current moment to calculate the skeleton motion state vector at the next moment.

[0029] The above-mentioned skeletal motion prediction processing method, device, computer equipment and storage medium respectively perform feature coding on a plurality of continuous historical skeletal motion state vectors to obtain corresponding skeletal motion feature vectors, thereby realizing skeletal motion feature extraction of each skeletal motion state vector. The skeletal motion implicit state vector of the previous moment of the current moment is determined; the skeletal motion implicit state vector is obtained by extracting motion intention from the skeletal motion feature vector obtained at the previous moment, thereby realizing extraction of motion intention according to the extracted skeletal motion features. Decoding is performed based on the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to calculate the skeletal motion state vector of the next moment. Decoding is performed by the extracted motion intention and the current skeletal motion state vector, thereby realizing mining of the extracted intention information, thereby calculating the skeletal motion state vector of the next moment based on the decoded mining intention information. The motion prediction processing of the prediction of the skeletal motion of the target object itself is realized.

[0030] A method for predicting and processing limb movements, the method comprising:

[0031] Acquire multiple continuous historical limb movement states;

[0032] Extracting motion features of the limb motion states in each history to obtain limb motion features corresponding to the limb motion states in each history;

[0033] Determining the implicit state of limb movement at a moment before the current moment; the implicit state of limb movement at the previous moment is obtained by extracting movement intention from the limb movement features at the previous moment;

[0034] Get the current limb movement status;

[0035] The limb movement state at the next moment is determined according to the limb movement implicit state at the previous moment and the limb movement state at the current moment.

[0036] A limb movement prediction and processing device, characterized in that the device comprises:

[0037] A motion feature extraction module is used to obtain a plurality of continuous historical limb motion states; extract motion features from each of the historical limb motion states to obtain limb motion features corresponding to each of the historical limb motion states;

[0038] An implicit state determination module is used to determine the implicit state of the limb movement at the previous moment before the current moment; the implicit state of the limb movement at the previous moment is obtained by extracting the movement intention of the limb movement features at the previous moment;

[0039] The limb movement state determination module is used to obtain the limb movement state at the current moment; and determine the limb movement state at the next moment based on the limb movement implicit state at the previous moment and the limb movement state at the current moment.

[0040] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0041] Acquire multiple continuous historical limb movement states;

[0042] Extracting motion features of the limb motion states in each history to obtain limb motion features corresponding to the limb motion states in each history;

[0043] Determining the implicit state of limb movement at a moment before the current moment; the implicit state of limb movement at the previous moment is obtained by extracting movement intention from the limb movement features at the previous moment;

[0044] Get the current limb movement status;

[0045] The limb movement state at the next moment is determined according to the limb movement implicit state at the previous moment and the limb movement state at the current moment.

[0046] A storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0047] Acquire multiple continuous historical limb movement states;

[0048] Extracting motion features of the limb motion states in each history to obtain limb motion features corresponding to the limb motion states in each history;

[0049] Determining the implicit state of limb movement at a moment before the current moment; the implicit state of limb movement at the previous moment is obtained by extracting movement intention from the limb movement features at the previous moment;

[0050] Get the current limb movement status;

[0051] The limb movement state at the next moment is determined according to the limb movement implicit state at the previous moment and the limb movement state at the current moment.

[0052] The above-mentioned limb movement prediction processing method, device, computer equipment and storage medium extract movement features from multiple continuous historical limb movement states to obtain corresponding limb movement features. Determine the implicit state of limb movement at the previous moment before the current moment; the implicit state of limb movement is obtained by extracting movement intention from the limb movement features obtained at the previous moment, thereby realizing the extraction of movement intention based on the extracted limb movement features. Decode the implicit state of limb movement at the previous moment and the limb movement state at the current moment to calculate the limb movement state at the next moment. By decoding the extracted movement intention and the limb movement state at the current moment, the extracted intention information is mined, thereby calculating the limb movement state at the next moment based on the decoded and mined intention information. A more detailed motion prediction process of predicting the limb movement of the target object itself is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 1 is a flow chart of a method for predicting and processing skeletal motion in one embodiment;

[0054] Figure 2 A schematic diagram of the effect of a skeletal motion prediction processing method according to an embodiment;

[0055] Figure 3 Schematic diagram of a processing framework of a skeletal motion prediction processing method in one embodiment;

[0056] Figure 4 Schematic diagram of the principle framework of a method for predicting and processing skeletal motion in one embodiment;

[0057] Figure 5 FIG1 is a schematic diagram of a process for calculating a skeletal motion state vector at the next moment in one embodiment;

[0058] Figure 6 A schematic diagram of controlling a target object through a behavior tag in one embodiment;

[0059] Figure 7 A diagram illustrating an application environment for skeletal motion prediction processing in one embodiment;

[0060] Figure 8 is a flowchart of a skeletal motion prediction processing method according to another embodiment;

[0061] Figure 9 1 is a flow chart of a method for predicting and processing limb movements in one embodiment;

[0062] Figure 10 is a block diagram of a skeletal motion prediction and processing device in one embodiment;

[0063] Figure 11 is a block diagram of a skeletal motion prediction and processing device in another embodiment;

[0064] Figure 12 is a block diagram of a limb movement prediction and processing device in one embodiment;

[0065] Figure 13 Schematic diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0066] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0067] Figure 1 FIG. 1 is a flow chart of a method for predicting and processing bone motion in an embodiment. This embodiment mainly uses the method for predicting and processing bone motion in a computer device as an example. Figure 1 , the method specifically comprises the following steps:

[0068] S102: Acquire multiple continuous historical skeletal motion state vectors.

[0069] The skeletal motion state vector is a vector representation of the skeletal motion state, that is, it describes the skeletal motion state in the form of a vector. The historical skeletal motion state vector is used to represent the skeletal motion state that has already occurred. Multiple consecutive historical skeletal motion state vectors are used to represent multiple historical skeletal motion states that have been continuously generated.

[0070] In one embodiment, the skeletal motion state vector can be a rotation vector that describes the skeletal motion state by expressing the rotation change between adjacent target skeleton joint points. Wherein, the target skeleton joint point is a skeletal joint point that is pre-specified, that needs reference when generating the skeletal motion state vector. A rotation vector refers to a vector whose direction is an axis of rotation and whose size is an angle of rotation. It will be appreciated that in the skeletal motion state vector, each vector element is the rotation data of each target skeleton joint point relative to a previous target skeleton joint point.

[0071] Specifically, the computer device can directly obtain multiple continuous historical skeletal motion state vectors. The computer device can also obtain multiple continuously collected skeletal motion state image frames, and analyze the multiple continuously collected skeletal motion state image frames to obtain multiple continuous historical skeletal motion state vectors.

[0072] In one implementation, the method also includes: acquiring multiple frames of continuously acquired skeletal motion state image frames; for each frame of the skeletal motion state image frame, identifying multiple target skeletal joint points from the skeletal motion state image frame; acquiring the rotation data of the subsequent target skeletal joint point relative to the previous target skeletal joint point according to the front-to-back order between the multiple target skeletal joint points; using each rotation data as a vector element, and splicing them according to the front-to-back order between the corresponding target skeletal joint points to obtain a skeletal motion state vector corresponding to the skeletal motion state image frame.

[0073] The skeleton motion state image frame is an image including the skeleton motion state information of the target object. The target skeleton joint point is a pre-specified skeleton joint point to be referenced when generating the skeleton motion state vector. It should be noted that the target skeleton joint points identified in each skeleton motion state image frame are identical.

[0074] It is understandable that the target skeleton joint point can be all or part of the skeleton joint points of the target object. Because the number of skeleton joint points of the target object is very large, some skeleton joint points can be selected from all skeleton joint points as the target skeleton joint point when generating the skeleton motion state vector.

[0075] Specifically, the computer device can perform skeletal motion state analysis on each skeletal motion state image frame to obtain skeletal motion state vectors corresponding to each skeletal motion state image frame. The computer device can obtain multiple continuous historical skeletal motion state vectors based on the skeletal motion state vectors corresponding to multiple continuous skeletal motion state image frames.

[0076] In one embodiment, the computer device can directly obtain the front-to-back order between multiple pre-set target skeletal joints, or it can determine the front-to-back order between multiple target skeletal joints by combining the skeletal information of each target skeletal joint in the target object in the skeletal motion state image frame.

[0077] The computer device can obtain rotation data of a subsequent target skeletal joint point relative to a previous target skeletal joint point according to the order of the target skeletal joint points. The rotation data is data describing the rotation change of the subsequent target skeletal joint point compared to the previous target skeletal joint point. In one embodiment, the rotation data can be a rotation vector, which is used to describe the rotation angle and rotation direction in the form of a vector.

[0078] Can be understood that, because rotation data is that rear target skeleton joint point produces rotational change with respect to previous target skeleton joint point and obtains, so rotation data is corresponding with this rear target skeleton joint point.Because leading target skeleton joint point does not have previous target skeleton joint point, the rotation data of leading target skeleton joint point can be set as default initial value.In one embodiment, this default initial value can be set to zero.

[0079] It should be noted that, in this embodiment, "the next one" and "the previous one" are relative relationships. For example, assuming that the target bone joint point A is adjacent to the target bone joint point B and is located before B, then the target bone joint point A is the previous target bone joint point of the target bone joint point B, and the target bone joint point B is the next target bone joint point of the target bone joint point A.

[0080] The computer device can use each rotation data as a vector element, and splice them according to the front-to-back order between the corresponding target bone joint points to obtain a bone motion state vector corresponding to the bone motion state image frame.

[0081] The steps of generating the skeletal motion state vector are now explained with an example. For example, the target skeletal joints are A, B, C, D, and E in sequence, and A is the first target skeletal joint. The rotation data 1 corresponding to A is the initial default value 0, and the computer device can respectively obtain the rotation data 2 of B compared to A (rotation data 2 corresponds to B), the rotation data 3 of C compared to B (rotation data 3 corresponds to C), the rotation data 4 of D compared to C (rotation data 4 corresponds to D), and the rotation data 5 of E compared to D (rotation data 5 corresponds to E). The computer device can use the obtained rotation data as vector elements, and splice them according to the front-to-back order between the corresponding target skeletal joints to obtain the skeletal motion state vector x(0, ​​rotation data 1, rotation data 2, rotation data 3, rotation data 4, rotation data 5).

[0082] S104, feature coding each skeletal motion state vector respectively through a machine learning model to generate skeletal motion feature vectors corresponding to each skeletal motion state vector respectively.

[0083] The machine learning model is a model obtained through pre-training. The skeletal motion feature vector is a feature vector obtained by extracting motion features from the skeletal motion state vector. It can be understood that the feature encoding process is the process of extracting motion features.

[0084] In one embodiment, the machine learning model may include a recurrent neural network model (RNN).

[0085] The computer device can input the acquired multiple continuous historical skeletal motion state vectors into the pre-trained machine learning model to perform feature encoding on each historical skeletal motion state vector to obtain the corresponding skeletal motion feature vector of each skeletal motion state vector. It can be understood that each historical skeletal motion state vector has a corresponding historical moment. The historical moment is the moment when the computer device observes and collects the skeletal motion state represented by the historical skeletal motion state vector.

[0086] In one embodiment, a computer device can input a plurality of continuous skeletal motion state vectors of historical observations into an encoder in a pre-trained recursive neural network model for feature encoding to obtain skeletal motion feature vectors corresponding to each of the skeletal motion state vectors. It is understood that the historical moments of the skeletal motion state vectors corresponding to each of the skeletal motion feature vectors correspond.

[0087] For example, the multiple continuous skeletal motion state vectors of historical observations are {x 1 , x 2 , x 3 ,…,x T'}, each skeletal motion state vector corresponds to the historical time 1 to T', and the corresponding skeletal motion feature vector {e 1 , e 2 , e 3 ,…,e T'}, each skeletal motion feature vector also corresponds to historical moments 1 to T'.

[0088] In one embodiment, the computer device may encode and obtain the skeletal motion feature vector corresponding to the historically observed skeletal motion state vector using the following formula:

[0089] e t' =f e (x t'); (Formula 1)

[0090] Among them, t' is the t'th historical moment; x t' is the skeletal motion state vector at the t'th historical moment; e t ' is the skeletal motion feature vector at the t'th historical moment (i.e. the skeletal motion state vector x at the t'th historical moment t' The skeleton motion feature vector obtained after encoding); f e is the feature encoding function.

[0091] In one embodiment, step S104 includes: according to the sequence of each historical skeletal motion state vector, the skeletal motion feature vector obtained in the previous encoding and the skeletal motion state vector to be encoded at the current time are input into the encoder in the pre-trained machine learning model for encoding, and the skeletal motion feature vector obtained in the current encoding corresponding to the skeletal motion state vector to be encoded at the current time is output.

[0092] In one embodiment, the computer device may encode according to the following formula to obtain the skeletal motion feature vector:

[0093] e t' =W e φ(U e e t'-1 +b e )+U x x t' +b x ;

[0094] Among them, e t' is the skeletal motion feature vector obtained by encoding the skeletal motion state vector at the t'th historical moment; e t'-1 is the skeletal motion feature vector obtained by the previous encoding of the skeletal motion state vector at the t'-1th historical moment; t' is the skeletal motion state vector to be encoded at the t'th historical moment; W e 、U e 、b e 、U x and b x are pre-trained parameters; φ represents the linear rectification function.

[0095] S106: Determine the skeletal motion implicit state vector at the moment before the current moment.

[0096] The current moment is the moment when the skeletal motion state prediction process is currently being performed.

[0097] It can be understood that the current moment is a moment of relative change, that is, after the skeletal motion state vector is predicted for the next moment of the current moment, the next moment of the predicted skeletal motion state vector can be used as the new current moment, and steps S106 to S110 are continued to be executed to predict the skeletal motion state vector. The first current moment is the historical moment corresponding to the last skeletal motion state vector among the multiple continuous skeletal motion state vectors in step S102, that is, the skeletal motion state prediction is performed from the historical moment corresponding to the last skeletal motion state vector to predict the skeletal motion state vector of the next next moment.

[0098] For example, multiple continuous skeletal motion state vectors are {x 1 , x 2 , x 3 ,…,x T'}, the first current moment is T', and the skeletal motion state vector is predicted for the moment after T', that is, T'+1. After the skeletal motion state vector at T'+1 is predicted, T'+1 is used as the new current moment to predict the skeletal motion state vector for the moment after T'+1, that is, T'+2. And so on, the prediction process of skeletal motion state vectors for future moments is continuously carried out.

[0099] The skeletal motion implicit state vector is derived by extracting motion intent from the skeletal motion feature vector. It is used to integrate historical skeletal motion state feature information to express the corresponding motion intent. Motion intent is the intention to perform a certain movement and is used to reflect the desired next movement.

[0100] It is understandable that since the corresponding skeletal motion states at different moments may be different, the motion intention may be different, and each moment has its own corresponding skeletal motion implicit state vector. The skeletal motion implicit state vector at the previous moment of the current moment is obtained by extracting the motion intention of the skeletal motion feature vector at the previous moment.

[0101] In one embodiment, the skeletal motion latent state vector is output by a hidden layer in a recurrent neural network model.

[0102] In one embodiment, for first current moment (i.e. the historical moment corresponding to last skeletal motion state vector is as current moment), the skeletal motion hidden state vector of its last moment can be preset default value.In one embodiment, the skeletal motion hidden state vector preset default value can be initialized to zero vector.

[0103] It can be understood that in other embodiments, for the first current moment (i.e., the historical moment corresponding to the last skeletal motion state vector is taken as the current moment), the skeletal motion implicit state vector at the previous moment may not be the preset default value.

[0104] S108, obtaining the skeleton motion state vector at the current moment.

[0105] It can be understood that the skeletal motion state vector of the first current moment (i.e., the historical moment corresponding to the last historical skeletal motion state vector is used as the current moment) is the last skeletal motion state vector of the historical observation. When the corresponding moment of the predicted skeletal motion state vector is used as the current moment, the skeletal motion state vector of the current moment is the predicted skeletal motion state vector.

[0106] For example, the skeletal motion state vectors of multiple consecutive histories are {x 1 , x 2 , x 3 ,…,x T'}, then the first current moment is T', and the skeletal motion state vector at the first current moment T' is x T' , the skeletal motion state vector of the next moment T', that is, T'+1, will be predicted to be, assuming that the predicted value is x T'+1 When T'+1 is taken as the new current moment, the skeletal motion state vector at the current moment T'+1 is the calculated x T'+1 .

[0107] S110 , decoding is performed based on the skeleton motion implicit state vector at the previous moment and the skeleton motion state vector at the current moment to calculate the skeleton motion state vector at the next moment.

[0108] It can be understood that since the implicit state vector of the skeleton movement at the previous moment is obtained by extracting the motion intention of the historical skeleton movement features, the skeleton movement state vector of the next moment can be predicted by decoding the implicit state vector of the skeleton movement at the previous moment based on the motion intention extraction of the historical skeleton movement features and the skeleton movement state vector at the current moment.

[0109] Figure 2 FIG. 1 is a schematic diagram showing the effect of a method for predicting and processing skeletal motion in an embodiment. Figure 2 , area 202 shows the skeletal motion state represented by multiple continuous skeletal motion state vectors observed from time 1 to T'. Based on the historically observed skeletal motion state vectors in 202, the skeletal motion state vectors at the future moment immediately after time T' can be predicted. Area 204 shows the skeletal motion state represented by the predicted skeletal motion state vectors at the future moment.

[0110] The above-mentioned skeletal motion prediction processing method performs feature coding on a plurality of continuous historical skeletal motion state vectors respectively to obtain corresponding skeletal motion feature vectors, thereby realizing the extraction of skeletal motion features of each skeletal motion state vector. Determine the skeletal motion implicit state vector of the previous moment of the current moment; The skeletal motion implicit state vector is obtained by extracting the motion intention of the skeletal motion feature vector obtained at the previous moment, thereby realizing the extraction of motion intention according to the extracted skeletal motion features. Decode the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to calculate the skeletal motion state vector of the next moment. Decode the extracted motion intention and the current skeletal motion state vector, thereby realizing the mining of the extracted intention information, thereby calculating the skeletal motion state vector of the next moment based on the decoded mining intention information. Realize the motion prediction processing of this detail of the prediction of the skeletal motion of the target object itself.

[0111] In one embodiment, step S106 includes: obtaining the estimated velocity feature vector of the previous moment before the current moment; respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector of the previous moment; determining the weight of each bone motion feature vector based on the motion correlation; the weight is positively correlated with the motion correlation; and performing weighted summation on each bone motion feature vector according to the corresponding weight to obtain the bone motion implicit state vector of the previous moment before the current moment.

[0112] Among them, the estimated speed feature vector is used to characterize the change between the estimated skeletal motion state vectors at adjacent moments. The size of the change between the skeletal motion state vectors at adjacent moments is positively correlated with the size of the estimated speed feature vector. The estimated speed feature vector of the previous moment before the current moment, that is, the speed feature reflecting the closest history, is used to characterize the change between the skeletal motion state vectors of the previous moment before the current moment and the previous moment before the current moment.

[0113] In one embodiment, for the first current moment, the computer device may perform velocity feature analysis based on the skeletal motion state vector of the previous moment and the skeletal motion state vector of the previous moment before the current moment, and estimate an estimated velocity feature vector.

[0114] The motion correlation between the skeletal motion feature vector and the estimated velocity feature vector is used to describe the correlation between the motion information represented by the skeletal motion feature vector and the skeletal motion information represented by the velocity feature vector.

[0115] Specifically, computer equipment can determine the weight of each skeletal motion eigenvector according to the motion correlation between each skeletal motion eigenvector and the estimated speed eigenvector of the previous moment; Weight is positively correlated with motion correlation. It can be understood that the estimated speed eigenvector of the previous moment is the speed feature of the closest history, that is, the motion correlation between the skeletal motion eigenvector of the history corresponding to each historical moment and the speed feature of the closest history, distribute the weight shared by the skeletal motion eigenvector of each history when participating in the motion intention extraction process, the skeletal motion eigenvector of each history is carried out weighted summation according to corresponding weight respectively, obtain the skeletal motion implicit state vector of the previous moment of the current moment. Wherein, the greater the motion correlation is, the greater the weight is, the greater the impact of motion intention extraction is, on the contrary, the less the motion correlation is, the less the weight is, the less the impact of motion intention extraction is.

[0116] In one embodiment, respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment includes: inputting the estimated velocity feature vector at the previous moment and each bone motion feature vector into an attention model in a machine learning model, and respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment according to the attention model.

[0117] The Attention Model is a machine learning model used to extract the skeletal motion feature vectors that are more critical for extracting motion intent from the historically observed skeletal motion feature vectors. It is understood that the Attention Model can analyze and determine the motion correlation between each skeletal motion feature vector and the estimated velocity feature vector at the previous moment to extract the skeletal motion feature vector that is more critical for extracting motion intent.

[0118] In one embodiment, the computer device may determine the motion correlation between the estimated velocity feature vector at the previous moment and each skeletal motion feature vector using the following formula:

[0119] β t' =W β tanh(U βν v t-1 +U βe e t' +b β ); (Formula 2)

[0120] Among them, t is the current moment; t-1 is the moment before the current moment; v t-1 is the estimated velocity feature vector of the previous moment; t' is the t'th historical moment; e t' is the skeletal motion feature vector at the t'th historical moment; β t'is the skeletal motion feature vector e at the t'th historical moment t' The estimated velocity feature vector v at the previous moment and the current moment t-1 The correlation between them; tanh() is the hyperbolic tangent function; W β 、U βν 、U βe and b β These are all parameters obtained by pre-training the attention model in the machine learning model.

[0121] It should be noted that, for the sake of clarity, the embodiments of this application distinguish between the moment of prediction processing and the historical moment, that is, a stroke (i.e., t') is added to the upper right corner of t to represent the historical moment, while the moment of prediction processing does not have a stroke (i.e., t) added to the upper right corner of t. It can be understood that the moment t of prediction processing is the moment value added on the basis of the last moment of the historical moment, that is, assuming that the last moment of the historical moment (i.e., the last historical moment) is T', then the moment of prediction processing is T'+t. In order to express it concisely and clearly, the embodiments of this application remove T' when representing the moment of prediction processing, and directly use t to represent it. Wherein, t is greater than or equal to 0.

[0122] It can be understood that the above tanh() function can be replaced by other activation functions, such as the sigmoid function (S-shaped curve function).

[0123] In one embodiment, the computer device may determine the weight of each skeletal motion feature vector using the following formula:

[0124]

[0125] Among them, α t' is the weight of the skeletal motion feature vector at the t'th historical moment; β t' is the correlation between the skeletal motion feature vector at the t'th historical moment and the estimated velocity feature vector at the previous moment; T' is the last historical moment.

[0126] In one embodiment, the computer device may obtain the skeletal motion implicit state vector at the previous moment by the following formula:

[0127]

[0128] Among them, t-1 is the moment before the current moment; h t-1 is the skeletal motion implicit state vector of the previous moment; e t' is the skeletal motion feature vector at the t'th historical moment; α t'is the weight of the skeletal motion feature vector at the t'th historical moment; T' is the last historical moment.

[0129] Figure 3 FIG. 1 is a schematic diagram of a processing framework of a method for predicting skeletal motion in an embodiment. t ' represents the skeletal motion state vector at the t'th historical moment, where t' ranges from 1 to T'. Represents multiple continuous skeletal motion state vectors of historical observations; each skeletal motion state vector is input into the encoder for encoding to obtain the corresponding skeletal motion feature vector e t' , similarly, e t' The value of t' in the equation ranges from 1 to T', and then the eigenvectors of the skeletal motion e t' Input the attention model and transform each skeletal motion feature vector e t 'And the estimated speed feature vector v at the previous moment of the current moment t-1 The correlation β between the two is obtained through the sigmoid function layer t' , β t' Input the softmax function layer for mapping and obtain the corresponding weight α t' , and each bone motion feature vector e t 'According to the corresponding weight α t' Perform weighted summation and output the skeletal motion implicit state vector h at the previous moment t-1 before the current moment t t-1 . Among them, t ranges from 0 to T. It can be understood that T here is any integer value greater than or equal to 0. The implicit state vector h of the skeletal motion at the previous moment t-1 is t-1 And the skeletal motion state vector of the current moment t is input into the decoder for decoding, and the skeletal motion state vector of the next moment t+1 of the current moment is output. It can be understood that in a new round of prediction processing, the next moment in the previous round of prediction processing is the new current moment, then the skeletal motion state vector of the new current moment, that is, the skeletal motion state vector predicted in the previous round, will participate in this new round of prediction processing to continue calculating the skeletal motion state vector of the next moment. Among them, the decoder is used to decode the encoded information. The softmax function is a function that maps multiple input values ​​so that the sum of the mapped values ​​is 1.

[0130] In the above-described embodiment, by the motion correlation between the skeletal motion characteristic vector of the history corresponding to the speed characteristics of the nearest history and each historical moment, the weight shared by the skeletal motion characteristic vector of each history when participating in the extraction process of motion intention is distributed, and the accuracy of weight distribution is guaranteed. Therefore, the skeletal motion characteristic vector of each history is carried out weighted summation by relevant weight, obtains the skeletal motion implicit state vector of the last moment, guarantees the accuracy of motion intention extraction, namely guarantees the accuracy of the skeletal motion implicit state vector of the last moment of the current moment.

[0131] In one embodiment, step S110 includes: inputting the implicit state vector of the skeleton movement at the previous moment and the skeleton movement state vector at the current moment into the decoder in the machine learning model for decoding to obtain the estimated speed feature vector at the current moment; calculating the skeleton movement state vector at the next moment based on the estimated speed feature vector at the current moment and the skeleton movement state vector at the current moment.

[0132] The decoder is used to decode the encoded information. It is understood that the next moment here refers to the moment after the current moment. In one embodiment, the decoder can be a Modified Highway Unit (MHU), which is used to integrate historical encoding and decoding information into the next encoding and decoding process to achieve high-speed encoding and decoding.

[0133] In one embodiment, the computer device may determine the estimated speed feature vector at the current moment according to the following formula:

[0134]

[0135] Among them, t is the current moment, t-1 is the moment before the current moment; v t is the estimated velocity feature vector at the current moment; h t-1 is the implicit state vector of the skeletal motion at the previous moment; φ represents the linear rectification function; W v 、U vh 、b v and b vh are all bias parameters obtained by pre-training in the machine learning model; is the skeletal motion state vector at the current moment.

[0136] Figure 4 FIG. 1 is a schematic diagram of the principle framework of a method for predicting and processing skeletal motion in an embodiment. Figure 4 Each time it passes through an MHU unit 402, it means that it has undergone a decoding prediction process. The decoding prediction process is the same each time, so only the first MHU will be explained. t-1and the skeletal motion state vector at the current time t The input is decoded in the MHU unit 402, and the output is the estimated velocity feature vector v at the current moment. t And predict the skeletal motion state vector at the next moment t+1 And v t and each bone motion feature vector e t ' (where t' ranges from 1 to T') is input into the attention model 404, and is processed by correlation comparison and the like, and the skeletal motion implicit state vector is output. It can be understood that when the next moment t+1 predicted by the previous round of prediction processing is used as the current moment t of the next new round, the skeletal motion implicit state vector outputted in the previous round is relatively recorded as the skeletal motion implicit state vector h of the previous moment in the time dimension. t-1 The skeletal motion state vector predicted in the previous round is relatively recorded as the skeletal motion state vector of the current time t in the new round in the time dimension Enter the next MHU unit.

[0137] In one embodiment, based on the estimated velocity feature vector at the current moment and the skeletal motion state vector at the current moment, calculating the skeletal motion state vector at the next moment includes: obtaining a first prediction weight vector corresponding to the skeletal motion state vector at the current moment through a decoder; determining a second prediction weight vector of the estimated velocity feature vector at the current moment based on the first prediction weight vector; the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector; performing dot multiplication on the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment with the corresponding first prediction weight vector and the second prediction weight vector, and then adding them to obtain the skeletal motion state vector at the next moment.

[0138] Among them, the first prediction weight vector is used to characterize the weight of the skeletal motion state vector at the current moment when it participates in the skeletal motion state prediction process. The second prediction weight vector is used to characterize the weight of the estimated velocity feature vector at the current moment when it participates in the skeletal motion state prediction process. The all-one vector is a vector whose vector elements are all 1. It should be noted that the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector.

[0139] In one embodiment, the computer device may obtain the skeletal motion state vector at the next moment according to the following formula:

[0140]

[0141]

[0142] Where t is the current time; z t is the first prediction weight vector; 1-z tis the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; v t is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx Both are bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeleton motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol. It can be understood that 1-z t The 1 in represents an all-one vector.

[0143] Figure 5 FIG. 1 is a schematic diagram of a process for calculating the skeletal motion state vector at the next moment in an embodiment. Figure 5 For each decoding unit, the input is the skeletal motion implicit state vector h at the previous moment t-1 and the current skeletal motion state vector h t-1 The result obtained after the linear rectification function layer and the linear layer processing is the same as The results of the linear layer processing are added to obtain the estimated speed feature vector v at the current moment t (corresponding to the above formula 5), ​​the obtained v t On the one hand, it outputs and participates in the decoding of the next decoding unit. In addition, the obtained v t The specific process of the decoding of the current decoding unit is as follows. The current skeletal motion state vector It also needs to be processed through the linear rectification function layer and the Sigmoid function layer to obtain the corresponding first prediction weight vector z t (corresponding to the above formula 6). According to the all-one vector and z t , get the same as v t The corresponding second prediction weight vector 1-z t , and the corresponding first prediction weight vector z t The result of the dot product plus v t The corresponding second prediction weight vector 1-z t The result of the dot product is the output of the skeletal motion state vector of the next moment at the current moment. (corresponding to the above formula 7).

[0144] In the above-described embodiment, the skeleton motion hidden state vector of the last moment is to carry out motion intention extraction to skeleton motion characteristic vector at the last moment and obtain, so describes motion intention to a certain extent.Decode according to the skeleton motion hidden state vector of the last moment describing motion intention and the skeleton motion state vector of the current moment, obtain the estimated speed characteristic vector of the current moment, the estimated speed characteristic vector of the current moment that should be obtained, then more accurately embodies the information of skeleton motion on motion speed dimension, so, based on the estimated speed characteristic vector of the current moment and the skeleton motion state vector of the current moment, can more accurately calculate the skeleton motion state vector of next moment.

[0145] In one embodiment, the method also includes a machine learning model training step, which specifically includes the following steps: dividing the actually generated skeletal motion state vectors into historical samples and predicted samples according to the order in which the skeletal motion states are generated; training the machine learning model based on the skeletal motion state vectors in the historical samples, and outputting the skeletal motion state vectors expressed by the model parameters of the machine learning model; constructing a loss function based on the skeletal motion state vectors expressed by the model parameters and the skeletal motion state vectors in the predicted samples; and using the model parameters when the loss function takes the minimum value as the stable model parameters of the machine learning model.

[0146] The actual skeletal motion state vector is the skeletal motion state vector that has actually been generated. The historical sample is the sample data of the skeletal motion state vector used in the training of the machine learning model. The predicted sample is the sample data of the skeletal motion state vector calculated based on the historical sample during the training of the machine learning model.

[0147] It can be understood that machine learning model training is a process of inputting known sample data into the machine learning model and continuously iterating and updating the model parameters until the model parameters are stable. Therefore, when historical samples are input into the machine learning model for model training, the output is the skeletal motion state vector expressed by the model parameters of the machine learning model.

[0148] The computer device can construct a loss function based on the skeletal motion state vector expressed by the model parameters and the skeletal motion state vector in the prediction sample. The loss function is used to represent the degree of difference between the skeletal motion state vector expressed by the model parameters and the corresponding skeletal motion state vector in the prediction sample. The computer device can minimize the constructed loss function and use the model parameters that achieve the minimum loss function as the stable model parameters of the machine learning model.

[0149] In the above embodiment, by constructing a loss function and determining the stable model parameters of the machine learning model, the accuracy of the machine learning model is improved, thereby making skeletal movement testing using the machine learning model more accurate.

[0150] In one embodiment, constructing a loss function based on the skeletal motion state vector expressed by the model parameters and the skeletal motion state vector in the prediction sample includes: splicing each pair of adjacent skeletal motion state vectors in the skeletal motion state vector expressed by the model parameters to obtain a first splicing vector; obtaining a first matrix based on the outer product of the first splicing vector and the vector obtained by the transposition of the first splicing vector; splicing each pair of adjacent skeletal motion state vectors in the prediction sample to obtain a second splicing vector; obtaining a second matrix based on the outer product of the vector obtained by the second splicing vector and the transposition of the second splicing vector; and obtaining a loss function based on the mean square error of the corresponding first matrix and the second matrix.

[0151] It can be understood that the splicing vector is the vector obtained by splicing the vectors. For example, the adjacent bone motion state vectors expressed by the model parameters are and Perform splicing and get the first splicing vector as For example, the two adjacent skeletal motion state vectors x in the prediction sample are t and x t-1 Splice and get the second splicing vector [x t ;x t-1 ].

[0152] In one embodiment, the first matrix and the second matrix may be Gram matrices.

[0153] In one embodiment, the computer device may obtain the first matrix and the second matrix according to the following formulas:

[0154]

[0155] G(x t ,x t-1 )=[x t ;x t-1 ][x t ;x t-1 ] T ;

[0156] Where G() represents the Gram matrix, represents the first matrix; and represents the adjacent skeleton motion state vector expressed by the model parameters; represents the first splicing vector; is the vector obtained by transposing the first concatenated vector; G(x t ,x t-1 ) represents the second matrix; x t and x t-1Represents the adjacent skeletal motion state vector x in the predicted sample t and x t-1 ;[x t ;x t-1 ] represents the second splicing vector; [x t ;x t-1 ] T The vector obtained by transposing the second concatenated vector.

[0157] In one embodiment, the computer device may obtain the loss function according to the following formula:

[0158]

[0159] Among them, L gram represents the loss function; N is the number of first matrices or the number of second matrices; Represents the first matrix; G(x t ,x t-1 ) represents the second matrix. It can be understood that the number of the first matrix is ​​the same as the number of the second matrix.

[0160] It should be noted that in the embodiments of the present application, the t in the formula involved in the machine learning model training process has nothing to do with the time t when the skeletal motion prediction process is actually performed. It is necessary to distinguish the time t during model training from the time t during skeletal motion prediction based on the pre-trained model.

[0161] In the above embodiment, the first splicing vector is obtained by splicing the two adjacent skeletal motion state vectors in the skeletal motion state vectors expressed by the model parameters; the first matrix is ​​obtained by taking the outer product of the first splicing vector and the vector obtained by the transposition of the first splicing vector; the second splicing vector is obtained by splicing the two adjacent skeletal motion state vectors in the prediction sample; the second matrix is ​​obtained by taking the outer product of the vector obtained by the transposition of the second splicing vector and the second splicing vector; and the loss function is constructed based on the mean square error of the corresponding first matrix and the second matrix. That is, the loss function is constructed based on the adjacent skeletal features to perform machine learning model training, thereby taking into account the degree of continuity between bones in the machine learning model training and improving the accuracy of the machine learning model.

[0162] In one embodiment, the method further includes: obtaining a behavior label vector corresponding to the skeletal movement implicit state vector of the previous moment; the behavior label vector is obtained by encoding the behavior label. In this embodiment, decoding according to the skeletal movement implicit state vector of the previous moment and the skeletal movement state vector of the current moment to calculate the skeletal movement state vector of the next moment includes: splicing the behavior label vector to the corresponding skeletal movement implicit state vector of the previous moment; decoding according to the spliced ​​skeletal movement implicit state vector of the previous moment and the skeletal movement state vector of the current moment to calculate the skeletal movement state vector of the next moment; the skeletal movement state represented by the calculated skeletal movement state vector of the next moment matches the behavior label.

[0163] The behavior label vector, or the vector representation of the behavior label, is the vector obtained by encoding the behavior label. Behavior is a general term for all actions performed. Behavior labels include labels for behaviors such as walking, sitting, and pointing. A behavior label is not limited to identifying a single behavior; it can also identify combined behaviors, such as walking and pointing.

[0164] In one embodiment, the computer device may perform one-hot encoding on the behavior label to obtain a behavior label vector. One-hot encoding is a binary encoding, and the resulting vector elements are either 1 or 0. For example, the encoding of the walking label is [0, 0, 1], and the encoding of the pointing label is [0, 1, 0].

[0165] Specifically, computer equipment can pre-store the behavior label vectors for each moment setting. Will be appreciated that, as mentioned above, in carrying out the skeletal motion prediction processing process, at each current moment output skeletal motion hidden state vector, can be used as the skeletal motion hidden state of the so-called previous moment in the next skeletal motion prediction process of input, so the behavior label vector and the skeletal motion hidden state corresponding to the moment setting are also corresponding. Will be appreciated that the behavior label vectors corresponding to different skeletal motion hidden states can be different.

[0166] It should be noted that the computer device can only configure behavior tags for moments when behavior transitions occur. The behavior tags corresponding to other moments without behavior tags are the behavior tags corresponding to the moments with the most recent behavior tags.

[0167] The computer device can obtain the behavior label vector corresponding to the skeletal motion implicit state vector at the previous moment, and splice the behavior label vector to the corresponding skeletal motion implicit state vector at the previous moment; decode the spliced ​​skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment to calculate the skeletal motion state vector at the next moment. For example, the behavior label is encoded as [0,0,1], which is the same as the skeletal motion implicit state vector h at the previous moment. t-1 After concatenation with [0,0,1], the skeletal motion implicit state vector at the previous moment can be [h t-1 , 0,0,1]. It should be noted that this is not limited to the front and back of the splicing position.

[0168] Can be understood that, due to splicing behavior label vector in the skeleton movement hidden state vector of the last moment after splicing, therefore carried the behavior information specified by behavior label in the skeleton movement hidden state vector of the last moment after splicing, so, based on the skeleton movement hidden state vector of the last moment after splicing and the skeleton movement state vector of the current moment, decode, the skeleton movement state represented by the skeleton movement state vector of the next moment predicted matches the behavior specified by behavior label to a great extent. Therefore, by configuring behavior label, can control and realize specified behavior action.

[0169] In one embodiment, the method further includes: generating a control instruction for the target object based on the calculated skeletal motion state vector that matches the behavior label; the control instruction is used to instruct the target object to perform corresponding movements according to the calculated skeletal motion state vector to execute the behavior represented by the behavior label.

[0170] Specifically, the computer device generates a control instruction for the target object based on the calculated skeletal motion state vector that matches the behavior tag. It can be understood that the control instruction is used to instruct the target object to perform corresponding movements according to the calculated skeletal motion state vector to perform the behavior represented by the behavior tag.

[0171] The computer device can output control instructions for the target object to the target object, and the target object can perform corresponding movements according to the calculated skeletal motion state vector. The process of the target object performing corresponding movements according to the calculated skeletal motion state vector is the process of executing the behavior represented by the behavior label.

[0172] It can be understood that the behavior label vectors corresponding to different implicit states of skeletal movement may be different, so the control instructions output according to the calculated skeletal movement state vector that matches the behavior label are different, and different behavior labels can be used to control the target object to perform different behavioral actions.

[0173] Figure 6 FIG. 1 is a schematic diagram of controlling a target object through a behavior tag in one embodiment. Figure 6 As shown, the skeletal motion states in 602, 604, and 606 correspond to the pointing label, the walking label, and the behavior represented by the pointing label, respectively, that is, the series of behaviors of "pointing-walking-pointing". The skeletal motion states in 608, 610, and 612 correspond to the walking label, the sitting label, and the behavior represented by the walking label, respectively, that is, the series of behaviors of "walking-sit-walking".

[0174] In one embodiment, the method also includes a motion behavior prediction processing step, which specifically includes the following steps: predicting the motion behavior of the target object based on multiple continuous skeletal motion state vectors obtained by calculation; determining the interactive behavior logic that realizes corresponding interaction with the predicted motion behavior based on the predicted motion behavior; and interacting with the target object according to the interactive behavior logic.

[0175] Among them, sports behavior prediction is to judge in advance the sports behavior to be performed in the future.

[0176] It can be understood that since the prediction calculation is performed one by one according to the time dimension, the calculated skeletal motion state vector is continuous.

[0177] Specifically, the computer device can pre-judge the next movement behavior of the target object based on the skeletal motion state features represented by the calculated multiple continuous skeletal motion state vectors.

[0178] In one embodiment, a correspondence between the target object's motion behavior and the interactive behavior logic is pre-set in the computer device. The interactive behavior logic is used to implement corresponding interactions with the target object's motion behavior. Based on the predicted motion behavior, the interactive behavior logic that implements the corresponding interaction with the predicted motion behavior is determined; and interaction with the target object is performed according to the interactive behavior logic. For example, if the target object's motion behavior is "waving," the corresponding interactive behavior logic can be the logic for implementing the interactive action of "walking toward the target object."

[0179] In one embodiment, the computer device may be a robot. Figure 7 FIG. 1 is an application environment diagram of skeletal motion prediction processing in one embodiment. Figure 7, the robot 702 can observe the target object 704 (for example, a person), collect and obtain multiple continuous skeletal motion state image frames, and then extract the corresponding skeletal motion state vectors from the skeletal motion state image frames to obtain multiple continuous historical skeletal motion state vectors. The robot 702 can perform feature code encoding on each historical skeletal motion state vector through a machine learning model, generate skeletal motion feature vectors corresponding to each historical skeletal motion state vector, and predict and calculate the skeletal motion state vector corresponding to the target object 704 at the next moment according to the method provided in each embodiment of the present application. In one embodiment, the robot 702 can predict the motion behavior of the target object 704 based on the multiple continuous skeletal motion state vectors obtained by calculation; determine the interactive behavior logic that realizes the corresponding interaction with the predicted motion behavior based on the predicted motion behavior; and interact with the target object 704 according to the interactive behavior logic.

[0180] In the above embodiment, by calculating the skeletal motion state vector at the future moment, the motion behavior of the target object is predicted, and the corresponding interaction behavior logic is determined to interact with the target object, thereby improving the intelligence and flexibility of the interaction and improving the efficiency of human-computer interaction.

[0181] In one embodiment, Figure 8 As shown, in another embodiment, a method for predicting and processing skeletal motion is provided, and the method specifically includes the following steps:

[0182] S802, obtain multiple continuous historical skeletal motion state vectors; according to the sequence of each historical skeletal motion state vector, input the skeletal motion feature vector obtained by the previous encoding and the skeletal motion state vector to be encoded at the current time into the encoder in the pre-trained machine learning model for encoding, and output the skeletal motion feature vector obtained by the current encoding corresponding to the skeletal motion state vector to be encoded at the current time.

[0183] In one embodiment, the method also includes: obtaining multiple frames of continuously collected bone motion state image frames; for each frame of the bone motion state image frame, identifying multiple target bone joint points from the bone motion state image frame; according to the front-to-back order between the multiple target bone joint points, respectively obtaining the rotation data of the subsequent target bone joint point relative to the previous target bone joint point; using each rotation data as a vector element, and splicing them according to the front-to-back order between the corresponding target bone joint points to obtain a bone motion state vector corresponding to the bone motion state image frame.

[0184] S804, obtain the estimated speed feature vector of the previous moment before the current moment; input the estimated speed feature vector of the previous moment and each bone motion feature vector into the attention model in the machine learning model, and determine the motion correlation between each bone motion feature vector and the estimated speed feature vector of the previous moment according to the attention model.

[0185] The estimated velocity feature vector is used to characterize the changes between the estimated skeletal motion state vectors at adjacent moments.

[0186] S806, determining the weight of each skeletal motion feature vector according to the motion correlation; performing weighted summation on each skeletal motion feature vector according to the corresponding weight to obtain the skeletal motion implicit state vector at the previous moment of the current moment.

[0187] Among them, the weight of each skeletal motion feature vector is positively correlated with the corresponding motion correlation.

[0188] It should be noted that the processing steps of the skeletal motion implicit state vector of the previous moment in the determination of the current moment described in steps S804~806 can be applicable to non-first current moment. For the first current moment, the skeletal motion implicit state vector of the previous moment in the current moment can be directly set to the initial default value.

[0189] In other embodiments, for the first current moment, the skeletal motion implicit state vector of the previous moment of the current moment can also be calculated according to steps S804 to 806. Specifically, in step S804, a velocity feature analysis can be performed based on the skeletal motion state vector of the previous moment of the current moment and the skeletal motion state vector of the previous moment, and an estimated velocity feature vector can be estimated, thereby calculating the skeletal motion implicit state vector of the previous moment of the current moment according to steps S804 to 806.

[0190] S808 , obtaining a behavior label vector corresponding to the skeletal motion implicit state vector at the previous moment; and concatenating the behavior label vector to the corresponding skeletal motion implicit state vector at the previous moment.

[0191] The behavior label vector is obtained by encoding the behavior label.

[0192] S810, obtain the skeleton motion state vector at the current moment; input the spliced ​​skeleton motion implicit state vector of the previous moment and the skeleton motion state vector of the current moment into the decoder in the machine learning model for decoding to obtain the estimated speed feature vector at the current moment.

[0193] S812, obtaining a first prediction weight vector corresponding to the skeletal motion state vector at the current moment through a decoder; and determining a second prediction weight vector of the estimated velocity feature vector at the current moment based on the first prediction weight vector.

[0194] The sum of the first prediction weight vector and the second prediction weight vector is an all-ones vector.

[0195] S814, the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are respectively multiplied with the corresponding first prediction weight vector and the second prediction weight vector and then added to obtain the skeletal motion state vector at the next moment; the skeletal motion state represented by the obtained skeletal motion state vector at the next moment matches the behavior label.

[0196] The above-mentioned skeletal motion prediction processing method performs feature coding on a plurality of continuous historical skeletal motion state vectors respectively to obtain corresponding skeletal motion feature vectors, thereby realizing the extraction of skeletal motion features of each skeletal motion state vector. Determine the skeletal motion implicit state vector of the previous moment of the current moment; The skeletal motion implicit state vector is obtained by extracting the motion intention of the skeletal motion feature vector obtained at the previous moment, thereby realizing the extraction of motion intention according to the extracted skeletal motion features. Decode the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to calculate the skeletal motion state vector of the next moment. Decode the extracted motion intention and the current skeletal motion state vector, thereby realizing the mining of the extracted intention information, thereby calculating the skeletal motion state vector of the next moment based on the decoded mining intention information. Realize the motion prediction processing of this detail of the prediction of the skeletal motion of the target object itself.

[0197] like Figure 9 As shown, in one embodiment, a limb movement prediction processing method is provided, which specifically includes the following steps:

[0198] S902: Acquire multiple continuous historical limb movement states.

[0199] The limb movement state is the state of the limb during movement. It is understood that the limb movement state can include dynamic and static states. The historical limb movement state is the limb movement state that has already occurred.

[0200] In one embodiment, limb movement state comprises skeletal movement state vector.Wherein, skeletal movement state vector is the vector representation of skeletal movement state, promptly describes skeletal movement state by the form of vector.Be understandable, because limb movement state can be represented by skeletal movement state, and skeletal movement state can be represented by skeletal movement state vector, therefore, limb movement state can be represented by multiple continuous skeletal movement state vectors, promptly limb movement state can comprise skeletal movement state vector.

[0201] It is understandable that in other embodiments, the limb movement state can also be represented by limb curve data. For example, different degrees of bending represented by the limb curve data can indicate different limb movement states.

[0202] It is understood that each historical skeletal motion state vector may have a corresponding historical moment, which is the moment when the computer device observes and collects the skeletal motion state represented by the historical skeletal motion state vector.

[0203] S904 , extracting motion features from each historical limb motion state to obtain limb motion features corresponding to each historical limb motion state.

[0204] The limb motion features are used to characterize the features of the limbs during movement. It is understood that the computer device can extract motion features from each historical limb motion state to obtain limb motion features corresponding to each historical limb motion state.

[0205] In one embodiment, the limb motion feature may include a skeletal motion feature vector, which is a feature vector obtained by extracting motion features from a skeletal motion state vector.

[0206] It is understood that in other embodiments, limb movement features can also be obtained by extracting features from limb curve data from aspects such as bending degree or bending turning points.

[0207] S906: Determine the implicit state of the limb movement at the moment before the current moment.

[0208] The current moment is the moment at which limb movement prediction is being performed. The implicit limb movement state is used to integrate historical limb movement feature information to express the corresponding movement intention. Movement intention is the intention to perform a certain movement and is used to reflect the next desired movement. The implicit limb movement state at the moment before the current moment is obtained by extracting movement intention from the limb movement features corresponding to the historical limb movement state at that moment.

[0209] In one embodiment, the limb movement implicit state may include a skeletal movement implicit state vector, which is a vector obtained by extracting movement intention from a skeletal movement feature vector and is used to integrate historical skeletal movement state feature information to express corresponding movement intention.

[0210] It can be understood that since the limb movement feature can include the skeletal movement feature vector, then the limb movement feature is subjected to motion intention extraction, and the skeletal movement feature vector can be subjected to motion intention extraction. Then, the skeletal movement implicit state vector obtained by performing motion intention extraction on the skeletal movement feature vector can be used to represent the limb movement implicit state, that is, the limb movement implicit state can include the skeletal movement implicit state vector. Then, the limb movement implicit state of the previous moment at the current moment can include the skeletal movement implicit state vector obtained by performing motion intention extraction on the skeletal movement feature vector at the previous moment, and the skeletal movement feature vector corresponds to the historical skeletal movement state vector.

[0211] S908: Obtain the current limb movement state.

[0212] It can be understood that the limb movement state at the first current moment (i.e., the historical moment corresponding to the last historical limb movement state is used as the current moment) is the last limb movement state observed in history. When the corresponding moment of the predicted limb movement state is used as the current moment, the limb movement state at that current moment is the calculated limb movement state.

[0213] In one embodiment, step S908 includes: obtaining the skeleton motion state vector at the current moment.

[0214] S910 , determining the limb movement state at the next moment according to the limb movement implicit state at the previous moment and the limb movement state at the current moment.

[0215] It can be understood that since the implicit state of limb movement at the previous moment is obtained by extracting the movement intention of historical limb movement features, the limb movement state at the next moment can be predicted and calculated based on the implicit state of limb movement at the previous moment obtained by extracting the movement intention of historical limb movement features and decoding it with the limb movement state at the current moment.

[0216] In one embodiment, step S910 includes: calculating the skeletal motion state vector at the next moment based on the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment, so as to determine the limb motion state at the next moment based on the calculated skeletal motion state vector.

[0217] The above-mentioned limb movement prediction processing method extracts movement features from multiple continuous historical limb movement states to obtain corresponding limb movement features. The implicit state of limb movement at the previous moment before the current moment is determined; the implicit state of limb movement is obtained by extracting movement intention from the limb movement features obtained at the previous moment, thereby realizing the extraction of movement intention based on the extracted limb movement features. Decoding is performed based on the implicit state of limb movement at the previous moment and the limb movement state at the current moment to calculate the limb movement state at the next moment. By decoding the extracted movement intention and the limb movement state at the current moment, the extracted intention information is mined, thereby calculating the limb movement state at the next moment based on the decoded and mined intention information. A more detailed motion prediction process of predicting the limb movement of the target object itself is realized.

[0218] In one embodiment, the limb movement state includes a skeletal movement state vector; the limb movement feature includes a skeletal movement feature vector; and the limb movement implicit state includes a skeletal movement implicit state vector. Step S904 includes: feature coding each historical skeletal movement state vector using a machine learning model to generate skeletal movement feature vectors corresponding to each historical skeletal movement state vector. Step S910 includes: decoding the skeletal movement implicit state vector at the previous moment and the skeletal movement state vector at the current moment, calculating the skeletal movement state vector at the next moment, and determining the limb movement state at the next moment based on the calculated skeletal movement state vector.

[0219] In one embodiment, the limb motion state includes a skeletal motion state vector; the limb motion feature includes a skeletal motion feature vector; and the limb motion implicit state includes a skeletal motion implicit state vector. Step S906 includes: obtaining an estimated velocity feature vector at the previous moment before the current moment; the estimated velocity feature vector is used to characterize the change between the estimated skeletal motion state vectors at adjacent moments; determining the motion correlation between each skeletal motion feature vector and the estimated velocity feature vector at the previous moment; determining the weight of each skeletal motion feature vector based on the motion correlation; the weight is positively correlated with the motion correlation; and performing weighted summation on each skeletal motion feature vector according to the corresponding weight to obtain the skeletal motion implicit state vector at the previous moment before the current moment.

[0220] In one embodiment, respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment includes: inputting the estimated velocity feature vector at the previous moment and each bone motion feature vector into an attention model in a machine learning model, and respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment according to the attention model.

[0221] In one embodiment, decoding the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment to calculate the skeletal motion state vector at the next moment includes:

[0222] The skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment are input into the decoder in the machine learning model for decoding to obtain the estimated velocity feature vector at the current moment;

[0223] The skeletal motion state vector at the next moment is calculated based on the estimated velocity feature vector at the current moment and the skeletal motion state vector at the current moment.

[0224] In one embodiment, the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment are input into a decoder in a machine learning model for decoding, and the estimated velocity feature vector at the current moment is obtained, including:

[0225] The estimated velocity feature vector at the current moment is determined according to the following formula:

[0226]

[0227] Among them, t is the current moment, t-1 is the moment before the current moment; v t is the estimated velocity feature vector at the current moment; h t-1 is the skeletal motion implicit state vector of the previous moment; φ represents the linear rectification function; W v 、U vh 、b v and b vh These are bias parameters obtained by pre-training in the machine learning model; is the skeletal motion state vector at the current moment.

[0228] In one embodiment, based on the estimated velocity feature vector at the current moment and the skeletal motion state vector at the current moment, calculating the skeletal motion state vector at the next moment includes: obtaining a first prediction weight vector corresponding to the skeletal motion state vector at the current moment through a decoder; determining a second prediction weight vector of the estimated velocity feature vector at the current moment based on the first prediction weight vector; the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector; performing dot multiplication on the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment with the corresponding first prediction weight vector and the second prediction weight vector, and then adding them to obtain the skeletal motion state vector at the next moment.

[0229] In one embodiment, the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are respectively multiplied with the corresponding first prediction weight vector and the second prediction weight vector and then added to obtain the skeletal motion state vector at the next moment, which includes:

[0230] The skeletal motion state vector at the next moment is calculated according to the following formula:

[0231]

[0232]

[0233] Where t is the current time; z t is the first prediction weight vector; 1-z t is the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; v t is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx Both are bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeletal motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol.

[0234] In one embodiment, the method also includes a training step of a machine learning model, which specifically includes the following steps: dividing the actually generated skeletal motion state vectors into historical samples and predicted samples according to the order in which the skeletal motion state is generated; training the machine learning model based on the skeletal motion state vectors in the historical samples, and outputting the skeletal motion state vectors expressed by the model parameters of the machine learning model; constructing a loss function based on the skeletal motion state vectors expressed by the model parameters and the skeletal motion state vectors in the predicted samples; the loss function is used to represent the degree of difference between the skeletal motion state vectors expressed by the model parameters and the corresponding skeletal motion state vectors in the predicted samples; and taking the model parameters when the loss function takes the minimum value as the stable model parameters of the machine learning model.

[0235] In one embodiment, constructing a loss function based on the skeletal motion state vector expressed by the model parameters and the skeletal motion state vector in the prediction sample includes: splicing each pair of adjacent skeletal motion state vectors in the skeletal motion state vector expressed by the model parameters to obtain a first splicing vector; obtaining a first matrix based on the outer product of the first splicing vector and the vector obtained by the transposition of the first splicing vector; splicing each pair of adjacent skeletal motion state vectors in the prediction sample to obtain a second splicing vector; obtaining a second matrix based on the outer product of the vector obtained by the second splicing vector and the transposition of the second splicing vector; and obtaining a loss function based on the mean square error of the corresponding first matrix and the second matrix.

[0236] In one embodiment, the limb movement prediction processing method further includes: obtaining a behavior label vector corresponding to the skeletal movement implicit state vector of the previous moment; the behavior label vector is obtained by encoding the behavior label. Decoding the skeletal movement implicit state vector of the previous moment and the skeletal movement state vector of the current moment to calculate the skeletal movement state vector of the next moment includes: splicing the behavior label vector to the corresponding skeletal movement implicit state vector of the previous moment; decoding the spliced ​​skeletal movement implicit state vector of the previous moment and the skeletal movement state vector of the current moment to calculate the skeletal movement state vector of the next moment; the skeletal movement state represented by the calculated skeletal movement state vector of the next moment matches the behavior label.

[0237] In one embodiment, the limb motion prediction and processing method also includes: generating a control instruction for the target object based on the calculated skeletal motion state vector that matches the behavior label; the control instruction is used to instruct the target object to perform corresponding movements according to the calculated skeletal motion state vector to execute the behavior represented by the behavior label.

[0238] In one embodiment, the limb motion prediction and processing method also includes: predicting the motion behavior of the target object based on multiple continuous skeletal motion state vectors obtained by calculation; determining the interactive behavior logic that realizes corresponding interaction with the predicted motion behavior based on the predicted motion behavior; and interacting with the target object according to the interactive behavior logic.

[0239] In one embodiment, the limb movement prediction and processing method also includes: obtaining multiple frames of continuously collected skeletal movement state image frames; for each frame of the skeletal movement state image frame, identifying multiple target skeletal joint points from the skeletal movement state image frame; according to the front-to-back order between the multiple target skeletal joint points, respectively obtaining the rotation data of the subsequent target skeletal joint point relative to the previous target skeletal joint point; using each rotation data as a vector element, and splicing them according to the front-to-back order between the corresponding target skeletal joint points to obtain a skeletal movement state vector corresponding to the skeletal movement state image frame.

[0240] In one embodiment, each historical skeletal motion state vector is feature-coded through a machine learning model to generate skeletal motion feature vectors corresponding to each historical skeletal motion state vector, including: according to the sequence of each historical skeletal motion state vector, the skeletal motion feature vector obtained by the previous encoding and the skeletal motion state vector of the history to be encoded are input into the encoder in the pre-trained machine learning model for encoding, and the skeletal motion feature vector obtained by the current encoding corresponding to the historical skeletal motion state vector to be encoded is output.

[0241] like Figure 10 As shown, in one embodiment, a skeletal motion prediction processing device 1000 is provided, which includes: an encoding module 1002, a hidden state vector determination module 1004, a skeletal motion state vector acquisition module 1006 and a decoding prediction module 1008, wherein:

[0242] The encoding module 1002 is used to obtain multiple continuous historical skeletal motion state vectors; each of the skeletal motion state vectors is feature-encoded through a machine learning model to generate skeletal motion feature vectors corresponding to each of the skeletal motion state vectors.

[0243] The implicit state vector determination module 1004 is used to determine the skeletal motion implicit state vector at the previous moment. The skeletal motion implicit state vector at the previous moment is obtained by extracting the motion intention of the skeletal motion feature vector at the previous moment.

[0244] The skeleton motion state vector acquisition module 1006 is used to acquire the skeleton motion state vector at the current moment.

[0245] The decoding prediction module 1008 is used to decode the skeleton motion implicit state vector at the previous moment and the skeleton motion state vector at the current moment to calculate the skeleton motion state vector at the next moment.

[0246] In one embodiment, the implicit state vector determination module 1004 is also used to obtain the estimated speed feature vector of the previous moment before the current moment; the estimated speed feature vector is used to characterize the changes between the estimated bone motion state vectors of adjacent moments; the motion correlation between each bone motion feature vector and the estimated speed feature vector of the previous moment is determined respectively; the weight of each bone motion feature vector is determined according to the motion correlation; the weight is positively correlated with the motion correlation; each bone motion feature vector is weighted and summed according to the corresponding weight to obtain the bone motion implicit state vector of the previous moment before the current moment.

[0247] In one embodiment, the implicit state vector determination module 1004 is also used to input the estimated velocity feature vector of the previous moment and each bone motion feature vector into the attention model in the machine learning model, and determine the motion correlation between each bone motion feature vector and the estimated velocity feature vector of the previous moment according to the attention model.

[0248] In one embodiment, the decoding prediction module 1008 is also used to input the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment into the decoder in the machine learning model for decoding to obtain the estimated speed feature vector of the current moment; and calculate the skeletal motion state vector of the next moment based on the estimated speed feature vector of the current moment and the skeletal motion state vector of the current moment.

[0249] In one embodiment, the decoding prediction module 1008 is further configured to determine the estimated speed feature vector at the current moment according to the following formula:

[0250]

[0251] Among them, t is the current moment, t-1 is the moment before the current moment; v t is the estimated velocity feature vector at the current moment; h t-1 is the implicit state vector of the skeletal motion at the previous moment; φ represents the linear rectification function; W v 、U vh 、b v and b vh These are bias parameters obtained by pre-training in the machine learning model; is the skeletal motion state vector at the current moment.

[0252] In one embodiment, the decoding prediction module 1008 is also used to obtain a first prediction weight vector corresponding to the skeletal motion state vector at the current moment through a decoder; based on the first prediction weight vector, determine the second prediction weight vector of the estimated velocity feature vector at the current moment; the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector; the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are respectively multiplied with the corresponding first prediction weight vector and the second prediction weight vector and then added to obtain the skeletal motion state vector at the next moment.

[0253] In one embodiment, the decoding prediction module 1008 is further configured to calculate the skeletal motion state vector at the next moment according to the following formula:

[0254]

[0255]

[0256] Where t is the current time; z t is the first prediction weight vector; 1-z t is the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; v t is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx Both are bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeletal motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol.

[0257] like Figure 11 As shown, in one embodiment, the apparatus 1000 further includes:

[0258] The machine learning model training module 1001 is used to divide the actual bone motion state vectors into historical samples and predicted samples according to the order in which the bone motion states are generated; perform machine learning model training based on the bone motion state vectors in the historical samples, and output the bone motion state vectors expressed by the model parameters of the machine learning model; construct a loss function based on the bone motion state vectors expressed by the model parameters and the bone motion state vectors in the predicted samples; the loss function is used to represent the degree of difference between the bone motion state vectors expressed by the model parameters and the corresponding bone motion state vectors in the predicted samples; and the model parameters when the loss function takes the minimum value are used as the stable model parameters of the machine learning model.

[0259] In one embodiment, the machine learning model training module 1001 is also used to splice each pair of adjacent bone motion state vectors in the bone motion state vector expressed by the model parameters to obtain a first splicing vector; obtain a first matrix based on the outer product of the first splicing vector and the vector obtained by the transpose of the first splicing vector; splice the pair of adjacent bone motion state vectors in the prediction sample to obtain a second splicing vector; obtain a second matrix based on the outer product of the vector obtained by the second splicing vector and the transpose of the second splicing vector; and obtain a loss function based on the mean square error of the corresponding first matrix and the second matrix.

[0260] In one embodiment, the apparatus 1000 further includes:

[0261] A behavior label vector acquisition module (not shown) is used to acquire a behavior label vector corresponding to the skeletal motion implicit state vector at the previous moment; the behavior label vector is obtained by encoding the behavior label;

[0262] The decoding prediction module 1008 is also used to splice the behavior label vector to the corresponding skeletal movement implicit state vector of the previous moment; decoding is performed based on the spliced ​​skeletal movement implicit state vector of the previous moment and the skeletal movement state vector of the current moment to calculate the skeletal movement state vector of the next moment; the skeletal movement state represented by the calculated skeletal movement state vector of the next moment matches the behavior label.

[0263] In one embodiment, the apparatus 1000 further includes:

[0264] The motion control module (not shown) is used to generate control instructions for the target object based on the predicted skeletal motion state vector that matches the behavior label; the control instructions are used to instruct the target object to perform corresponding movements according to the predicted skeletal motion state vector to execute the behavior represented by the behavior label.

[0265] In one embodiment, the apparatus further comprises:

[0266] The interactive control module (not shown) is used to predict the motion behavior of the target object based on multiple continuous skeletal motion state vectors obtained by calculation; determine the interactive behavior logic that realizes corresponding interaction with the predicted motion behavior based on the predicted motion behavior; and interact with the target object according to the interactive behavior logic.

[0267] In one embodiment, the skeletal motion state vector acquisition module 1006 is also used to acquire multiple frames of continuously collected skeletal motion state image frames; for each frame of the skeletal motion state image frame, multiple target skeletal joint points are identified from the skeletal motion state image frame; according to the front-to-back order between the multiple target skeletal joint points, the rotation data of the subsequent target skeletal joint point relative to the previous target skeletal joint point are respectively acquired; each rotation data is used as a vector element, and according to the front-to-back order between the corresponding target skeletal joint points, the skeletal motion state vector corresponding to the skeletal motion state image frame is spliced.

[0268] In one embodiment, the encoding module 1002 is also used to input the skeletal motion feature vector obtained in the previous encoding and the skeletal motion state vector to be encoded at the current time into the encoder in the pre-trained machine learning model for encoding in the order of the historical skeletal motion state vectors, and output the skeletal motion feature vector obtained in the current encoding corresponding to the skeletal motion state vector to be encoded at the current time.

[0269] like Figure 12 As shown, in one embodiment, a limb motion prediction device 1200 is provided. The limb motion prediction device 1200 includes: a motion feature extraction module 1202, a hidden state determination module 1204, and a limb motion state determination module 1206, wherein:

[0270] The motion feature extraction module 1202 is used to obtain multiple continuous historical limb motion states; extract motion features from each of the historical limb motion states respectively, and obtain limb motion features corresponding to each of the historical limb motion states.

[0271] The implicit state determination module 1204 is used to determine the implicit state of the limb movement at the previous moment before the current moment; the implicit state of the limb movement at the previous moment is obtained by extracting the movement intention of the limb movement features at the previous moment.

[0272] The limb movement state determination module 1206 is used to obtain the limb movement state at the current moment; and determine the limb movement state at the next moment based on the limb movement implicit state at the previous moment and the limb movement state at the current moment.

[0273] In one embodiment, the limb movement state includes a skeletal movement state vector; the limb movement feature includes a skeletal movement feature vector; and the limb movement hidden state includes a skeletal movement hidden state vector. The movement feature extraction module 1202 is also used for encoding the skeletal movement state vectors of each history respectively by a machine learning model, and generating skeletal movement feature vectors corresponding to the skeletal movement state vectors of each history respectively. The limb movement state determination module 1206 is also used for decoding according to the skeletal movement hidden state vector of the previous moment and the skeletal movement state vector of the current moment, calculating the skeletal movement state vector of the next moment, and determining the limb movement state of the next moment according to the calculated skeletal movement state vector.

[0274] In one embodiment, the limb motion state includes a skeletal motion state vector; the limb motion feature includes a skeletal motion feature vector; and the limb motion implicit state includes a skeletal motion implicit state vector. The implicit state determination module 1204 is further configured to obtain an estimated velocity feature vector for the moment preceding the current moment; the estimated velocity feature vector is used to characterize the change between the estimated skeletal motion state vectors at adjacent moments; determine the motion correlation between each skeletal motion feature vector and the estimated velocity feature vector for the previous moment; determine the weight of each skeletal motion feature vector based on the motion correlation; the weight is positively correlated with the motion correlation; and perform weighted summation on each skeletal motion feature vector according to the corresponding weight to obtain the skeletal motion implicit state vector for the moment preceding the current moment.

[0275] In one embodiment, the implicit state determination module 1204 is also used to input the estimated speed feature vector of the previous moment and each bone motion feature vector into the attention model in the machine learning model, and determine the motion correlation between each bone motion feature vector and the estimated speed feature vector of the previous moment according to the attention model.

[0276] In one embodiment, the limb motion state determination module 1206 is also used to input the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment into the decoder in the machine learning model for decoding to obtain the estimated speed feature vector of the current moment; and calculate the skeletal motion state vector of the next moment based on the estimated speed feature vector of the current moment and the skeletal motion state vector of the current moment.

[0277] In one embodiment, the limb motion state determination module 1206 is further configured to determine the estimated velocity feature vector at the current moment according to the following formula:

[0278]

[0279] Among them, t is the current moment, t-1 is the moment before the current moment; v t is the estimated velocity feature vector at the current moment; h t-1 is the skeletal motion implicit state vector of the previous moment; φ represents the linear rectification function; W v 、U vh 、b v and b vh These are bias parameters obtained by pre-training in the machine learning model; is the skeletal motion state vector at the current moment.

[0280] In one embodiment, the limb motion state determination module 1206 is also used to obtain a first prediction weight vector corresponding to the skeletal motion state vector at the current moment through a decoder; determine a second prediction weight vector of the estimated velocity feature vector at the current moment based on the first prediction weight vector; the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector; the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are respectively multiplied with the corresponding first prediction weight vector and the second prediction weight vector and then added to obtain the skeletal motion state vector at the next moment.

[0281] In one embodiment, the limb motion state determination module 1206 is further configured to calculate the skeletal motion state vector at the next moment according to the following formula:

[0282]

[0283]

[0284] Where t is the current time; z t is the first prediction weight vector; 1-z t is the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; vt is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx Both are bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeletal motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol.

[0285] In one embodiment, the apparatus 1200 further includes:

[0286] A machine learning model training module (not shown) is used to divide the actual skeletal motion state vectors into historical samples and predicted samples according to the order in which the skeletal motion states are generated; to train the machine learning model based on the skeletal motion state vectors in the historical samples, and to output the skeletal motion state vectors expressed by the model parameters of the machine learning model; to construct a loss function based on the skeletal motion state vectors expressed by the model parameters and the skeletal motion state vectors in the predicted samples; the loss function is used to represent the degree of difference between the skeletal motion state vectors expressed by the model parameters and the corresponding skeletal motion state vectors in the predicted samples; and the model parameters when the loss function takes the minimum value are used as the stable model parameters of the machine learning model.

[0287] In one embodiment, the machine learning model training module is also used to splice each pair of adjacent bone motion state vectors in the bone motion state vector expressed by the model parameters to obtain a first splicing vector; obtain a first matrix based on the outer product of the first splicing vector and the vector obtained by the transpose of the first splicing vector; splice the pair of adjacent bone motion state vectors in the prediction sample to obtain a second splicing vector; obtain a second matrix based on the outer product of the vector obtained by the second splicing vector and the transpose of the second splicing vector; and obtain a loss function based on the mean square error of the corresponding first matrix and the second matrix.

[0288] In one embodiment, the apparatus 1200 further includes:

[0289] The behavior label vector acquisition module (not shown) is used to acquire the behavior label vector corresponding to the skeletal motion implicit state vector at the previous moment; the behavior label vector is obtained by encoding the behavior label.

[0290] The limb movement state determination module 1206 is also used to splice the behavior label vector to the corresponding skeletal movement implicit state vector of the previous moment; decode according to the spliced ​​skeletal movement implicit state vector of the previous moment and the skeletal movement state vector of the current moment to calculate the skeletal movement state vector of the next moment; the skeletal movement state represented by the calculated skeletal movement state vector of the next moment matches the behavior label.

[0291] In one embodiment, the apparatus 1200 further includes:

[0292] The motion control module (not shown) is used to generate control instructions for the target object based on the calculated skeletal motion state vector that matches the behavior label; the control instructions are used to instruct the target object to perform corresponding movements according to the calculated skeletal motion state vector to execute the behavior represented by the behavior label.

[0293] In one embodiment, the apparatus 1200 further includes:

[0294] The interactive control module (not shown) is used to predict the motion behavior of the target object based on multiple continuous skeletal motion state vectors obtained by calculation; determine the interactive behavior logic that realizes corresponding interaction with the predicted motion behavior based on the predicted motion behavior; and interact with the target object according to the interactive behavior logic.

[0295] In one embodiment, the motion feature extraction module 1202 is also used to obtain multiple frames of continuously collected skeletal motion state image frames; for each frame of the skeletal motion state image frame, multiple target skeletal joint points are identified from the skeletal motion state image frame; according to the front-to-back order between the multiple target skeletal joint points, the rotation data of the subsequent target skeletal joint point relative to the previous target skeletal joint point are respectively obtained; each rotation data is used as a vector element, and according to the front-to-back order between the corresponding target skeletal joint points, the skeletal motion state vector corresponding to the skeletal motion state image frame is spliced.

[0296] In one embodiment, the motion feature extraction module 1202 is also used to input the skeletal motion feature vector obtained in the previous encoding and the skeletal motion state vector of the history to be encoded at the current time into the encoder in the pre-trained machine learning model for encoding in the order of the historical skeletal motion state vectors, and output the skeletal motion feature vector obtained in the current encoding corresponding to the historical skeletal motion state vector to be encoded at the current time.

[0297] Figure 13 FIG. 1 is a schematic diagram of the internal structure of a computer device in one embodiment. Figure 13, the computer device can be a terminal or a server. The terminal can be a personal computer, a mobile terminal, a vehicle-mounted device or a robot, and the mobile terminal includes at least one of a mobile phone, a tablet computer, a personal digital assistant or a wearable device. The server can be implemented as an independent server or a server cluster composed of multiple physical servers. The computer device includes a processor, a memory and a network interface connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device can store an operating system and a computer program. When the computer program is executed, the processor can execute a skeletal motion prediction processing method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory can store a computer program, and when the computer program is executed by the processor, the processor can execute a skeletal motion prediction processing method. The network interface of the computer device is used for network communication.

[0298] Those skilled in the art will understand that Figure 13 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0299] In one embodiment, the skeletal motion prediction processing device provided by the present application can be implemented in the form of a computer program. Figure 13 The computer device shown in FIG. 1 is run on the computer device, and the non-volatile storage medium of the computer device can store various program modules constituting the skeletal motion prediction processing device, such as: Figure 10 The computer program composed of the various program modules is used to enable the computer device to execute the steps of the skeletal motion prediction processing method of each embodiment of the present application described in this specification. For example, the computer device can Figure 10The coding module 1002 in the skeletal motion prediction processing device 1000 shown obtains a plurality of continuous historical skeletal motion state vectors; each skeletal motion state vector is respectively encoded with a feature code through a machine learning model to generate a skeletal motion feature vector corresponding to each skeletal motion state vector. The computer device can determine the skeletal motion implicit state vector of the previous moment of the current moment through the implicit state vector determination module 1004; the skeletal motion implicit state vector of the previous moment is obtained by extracting the motion intention of the skeletal motion feature vector at the previous moment. The computer device can obtain the skeletal motion state vector of the current moment through the skeletal motion state vector acquisition module 1006, and decode according to the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment through the decoding prediction module 1008 to calculate the skeletal motion state vector of the next moment.

[0300] In one embodiment, the limb movement prediction processing device provided by the present application can be implemented in the form of a computer program. Figure 13 The computer device shown in FIG. 1 is run on the computer device, and the non-volatile storage medium of the computer device can store various program modules constituting the limb movement prediction processing device, such as: Figure 12 The motion feature extraction module 1202, the implicit state determination module 1204 and the limb motion state determination module 1206 are shown. The computer program composed of each program module is used to enable the computer device to execute the steps of the skeletal motion prediction processing method of each embodiment of the present application described in this specification. For example, the computer device can Figure 12 The motion feature extraction module 1202 in the limb motion prediction and processing device 1200 shown obtains multiple continuous historical limb motion states; motion features are extracted for each historical limb motion state to obtain limb motion features corresponding to each historical limb motion state. The computer device can determine the implicit limb motion state at the previous moment before the current moment through the implicit state determination module 1204; the implicit limb motion state at the previous moment is obtained by extracting the motion intention of the limb motion features at the previous moment. The computer device can obtain the limb motion state at the current moment through the limb motion state determination module 1206; and determine the limb motion state at the next moment based on the implicit limb motion state at the previous moment and the limb motion state at the current moment.

[0301] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor performs the following steps: obtaining multiple continuous historical bone motion state vectors; encoding each of the bone motion state vectors with a feature code through a machine learning model to generate a bone motion feature vector corresponding to each of the bone motion state vectors; determining the bone motion implicit state vector at the previous moment of the current moment; the bone motion implicit state vector at the previous moment is obtained by extracting the motion intention of the bone motion feature vector at the previous moment; obtaining the bone motion state vector at the current moment; decoding according to the bone motion implicit state vector at the previous moment and the bone motion state vector at the current moment to calculate the bone motion state vector at the next moment.

[0302] In one embodiment, determining the implicit state vector of the skeletal motion at the moment before the current moment includes: obtaining the estimated velocity feature vector at the moment before the current moment; the estimated velocity feature vector is used to characterize the changes between the estimated skeletal motion state vectors at adjacent moments; respectively determining the motion correlation between each skeletal motion feature vector and the estimated velocity feature vector at the previous moment; determining the weight of each skeletal motion feature vector based on the motion correlation; the weight is positively correlated with the motion correlation; and weightedly summing each skeletal motion feature vector according to the corresponding weight to obtain the implicit state vector of the skeletal motion at the moment before the current moment.

[0303] In one embodiment, respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment includes: inputting the estimated velocity feature vector at the previous moment and each bone motion feature vector into an attention model in a machine learning model, and respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment according to the attention model.

[0304] In one embodiment, decoding is performed based on the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to calculate the skeletal motion state vector of the next moment, including: inputting the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment into a decoder in a machine learning model for decoding to obtain an estimated speed feature vector of the current moment; calculating the skeletal motion state vector of the next moment based on the estimated speed feature vector of the current moment and the skeletal motion state vector of the current moment.

[0305] In one embodiment, the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment are input into a decoder in a machine learning model for decoding to obtain an estimated velocity feature vector at the current moment, including: determining the estimated velocity feature vector at the current moment according to the following formula:

[0306]

[0307] Among them, t is the current moment, t-1 is the moment before the current moment; v t is the estimated velocity feature vector at the current moment; h t-1 is the implicit state vector of the skeletal motion at the previous moment; φ represents the linear rectification function; W v 、U vh 、b v and b vh These are bias parameters obtained by pre-training in the machine learning model; is the skeletal motion state vector at the current moment.

[0308] In one embodiment, based on the estimated velocity feature vector at the current moment and the skeletal motion state vector at the current moment, calculating the skeletal motion state vector at the next moment includes: obtaining a first prediction weight vector corresponding to the skeletal motion state vector at the current moment through a decoder; determining a second prediction weight vector of the estimated velocity feature vector at the current moment based on the first prediction weight vector; the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector; performing dot multiplication on the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment with the corresponding first prediction weight vector and the second prediction weight vector, and then adding them to obtain the skeletal motion state vector at the next moment.

[0309] In one embodiment, the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are respectively multiplied with the corresponding first prediction weight vector and the second prediction weight vector and then added to obtain the predicted skeletal motion state vector at the next moment, including: calculating the skeletal motion state vector at the next moment according to the following formula:

[0310]

[0311]

[0312] Where t is the current time; z t is the first prediction weight vector; 1-z t is the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; v t is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx Both are bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeletal motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol.

[0313] In one embodiment, the computer program also causes the processor to perform the following steps: dividing the actually generated skeletal motion state vectors into historical samples and predicted samples according to the order in which the skeletal motion states are generated; training the machine learning model based on the skeletal motion state vectors in the historical samples, and outputting the skeletal motion state vectors expressed by the model parameters of the machine learning model; constructing a loss function based on the skeletal motion state vectors expressed by the model parameters and the skeletal motion state vectors in the predicted samples; the loss function is used to represent the degree of difference between the skeletal motion state vectors expressed by the model parameters and the corresponding skeletal motion state vectors in the predicted samples; and taking the model parameters when the loss function takes the minimum value as the stable model parameters of the machine learning model.

[0314] In one embodiment, constructing a loss function based on the skeletal motion state vector expressed by the model parameters and the skeletal motion state vector in the prediction sample includes: splicing each pair of adjacent skeletal motion state vectors in the skeletal motion state vector expressed by the model parameters to obtain a first splicing vector; obtaining a first matrix based on the outer product of the first splicing vector and the vector obtained by the transposition of the first splicing vector; splicing each pair of adjacent skeletal motion state vectors in the prediction sample to obtain a second splicing vector; obtaining a second matrix based on the outer product of the vector obtained by the second splicing vector and the transposition of the second splicing vector; and obtaining a loss function based on the mean square error of the corresponding first matrix and the second matrix.

[0315] In one embodiment, the computer program also enables the processor to perform the following steps: obtaining a behavior label vector corresponding to the skeletal movement implicit state vector at the previous moment; the behavior label vector is obtained by encoding the behavior label; decoding is performed according to the skeletal movement implicit state vector at the previous moment and the skeletal movement state vector at the current moment to calculate the skeletal movement state vector at the next moment, including: splicing the behavior label vector to the corresponding skeletal movement implicit state vector at the previous moment; decoding is performed according to the spliced ​​skeletal movement implicit state vector at the previous moment and the skeletal movement state vector at the current moment to calculate the skeletal movement state vector at the next moment; the skeletal movement state represented by the calculated skeletal movement state vector at the next moment matches the behavior label.

[0316] In one embodiment, the computer program also causes the processor to perform the following steps: generating a control instruction for the target object based on the predicted skeletal motion state vector that matches the behavior label; the control instruction is used to instruct the target object to perform corresponding movements according to the predicted skeletal motion state vector to execute the behavior represented by the behavior label.

[0317] In one embodiment, the computer program also causes the processor to perform the following steps: predicting the motion behavior of the target object based on the calculated multiple continuous skeletal motion state vectors; determining the interactive behavior logic that realizes corresponding interaction with the predicted motion behavior based on the predicted motion behavior; and interacting with the target object according to the interactive behavior logic.

[0318] In one embodiment, the computer program also enables the processor to perform the following steps: obtaining multiple frames of continuously collected skeletal motion state image frames; for each frame of the skeletal motion state image frame, identifying multiple target skeletal joint points from the skeletal motion state image frame; according to the front-to-back order between the multiple target skeletal joint points, respectively obtaining the rotation data of the subsequent target skeletal joint point relative to the previous target skeletal joint point; using each rotation data as a vector element, and splicing them according to the front-to-back order between the corresponding target skeletal joint points to obtain a skeletal motion state vector corresponding to the skeletal motion state image frame.

[0319] In one embodiment, each skeletal motion state vector is feature-coded by a machine learning model to generate skeletal motion feature vectors corresponding to each skeletal motion state vector, including: according to the sequence of each historical skeletal motion state vector, the skeletal motion feature vector obtained by the previous encoding and the skeletal motion state vector to be encoded at the current time are input into the encoder in the pre-trained machine learning model for encoding, and the skeletal motion feature vector obtained by the current encoding corresponding to the skeletal motion state vector to be encoded at the current time is output.

[0320] In one embodiment, a storage medium storing a computer program is provided. When the computer program is executed by a processor, the processor performs the following steps: obtaining multiple continuous historical bone motion state vectors; encoding each of the bone motion state vectors with a feature code through a machine learning model to generate a bone motion feature vector corresponding to each of the bone motion state vectors; determining the bone motion implicit state vector at the previous moment of the current moment; the bone motion implicit state vector at the previous moment is obtained by extracting the motion intention of the bone motion feature vector at the previous moment; obtaining the bone motion state vector at the current moment; decoding according to the bone motion implicit state vector at the previous moment and the bone motion state vector at the current moment to calculate the bone motion state vector at the next moment.

[0321] In one embodiment, determining the implicit state vector of the skeletal motion at the moment before the current moment includes: obtaining the estimated velocity feature vector at the moment before the current moment; the estimated velocity feature vector is used to characterize the changes between the estimated skeletal motion state vectors at adjacent moments; respectively determining the motion correlation between each skeletal motion feature vector and the estimated velocity feature vector at the previous moment; determining the weight of each skeletal motion feature vector based on the motion correlation; the weight is positively correlated with the motion correlation; and weightedly summing each skeletal motion feature vector according to the corresponding weight to obtain the implicit state vector of the skeletal motion at the moment before the current moment.

[0322] In one embodiment, respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment includes: inputting the estimated velocity feature vector at the previous moment and each bone motion feature vector into an attention model in a machine learning model, and respectively determining the motion correlation between each bone motion feature vector and the estimated velocity feature vector at the previous moment according to the attention model.

[0323] In one embodiment, decoding is performed based on the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to calculate the skeletal motion state vector of the next moment, including: inputting the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment into a decoder in a machine learning model for decoding to obtain an estimated speed feature vector of the current moment; calculating the skeletal motion state vector of the next moment based on the estimated speed feature vector of the current moment and the skeletal motion state vector of the current moment.

[0324] In one embodiment, the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment are input into a decoder in a machine learning model for decoding to obtain an estimated velocity feature vector at the current moment, including: determining the estimated velocity feature vector at the current moment according to the following formula:

[0325]

[0326] Among them, t is the current moment, t-1 is the moment before the current moment; v t is the estimated velocity feature vector at the current moment; h t-1 is the implicit state vector of the skeletal motion at the previous moment; φ represents the linear rectification function; W v 、U vh 、b v and b vh These are bias parameters obtained by pre-training in the machine learning model; is the skeletal motion state vector at the current moment.

[0327] In one embodiment, based on the estimated velocity feature vector at the current moment and the skeletal motion state vector at the current moment, calculating the skeletal motion state vector at the next moment includes: obtaining a first prediction weight vector corresponding to the skeletal motion state vector at the current moment through a decoder; determining a second prediction weight vector of the estimated velocity feature vector at the current moment based on the first prediction weight vector; the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector; performing dot multiplication on the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment with the corresponding first prediction weight vector and the second prediction weight vector, and then adding them to obtain the skeletal motion state vector at the next moment.

[0328] In one embodiment, the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are respectively multiplied with the corresponding first prediction weight vector and the second prediction weight vector and then added to obtain the predicted skeletal motion state vector at the next moment, including: calculating the skeletal motion state vector at the next moment according to the following formula:

[0329]

[0330]

[0331] Where t is the current time; z t is the first prediction weight vector; 1-z t is the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; v t is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx Both are bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeletal motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol.

[0332] In one embodiment, the computer program also causes the processor to perform the following steps: dividing the actually generated skeletal motion state vectors into historical samples and predicted samples according to the order in which the skeletal motion states are generated; training the machine learning model based on the skeletal motion state vectors in the historical samples, and outputting the skeletal motion state vectors expressed by the model parameters of the machine learning model; constructing a loss function based on the skeletal motion state vectors expressed by the model parameters and the skeletal motion state vectors in the predicted samples; the loss function is used to represent the degree of difference between the skeletal motion state vectors expressed by the model parameters and the corresponding skeletal motion state vectors in the predicted samples; and taking the model parameters when the loss function takes the minimum value as the stable model parameters of the machine learning model.

[0333] In one embodiment, constructing a loss function based on the skeletal motion state vector expressed by the model parameters and the skeletal motion state vector in the prediction sample includes: splicing each pair of adjacent skeletal motion state vectors in the skeletal motion state vector expressed by the model parameters to obtain a first splicing vector; obtaining a first matrix based on the outer product of the first splicing vector and the vector obtained by the transposition of the first splicing vector; splicing each pair of adjacent skeletal motion state vectors in the prediction sample to obtain a second splicing vector; obtaining a second matrix based on the outer product of the vector obtained by the second splicing vector and the transposition of the second splicing vector; and obtaining a loss function based on the mean square error of the corresponding first matrix and the second matrix.

[0334] In one embodiment, the computer program also enables the processor to perform the following steps: obtaining a behavior label vector corresponding to the skeletal movement implicit state vector at the previous moment; the behavior label vector is obtained by encoding the behavior label; decoding is performed according to the skeletal movement implicit state vector at the previous moment and the skeletal movement state vector at the current moment to calculate the skeletal movement state vector at the next moment, including: splicing the behavior label vector to the corresponding skeletal movement implicit state vector at the previous moment; decoding is performed according to the spliced ​​skeletal movement implicit state vector at the previous moment and the skeletal movement state vector at the current moment to calculate the skeletal movement state vector at the next moment; the skeletal movement state represented by the calculated skeletal movement state vector at the next moment matches the behavior label.

[0335] In one embodiment, the computer program also causes the processor to perform the following steps: generating a control instruction for the target object based on the predicted skeletal motion state vector that matches the behavior label; the control instruction is used to instruct the target object to perform corresponding movements according to the predicted skeletal motion state vector to execute the behavior represented by the behavior label.

[0336] In one embodiment, the computer program also causes the processor to perform the following steps: predicting the motion behavior of the target object based on the calculated multiple continuous skeletal motion state vectors; determining the interactive behavior logic that realizes corresponding interaction with the predicted motion behavior based on the predicted motion behavior; and interacting with the target object according to the interactive behavior logic.

[0337] In one embodiment, the computer program also enables the processor to perform the following steps: obtaining multiple frames of continuously collected skeletal motion state image frames; for each frame of the skeletal motion state image frame, identifying multiple target skeletal joint points from the skeletal motion state image frame; according to the front-to-back order between the multiple target skeletal joint points, respectively obtaining the rotation data of the subsequent target skeletal joint point relative to the previous target skeletal joint point; using each rotation data as a vector element, and splicing them according to the front-to-back order between the corresponding target skeletal joint points to obtain a skeletal motion state vector corresponding to the skeletal motion state image frame.

[0338] In one embodiment, each skeletal motion state vector is feature-coded by a machine learning model to generate skeletal motion feature vectors corresponding to each skeletal motion state vector, including: according to the sequence of each historical skeletal motion state vector, the skeletal motion feature vector obtained by the previous encoding and the skeletal motion state vector to be encoded at the current time are input into the encoder in the pre-trained machine learning model for encoding, and the skeletal motion feature vector obtained by the current encoding corresponding to the skeletal motion state vector to be encoded at the current time is output.

[0339] It should be understood that although each step in each embodiment of the present application is not necessarily performed in sequence according to the order indicated by the step number.Unless clear instructions are arranged in this article, the execution of these steps does not have strict order restriction, and these steps can be performed in other order.And, in each embodiment, at least a portion of steps can include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or the sub-steps of other steps or the stage.

[0340] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0341] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0342] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A method for predicting skeletal motion, comprising: Performing feature code encoding on a plurality of continuous historical skeletal motion state vectors respectively to obtain skeletal motion feature vectors corresponding to each skeletal motion state vector; Obtaining an estimated velocity feature vector of the moment before the current moment; the estimated velocity feature vector is used to represent the change between the estimated skeletal motion state vectors at adjacent moments; respectively determining the motion correlation between each of the skeletal motion feature vectors and the estimated velocity feature vector at the previous moment; Determining the weight of each of the skeletal motion feature vectors according to the motion correlation; Performing weighted summation on each of the skeletal motion feature vectors according to corresponding weights to obtain the skeletal motion implicit state vector at the previous moment; Decoding is performed based on the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to predict the skeletal motion state vector of the next moment; the skeletal motion state vector of the next moment is used to characterize the skeletal motion state of the next moment.

2. The method according to claim 1, characterized in that Each of the skeletal motion feature vectors is obtained by performing running feature encoding on a plurality of continuous historical skeletal motion state vectors through a machine learning model; and determining the motion correlation between each of the skeletal motion feature vectors and the estimated velocity feature vector at the previous moment includes: The estimated speed feature vector of the previous moment and each of the bone motion feature vectors are input into the attention model in the machine learning model, and the motion correlation between each bone motion feature vector and the estimated speed feature vector of the previous moment is determined respectively according to the attention model.

3. The method according to claim 1, characterized in that Decoding the skeleton motion implicit state vector at the previous moment and the skeleton motion state vector at the current moment to predict the skeleton motion state vector at the next moment includes: Decoding the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment to predict the estimated velocity feature vector at the current moment; The skeletal motion state vector at the next moment is predicted based on the estimated velocity feature vector at the current moment and the skeletal motion state vector at the current moment.

4. The method according to claim 3, characterized in that Decoding the estimated velocity feature vector at the current moment and the skeletal motion state vector at the current moment to predict the skeletal motion state vector at the next moment includes: Determine, by a decoder, a first prediction weight vector corresponding to the skeletal motion state vector at the current moment; determining a second prediction weight vector of the estimated speed feature vector at the current moment based on the first prediction weight vector; wherein the sum of the first prediction weight vector and the second prediction weight vector is an all-ones vector; The current moment's skeletal motion state vector and the current moment's estimated velocity feature vector are respectively multiplied with the corresponding first prediction weight vector and second prediction weight vector and then added to obtain the next moment's skeletal motion state vector.

5. The method according to claim 4, characterized in that Each of the skeletal motion feature vectors is obtained by performing running feature encoding on a plurality of continuous historical skeletal motion state vectors through a machine learning model; the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are respectively multiplied with the corresponding first prediction weight vector and the second prediction weight vector and then added to obtain the skeletal motion state vector at the next moment, comprising: The skeletal motion state vector at the next moment is calculated according to the following formula: Where t is the current time; z t is the first prediction weight vector; 1-z t is the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; v t is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx are all bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeletal motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol.

6. The method according to claim 1, characterized in that Also includes: The actual skeletal motion state vectors are divided into historical samples and predicted samples according to the order in which the skeletal motion states are generated; Performing machine learning model training based on the skeletal motion state vector in the historical samples, and outputting the skeletal motion state vector expressed by the model parameters of the machine learning model; Constructing a loss function based on the skeletal motion state vector expressed by the model parameters and the skeletal motion state vector in the predicted sample; the loss function is used to represent the degree of difference between the skeletal motion state vector expressed by the model parameters and the corresponding skeletal motion state vector in the predicted sample; The model parameters when the loss function takes the minimum value are regarded as the stable model parameters of the machine learning model.

7. The method according to claim 6, characterized in that The loss function constructed based on the skeletal motion state vector expressed by the model parameters and the skeletal motion state vector in the predicted sample includes: Splicing every two adjacent skeletal motion state vectors in the skeletal motion state vectors expressed by the model parameters to obtain a first splicing vector; Obtaining a first matrix according to the outer product of the first splicing vector and a vector obtained by transposing the first splicing vector; Splicing two adjacent skeletal motion state vectors in the prediction sample to obtain a second splicing vector; Obtain a second matrix according to the outer product of the second splicing vector and a vector obtained by transposing the second splicing vector; A loss function is obtained according to the corresponding mean square error of the first matrix and the second matrix.

8. The method according to any one of claims 1 to 7, characterized in that Also includes: Obtaining a behavior label vector corresponding to the skeletal motion implicit state vector at the previous moment; the behavior label vector is obtained by encoding the behavior label; The step of predicting the skeletal motion state vector at the next moment based on the skeletal motion implicit state vector at the previous moment and the skeletal motion state vector at the current moment includes: Splicing the behavior label vector to the corresponding skeletal motion implicit state vector at the previous moment; Decoding is performed based on the spliced ​​skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to calculate the skeletal motion state vector of the next moment; the skeletal motion state represented by the calculated skeletal motion state vector of the next moment matches the behavior label.

9. The method according to claim 8, characterized in that Also includes: generating a control instruction for a target object according to the calculated skeletal motion state vector that matches the behavior label; The control instruction is used to instruct the target object to perform corresponding movement according to the calculated skeletal motion state vector to execute the behavior represented by the behavior label.

10. The method according to any one of claims 1 to 7, characterized in that Also includes: Predict the target object's motion behavior based on the calculated multiple continuous skeletal motion state vectors; Determining, based on the predicted movement behavior, an interactive behavior logic for implementing corresponding interactions with the predicted movement behavior; Interact with the target object according to the interaction behavior logic.

11. The method according to any one of claims 1 to 7, characterized in that Also includes: Acquire multiple frames of continuously acquired skeletal motion state image frames; For each frame of the skeleton motion state image frame, identifying a plurality of target skeleton joint points from the skeleton motion state image frame; According to the front-to-back order of the plurality of target skeletal joint points, respectively obtain the rotation data of the subsequent target skeletal joint point relative to the previous target skeletal joint point; Each of the rotation data is used as a vector element, and is spliced ​​in a front-to-back order between the corresponding target bone joint points to obtain a bone motion state vector corresponding to the bone motion state image frame.

12. The method according to any one of claims 1 to 7, characterized in that The step of encoding the plurality of continuous historical skeletal motion state vectors with characteristic codes to obtain skeletal motion characteristic vectors corresponding to the respective skeletal motion state vectors comprises: In the process of using a machine learning model to perform feature encoding on multiple continuous historical skeletal motion state vectors, the skeletal motion feature vector obtained by the previous encoding and the skeletal motion state vector to be encoded at the current time are input into the encoder in the pre-trained machine learning model for encoding in the order of the historical skeletal motion state vectors, and the output is the skeletal motion feature vector obtained by the current encoding corresponding to the skeletal motion state vector to be encoded at the current time.

13. A method for predicting and processing limb movements, the method comprising: Extracting motion features from a plurality of continuous historical limb motion states to obtain limb motion features corresponding to the plurality of continuous historical limb motion states; The limb motion state includes a skeletal motion state vector; the limb motion feature includes a skeletal motion feature vector; Get the estimated speed feature vector of the previous moment before the current moment; Estimate Velocity feature vector, used to characterize the changes between the estimated skeletal motion state vectors at adjacent moments; Determine the motion correlation between each skeletal motion feature vector and the estimated velocity feature vector at the previous moment; Determine the weight of each skeletal motion feature vector according to the motion correlation, and perform weighted summation on each skeletal motion feature vector according to the corresponding weight to obtain the skeletal motion implicit state vector at the previous moment of the current moment; Get the current limb movement status; The limb movement state at the next moment is predicted based on the skeletal movement implicit state vector at the previous moment and the limb movement state at the current moment.

14. A skeletal motion prediction and processing device, characterized in that: The device comprises: An encoding module, configured to perform feature code encoding on a plurality of continuous historical skeletal motion state vectors respectively, to obtain skeletal motion feature vectors corresponding to each skeletal motion state vector; An implicit state vector determination module is configured to obtain an estimated velocity feature vector at a previous moment before the current moment; the estimated velocity feature vector is configured to characterize changes between estimated skeletal motion state vectors at adjacent moments; determine motion correlations between each of the skeletal motion feature vectors and the estimated velocity feature vector at the previous moment; determine weights of each of the skeletal motion feature vectors based on the motion correlations; and perform weighted summation of each of the skeletal motion feature vectors according to corresponding weights to obtain a skeletal motion implicit state vector at the previous moment; A skeleton motion state vector acquisition module is used to obtain the skeleton motion state vector at the current moment; The decoding prediction module is used to decode the skeleton motion implicit state vector at the previous moment and the skeleton motion state vector at the current moment, and predict the skeleton motion state vector at the next moment.

15. The skeletal motion prediction processing device according to claim 14, characterized in that: The machine learning model used by the encoding module is also used to run feature encoding; the implicit state vector determination module is also used to input the estimated speed feature vector of the previous moment and each of the bone motion feature vectors into the attention model in the machine learning model, and determine the motion correlation between each bone motion feature vector and the estimated speed feature vector of the previous moment according to the attention model.

16. The skeletal motion prediction and processing device according to claim 14, characterized in that: The decoding prediction module is also used to decode according to the skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment, and predict the estimated speed feature vector of the current moment; and predict the skeletal motion state vector of the next moment according to the estimated speed feature vector of the current moment and the skeletal motion state vector of the current moment.

17. The skeletal motion prediction processing device according to claim 16, characterized in that: The decoding prediction module is also used to determine the first prediction weight vector corresponding to the skeletal motion state vector at the current moment through the decoder; determine the second prediction weight vector of the estimated velocity feature vector at the current moment based on the first prediction weight vector; the sum of the first prediction weight vector and the second prediction weight vector is an all-one vector; the skeletal motion state vector at the current moment and the estimated velocity feature vector at the current moment are point multiplied with the corresponding first prediction weight vector and the second prediction weight vector, respectively, and then added to obtain the skeletal motion state vector at the next moment.

18. The skeletal motion prediction processing device according to claim 17, characterized in that: Each of the skeletal motion feature vectors is obtained by performing feature encoding on a plurality of continuous historical skeletal motion state vectors through a machine learning model; the decoding prediction module is further used to calculate the skeletal motion state vector at the next moment according to the following formula: Where t is the current time; z t is the first prediction weight vector; 1-z t is the second prediction weight vector; σ represents the Sigmoid function; φ represents the linear rectification function; is the skeleton motion state vector at the current moment; v t is the estimated velocity feature vector at the current moment; W z 、U zx 、b z and b zx are all bias parameters obtained by pre-training in the machine learning model; t+1 is the next moment after the current moment; is the skeletal motion state vector of the next moment from the current moment; ⊙ is the vector dot product symbol.

19. The skeletal motion prediction and processing device according to claim 14, wherein: The device also includes a model training module, which is used to divide the actual bone motion state vectors into historical samples and predicted samples according to the order in which the bone motion states are generated; perform machine learning model training based on the bone motion state vectors in the historical samples, and output the bone motion state vectors expressed by the model parameters of the machine learning model; construct a loss function based on the bone motion state vectors expressed by the model parameters and the bone motion state vectors in the predicted samples; the loss function is used to represent the degree of difference between the bone motion state vectors expressed by the model parameters and the corresponding bone motion state vectors in the predicted samples; and the model parameters when the loss function takes the minimum value are used as the stable model parameters of the machine learning model.

20. The skeletal motion prediction and processing device according to claim 19, wherein: The model training module is also used to splice every two adjacent bone motion state vectors in the bone motion state vector expressed by the model parameters to obtain a first splicing vector; obtain a first matrix based on the outer product of the first splicing vector and the vector obtained by the transposition of the first splicing vector; splice the two adjacent bone motion state vectors in the prediction sample to obtain a second splicing vector; obtain a second matrix based on the outer product of the second splicing vector and the vector obtained by the transposition of the second splicing vector; and obtain a loss function based on the mean square error of the corresponding first matrix and second matrix.

21. The skeletal motion prediction and processing device according to claim 14, wherein: The device further includes a behavior label processing module, the behavior label processing module being configured to obtain a behavior label vector corresponding to the skeletal motion implicit state vector at the previous moment; the behavior label vector being obtained by encoding the behavior label; and concatenating the behavior label vector to the corresponding skeletal motion implicit state vector at the previous moment; The decoding prediction module is also used to decode based on the spliced ​​skeletal motion implicit state vector of the previous moment and the skeletal motion state vector of the current moment to calculate the skeletal motion state vector of the next moment; the skeletal motion state represented by the calculated skeletal motion state vector of the next moment matches the behavior label.

22. The skeletal motion prediction and processing device according to claim 21, wherein: The device also includes a control instruction generation module, which is used to generate control instructions for the target object based on the calculated skeletal motion state vector that matches the behavior label; the control instruction is used to instruct the target object to perform corresponding movements according to the calculated skeletal motion state vector to execute the behavior represented by the behavior label.

23. The skeletal motion prediction and processing device according to claim 14, characterized in that: The device also includes an interactive control module, which is used to predict the movement behavior of the target object based on multiple continuous skeletal motion state vectors obtained by calculation; determine the interactive behavior logic that realizes corresponding interaction with the predicted movement behavior based on the predicted movement behavior; and interact with the target object according to the interactive behavior logic.

24. The skeletal motion prediction and processing device according to claim 14, characterized in that: The skeleton motion state vector acquisition module is also used to obtain multiple frames of continuously collected skeleton motion state image frames; for each frame of the skeleton motion state image frame, multiple target skeleton joint points are identified from the skeleton motion state image frame; according to the front-to-back order between the multiple target skeleton joint points, the rotation data of the subsequent target skeleton joint point relative to the previous target skeleton joint point are respectively obtained; each of the rotation data is used as a vector element, and according to the front-to-back order between the corresponding target skeleton joint points, the skeleton motion state vector corresponding to the skeleton motion state image frame is spliced.

25. The skeletal motion prediction and processing device according to claim 14, wherein: The encoding module is also used to input the bone motion feature vector obtained by the previous encoding and the bone motion state vector to be encoded at the current time into the encoder in the pre-trained machine learning model for encoding in the process of using the machine learning model to perform feature encoding on multiple continuous historical bone motion state vectors according to the sequence of the historical bone motion state vectors, and output the bone motion feature vector obtained by the current encoding corresponding to the bone motion state vector to be encoded at the current time.

26. A limb movement prediction and processing device, characterized in that: The device comprises: a motion feature extraction module, configured to extract motion features from a plurality of continuous historical limb motion states, respectively, to obtain limb motion features corresponding to the plurality of continuous historical limb motion states, wherein the limb motion states include skeletal motion state vectors; and the limb motion features include skeletal motion feature vectors; An implicit state determination module is used to obtain an estimated velocity feature vector at the moment before the current moment; the estimated velocity feature vector is used to characterize the change between the estimated skeletal motion state vectors at adjacent moments; the motion correlation between each skeletal motion feature vector and the estimated velocity feature vector at the previous moment is determined; the weight of each skeletal motion feature vector is determined based on the motion correlation, and each skeletal motion feature vector is weighted and summed according to the corresponding weight to obtain the skeletal motion implicit state vector at the moment before the current moment; The limb movement state determination module is used to obtain the limb movement state at the current moment; and predict the limb movement state at the next moment based on the skeletal movement implicit state vector at the previous moment and the limb movement state at the current moment.

27. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 13 is implemented.

28. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Passerby face detection and tracing algorithm based on video

    CN101216885A

  • Behavior identification method based on recurrent neural network and human skeleton movement sequences

    CN104615983A