Training trajectory prediction model and method and device for predicting motion trajectory of image element

By training a trajectory prediction model to generate editable intermediate animation products, the problem of uneditable animation in existing technologies is solved, and the accuracy and flexibility of animation generation are improved.

CN121120692APending Publication Date: 2025-12-12ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511189804.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Animations generated by existing technologies are usually uneditable, resulting in defects such as unnatural transitions between adjacent frames and distortion of image elements, which cannot be adjusted.

Method used

The training trajectory prediction model, including the generative network and the diffusion network, generates editable intermediate animation products by generating transformation parameters of the motion trajectory of image elements and performing denoising processing, and then uses software or tools to generate animations.

Benefits of technology

It improves the accuracy and flexibility of animation generation, enabling the editing of animations to improve transitions between adjacent frames and deformation of image elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120692A_ABST
    Figure CN121120692A_ABST
Patent Text Reader

Abstract

An embodiment of the present specification provides a method for training a trajectory prediction model, the trajectory prediction model comprising a generative network and a diffusion network, the method comprising: inputting a sample image layer extracted from a sample animation and a sample text describing an animation effect of an image element in the sample image layer into the generative network, and obtaining a plurality of regulation and control coefficients corresponding to the plurality of transformation parameters of the motion trail of the image element. And inputting the plurality of regulation and control coefficients and the noise parameter sequence into a diffusion network, and performing denoising processing on the noise parameter sequence based on the plurality of regulation and control coefficients to obtain a target parameter sequence. And determining the basic loss according to the first difference degree between the target parameter sequence and the real parameter sequence of the sample image layer. Parameters of the trajectory prediction model are adjusted according to the comprehensive loss, and the comprehensive loss comprises the basic loss.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One or more embodiments of the present specification relate to the technical field of artificial intelligence, and in particular, to a method and device for training a trajectory prediction model and predicting a motion trajectory of an image element. BACKGROUND

[0002] In order to improve user experience or competitiveness, many business products will add animations. In order to improve the generation efficiency of the animation, artificial intelligence technology is currently mainly used to generate the animation. However, the animation generated by the current artificial intelligence technology is usually not editable, which leads to the fact that when the generated animation has defects (such as unnatural transition between adjacent frames, image element deformation, etc.), it cannot be adjusted.

[0003] Therefore, there is an urgent need to provide a more effective animation generation scheme. SUMMARY

[0004] One or more embodiments of the present specification describe a method and device for training a trajectory prediction model and predicting a motion trajectory of an image element. The trained trajectory prediction model can generate an editable intermediate product (i.e., a series of transformation parameters describing the motion trajectory of an image element) of an animation, which can greatly improve the accuracy and flexibility of animation generation.

[0005] In a first aspect, a method for training a trajectory prediction model is provided. The trajectory prediction model includes a generation network and a diffusion network. The method includes:

[0006] inputting a sample image layer extracted from a sample animation and a sample text describing an animation effect of an image element in the sample image layer into the generation network to obtain a plurality of control coefficients corresponding to a plurality of transformation parameters describing a motion trajectory of the image element;

[0007] inputting the plurality of control coefficients and a noise parameter sequence into the diffusion network to perform denoising processing on the noise parameter sequence based on the plurality of control coefficients, to obtain a target parameter sequence; the noise parameter sequence includes a plurality of noise parameter groups of a plurality of key frames of the sample image layer, and a single noise parameter group includes a noisy parameter value of the plurality of transformation parameters corresponding to a key frame;

[0008] determining a basic loss according to a first difference between the target parameter sequence and a real parameter sequence of the sample image layer;

[0009] adjusting parameters of the trajectory prediction model according to a comprehensive loss; the comprehensive loss includes the basic loss.

[0010] In a second aspect, a method for predicting a motion trajectory of an image element is provided. The method includes:

[0011] obtain a trajectory prediction model trained according to the method of the first aspect;

[0012] input a target image layer and target text describing animation effects of target image elements in the target image layer into a generation network in the trajectory prediction model to obtain a plurality of first control coefficients;

[0013] input the plurality of first control coefficients and a noise data sequence into a diffusion network of the trajectory prediction model, so that the diffusion network performs denoising processing on the noise data sequence based on the plurality of first control coefficients to obtain a first data sequence used to determine motion trajectories of the target image elements, and the noise data sequence includes a plurality of noise data groups, and each noise data in a single noise data group corresponds to a first control coefficient.

[0014] In a third aspect, a device for training a trajectory prediction model is provided, the trajectory prediction model including a generation network and a diffusion network; and the device includes:

[0015] a generation unit configured to input a sample image layer extracted from a sample animation and sample text describing animation effects of image elements in the sample image layer into the generation network to obtain a plurality of control coefficients corresponding to a plurality of transformation parameters describing motion trajectories of the image elements;

[0016] a denoising unit configured to input the plurality of control coefficients and a noise parameter sequence into the diffusion network, so that the diffusion network performs denoising processing on the noise parameter sequence based on the plurality of control coefficients to obtain a target parameter sequence; and the noise parameter sequence includes a plurality of noise parameter groups of a plurality of key frames of the sample image layer, and a single noise parameter group includes noisy parameter values of the plurality of transformation parameters corresponding to a key frame;

[0017] a determination unit configured to determine a basic loss according to a first difference degree between the target parameter sequence and a real parameter sequence of the sample image layer;

[0018] an adjustment unit configured to adjust parameters of the trajectory prediction model according to a comprehensive loss; and the comprehensive loss includes the basic loss.

[0019] In a fourth aspect, a device for predicting motion trajectories of image elements is provided, and the device includes:

[0020] a obtaining unit configured to obtain a trajectory prediction model trained according to the method of the first aspect;

[0021] a generation unit configured to input a target image layer and target text describing animation effects of target image elements in the target image layer into a generation network in the trajectory prediction model to obtain a plurality of first control coefficients;

[0022] The denoising unit is configured to input the plurality of first regulation coefficients and a noise data sequence into a diffusion network of the trajectory prediction model, and perform denoising processing on the noise data sequence based on the plurality of first regulation coefficients, to obtain a first data sequence used to determine the motion trajectory of the target image element, wherein the noise data sequence includes a plurality of noise data groups, and each noise data in a single noise data group corresponds to a first regulation coefficient.

[0023] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed in a computer, the computer program causes the computer to perform the method of the first aspect.

[0024] In a sixth aspect, a computing device is provided, which includes a memory and a processor. The memory stores executable code, and the processor executes the executable code to implement the method of the first aspect.

[0025] The method for training a trajectory prediction model provided by one or more embodiments of the present specification first generates a plurality of regulation coefficients corresponding to a plurality of transformation parameters describing the motion trajectory of an image element based on a sample image layer and a sample text thereof through a generation network, and then performs denoising processing on a noise parameter sequence based on the plurality of regulation coefficients through a diffusion network to obtain a target parameter sequence. Finally, at least according to a basic loss determined based on the difference between the noise parameter sequence and the target parameter sequence, the parameters of the trajectory prediction model are adjusted. That is, in the present solution, a control condition (i.e., a plurality of regulation coefficients) for the diffusion network is generated based on multi-modal data such as images and texts, so that the accuracy of the control condition generation can be improved, and in turn the accuracy of the denoising processing result can be improved. In addition, the trajectory prediction model trained in the present solution can generate an editable intermediate product (i.e., a series of transformation parameters describing the motion trajectory of an image element) of an animation, so that the accuracy and flexibility of the animation generation can be greatly improved. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present specification, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can also be obtained by those skilled in the art without creative labor.

[0027] Figure 1 A schematic diagram of a trajectory prediction model shown in an example of the present specification;

[0028] Figure 2 A flowchart of a method for training a trajectory prediction model according to an embodiment of the present specification;

[0029] Figure 3 A schematic diagram of a feature extraction module is shown in one example of the present specification.

[0030] Figure 4 A schematic diagram of a method of predicting a motion trajectory of an image element is shown according to one embodiment of the present specification.

[0031] Figure 5 A schematic diagram of an apparatus of training a trajectory prediction model is shown according to one embodiment of the present specification.

[0032] Figure 6 A schematic diagram of an apparatus of predicting a motion trajectory of an image element is shown according to one embodiment of the present specification. DETAILED DESCRIPTION

[0033] The scheme provided in the present specification will be described below in conjunction with the accompanying drawings.

[0034] As mentioned before, the prior art has the problem that the generated animation cannot be edited. To this end, the present scheme proposes to train a trajectory prediction model, which can generate an editable intermediate product (i.e. a series of transformation parameters describing the motion trajectory of an image element) of the animation, and then generate the animation based on the intermediate product by using software or tools, so as to greatly improve the accuracy and flexibility of animation generation.

[0035] Figure 1 A schematic diagram of a trajectory prediction model is shown in one example of the present specification. Figure 1 In this embodiment, the trajectory prediction model can include a generation network and a diffusion network.

[0036] In one embodiment, the generation network described above can be implemented as a large language model.

[0037] In another embodiment, the generation network described above can include a feature extraction module, a linear layer and a gating module. The feature extraction module is configured to extract aggregated representations from the inputted multi-modal data. The linear layer is configured to perform linear transformation on the aggregated representations, thereby obtaining a linear transformation result. The gating module is configured to generate a plurality of independent gating signals according to a sample text in combination with the linear transformation result, and inject the gating signals into the diffusion network as control conditions.

[0038] Further, the generation network described above can further include a first image encoder and a text encoder, which are respectively configured to perform feature extraction on the image and the text in the multi-modal data.

[0039] In addition, the generation network can further include a second image encoder configured to extract features of the other image, and a representation fusion network configured to fuse the image representation extracted by the first image encoder and the image representation extracted by the second image encoder to obtain a target image representation, so that the feature extraction module can determine the aggregated representation based on the target image representation and the text representation.

[0040] It should be noted that in the process of training or using the trajectory prediction model, multi-step prediction is usually performed, so that the generation network can generate a plurality of control signals for the diffusion model iteratively.

[0041] The diffusion network can be implemented as a denoising diffusion probabilistic model (DDPM), a stable diffusion, or a model based on a Transformer architecture, i.e., a diffusion Transformer (DiT).

[0042] Taking the DiT as an example, the diffusion network can include a preset number of denoising modules arranged in series, and each denoising module can include a normalization layer configured to normalize the input received from the generation network according to the gating signal.

[0043] In one embodiment, each denoising module can include an attention layer, a feedforward neural network layer, and a normalization layer corresponding to each of the two layers. That is, each denoising module includes two normalization layers. Specifically, the normalization layer corresponding to the attention layer (hereinafter referred to as the first normalization layer) is configured to normalize the input of the diffusion network (or the output of the previous denoising module) according to the gating signal received from the generation network, and the attention layer is configured to process the input (i.e., the output of the first normalization layer) based on a multi-head attention mechanism. The normalization layer corresponding to the feedforward neural network (hereinafter referred to as the second normalization layer) is configured to normalize the sum of the output of the attention layer and the input of the diffusion network (or the output of the previous denoising module) according to the gating signal received from the generation network, and the feedforward neural network layer is configured to further process the output of the second normalization layer to obtain the output of the denoising module, which then enters the next denoising module. In this way, the single-step prediction of the diffusion network is completed.

[0044] It should be understood that after the above multi-step prediction is completed, the final denoising result can be obtained.

[0045] The following description is directed to Figure 1 The process of training the trajectory prediction model shown in the figure.

[0046] Figure 2 A flowchart of a method for training a trajectory prediction model is shown. It is noted that the method comprises multiple rounds of iterations, Figure 2 Steps of the method in which the k-th (k is a positive integer) round of iteration is included are shown. It is understood that multiple rounds of iterative updating of the trajectory prediction model can be achieved by repeatedly performing the steps shown therein.

[0047] As Figure 2 shown, the method can comprise the following steps:

[0048] In step S202, a sample image layer extracted from a sample animation, a sample text describing animation effects of image elements in the sample image layer are input into a generation network to obtain a plurality of control coefficients corresponding to a plurality of transformation parameters describing the motion trajectory of the image elements.

[0049] The sample animation can be obtained by sequentially superimposing a plurality of sample image layers. A single sample image layer can include, but is not limited to, any one of the following: an icon layer, a text layer, a foreground layer, a background layer, etc.

[0050] The description dimension of the sample text can include, but is not limited to, one or more of the following: elements (or objects), actions, directions, and ways, etc. For example, the sample text can be, for example, "left arm swings left and right", "concrete arms jump up and down", etc.

[0051] In one example, the plurality of transformation parameters can include a position parameter, a scaling parameter, a rotation parameter, and a transparency.

[0052] In a more specific example, the position parameter, the scaling parameter, and the rotation parameter are represented by three-dimensional coordinates (i.e., x, y, and z coordinates), so that 10 transformation parameters can be obtained: position x, position y, position z, scaling x, scaling y, scaling z, rotation x, rotation y, rotation z, and transparency.

[0053] Of course, in practice, the number of transformation parameters can be more or less, and the present specification does not limit this.

[0054] In one embodiment, the generation network can be implemented as a large language model, so that the target prompt word can be input into the large language model, and the large language model can output a plurality of control coefficients corresponding to a plurality of transformation parameters, and the target prompt word can include a sample image layer and a sample text.

[0055] In another embodiment, the generation network can include a feature extraction module, a linear layer and a gating module. The feature extraction module can be used to process the sample image layer and the sample text to obtain the target aggregated representation. The linear layer can be used to process the target aggregated representation to obtain the overall regulation coefficient shared by the plurality of transformation parameters. The gating module can be used to determine the plurality of target weights for the plurality of transformation parameters based on the sample text, and determine the plurality of regulation coefficients corresponding to the plurality of transformation parameters according to the product of the plurality of target weights and the overall regulation coefficient.

[0056] The feature extraction module is configured to fuse the multi-modal data.

[0057] In one embodiment, the feature extraction module can include an embedding layer, so that the sample text can be directly input into the feature extraction module for embedding processing to obtain the text representation.

[0058] In another embodiment, the sample text can also be processed by a text encoder to obtain the text representation. Then, the text representation is input into the feature extraction module.

[0059] The text encoder can be selected from CLIP (a model trained based on ViT architecture). In CLIP, an image encoder and a text encoder are included, and the image and the text are trained in the same vector space by contrastive learning to establish a corresponding relationship. Of course, the text encoder can also be implemented as BERT, etc.

[0060] As for the sample image layer, the embedding layer in the feature extraction module can be directly used for embedding processing, or a first image encoder can be used to extract features of the sample image layer to obtain an image representation V1. Then, the image representation V1 is input into the feature extraction module.

[0061] The first image encoder can be selected from CLIP, or implemented as a Visual Geometry Group (VGG) model or a Residual Network (ResNet) model, etc.

[0062] In the case where the first image encoder is selected from CLIP, since CLIP is actually a Transformer-based neural network model, it processes sequence data. Therefore, the sample image layer can be first divided into a plurality of image blocks. Then, the plurality of image blocks can be tiled to obtain the input sequence of the first image encoder. Of course, in practice, the position information of each image block can also be added to the input sequence to facilitate the first image encoder to distinguish image blocks at different positions.

[0063] In addition, in order to improve the accuracy of the control coefficient generated by the generation network, the initial state (such as the initial key frame) of the sample animation can also be input to the generation network, and at this time, the generation network can also include a second image encoder and a feature fusion network.

[0064] In one example, the second image encoder and the feature fusion network can both be implemented as a multi-layer perception machine. Of course, the second image encoder can also be implemented as a VGG model, etc.

[0065] Specifically, the initial key frame of the sample animation can be feature extracted by the second image encoder to obtain an image feature V2. Then, the combination of the image feature V1 and the image feature V2 can be processed by the feature fusion network to obtain a target feature V'.

[0066] The above feature extraction module can aggregate the image feature V1 (or the target feature V') and the text feature to obtain a target aggregated feature.

[0067] Figure 3 A feature extraction module diagram shown in one example of the present specification. Figure 3 In the above example, the feature extraction module can include a first attention layer, a second attention layer, and a third attention layer.

[0068] Specifically, in the first attention layer, the text feature T can be aggregated based on a self-attention mechanism to obtain a corresponding updated feature T'.

[0069] Wherein, the number of text features T is usually multiple, at this time, the first attention layer can be used to map each text feature T to a query vector, a key vector and a value vector respectively by using the query matrix, the key matrix and the value matrix of the first attention layer; then for any query vector, according to the query vector and each key vector, each weight corresponding to each value vector is obtained, and based on each weight and each value vector, the updated feature T' corresponding to the query vector is determined, so that each updated feature T' can be obtained.

[0070] Then, in the second attention layer, the cross-attention mechanism can be used to update the learnable target vector based on the image feature V1 (or the target feature V') to obtain an intermediate vector after fusing image information.

[0071] Wherein, the number of target vectors is usually multiple, and the number of image representations V1 (or target representations V') is also multiple, at this time, the query matrix of the second attention layer can be used to map each target vector to each query vector, and the key matrix and the value matrix of the second attention layer can be used to map each image representation V1 (or each target representation V') to each key vector and each value vector; then for any query vector, according to the query vector and each key vector, each weight corresponding to each value vector is obtained, and based on each weight and each value vector, the intermediate vector corresponding to the query vector is determined. In this way, each intermediate vector corresponding to each target vector can be obtained.

[0072] Finally, in the third attention layer, the intermediate vector can be updated based on the updated representation T' using the cross-attention mechanism to obtain the target aggregated representation that fuses image and text information.

[0073] That is, the query matrix of the third attention layer is used to map each intermediate vector to each query vector, and the key matrix and the value matrix of the second attention layer are used to map each image representation V1 (or each target representation V') to each key vector and each value vector; then for any query vector, according to the query vector and each key vector, each weight corresponding to each value vector is obtained, and based on each weight and each value vector, the final vector corresponding to the query vector is determined. In this way, each final vector corresponding to each intermediate vector can be obtained.

[0074] After that, the combination of each final vector can be directly used as the target aggregated representation, of course, each final vector can also be mapped and processed using a multi-layer perception, and the mapping result is used as the target aggregated representation.

[0075] It should be noted that when the image layer, the text and the initial key frame are processed using the feature extraction module including three attention layers, the precise semantic alignment of the three can be realized.

[0076] It should be understood that the above is only one implementation manner of the feature extraction module, which can also be implemented as CLIP, etc., and the present specification does not limit this.

[0077] The above is a description of the feature extraction module, and as for the linear layer, it can be implemented as a simple neural network model for linear transformation processing of the input target aggregated representation.

[0078] In practice, the number of linear layers can be adjusted according to the number of overall control coefficients to be generated. For example, when the number of overall control coefficients is two, and they are overall scaling coefficients and overall offset coefficients respectively, two linear layers can be set, one of which is used to generate the overall scaling coefficient, and the other is used to generate the overall offset coefficient.

[0079] Finally, regarding the above-mentioned gating module, it can determine a plurality of target weights for a plurality of transformation parameters according to the following formula:

[0080]

[0081] wherein x is a text representation of the sample text, which can be obtained by processing the sample text by using a text encoder, is a parameter of the gating module, is a target weight determined by the gating module for the kth transformation parameter, 1≤k≤K s , K s is the number of transformation parameters, and the softmax() function is a normalization function.

[0082] Taking the number of overall regulation coefficients as two, and the overall scaling coefficient and the overall offset coefficient as examples, for any transformation parameter, multiplying the target weight corresponding to the transformation parameter by the overall scaling coefficient can obtain the individual scaling coefficient corresponding to the transformation parameter, and multiplying the target weight corresponding to the transformation parameter by the overall offset coefficient can obtain the individual offset coefficient corresponding to the transformation parameter.

[0083] Step S204, inputting the plurality of regulation coefficients and the noise parameter sequence into the diffusion network, so as to perform denoising processing on the noise parameter sequence based on the plurality of regulation coefficients, to obtain a target parameter sequence.

[0084] The noise parameter sequence here can be obtained by adding noise to the original parameter sequence of the above-mentioned sample image layer. The original parameter sequence can include a plurality of original parameter groups of a plurality of key frames of the above-mentioned sample image layer, wherein each original parameter group includes parameter values of a plurality of transformation parameters corresponding to the key frame.

[0085] In one embodiment, the noise parameter sequence can be obtained by adding noise to each parameter value in the original parameter sequence respectively.

[0086] Taking adding noise to any parameter value (hereinafter referred to as original parameter value) as an example, the corresponding noise-added parameter value (hereinafter referred to as noise parameter value) can be determined according to the following formula:

[0087]

[0088] wherein t here is a time step randomly sampled for the sample image element, and its value range is [1, T], T is a preset total diffusion step number. x0 is the original parameter value, x t is the noise parameter value, ∈ is noise sampled from a standard Gaussian distribution, controls the degree of preservation of the original parameter value, controls the intensity of the noise. The noise parameter sequence can be calculated according to the following formula:

[0089]

[0090] wherein β s is a noise variance that increases with t and can be calculated by a predefined rule. Of course, in practice, β s may not be defined, and the calculation rule of a s may be directly defined.

[0091] In addition, in practice, noise can also be added to the original parameter value in other ways, which are not limited in this specification.

[0092] In summary, the above noise parameter sequence includes a plurality of noise parameter groups of a plurality of key frames of a sample image layer, and a single noise parameter group includes a plurality of noise parameter values of a plurality of transformation parameters corresponding to a key frame.

[0093] In addition to a plurality of control coefficients and a noise parameter sequence, the input of the diffusion network can also include a randomly sampled time step t, at this time, the time step t can be encoded (for example, the time step t is encoded by sine and cosine functions of different frequencies), and then the corresponding encoding result is fused with the noise parameter sequence to obtain a fused parameter sequence. Since the denoising process of the noise parameter sequence and the fused parameter sequence is similar, the above denoising process is described below by taking the noise parameter sequence as an example.

[0094] It should be noted that by inputting the time step t, the dependence relationship between the key frames can be effectively established.

[0095] Regarding the above diffusion network, it can be implemented as a denoising diffusion probabilistic model (DDPM), a stable diffusion, etc., or as a model based on a Transformer architecture, namely a diffusion Transformer (DiT for short).

[0096] Taking the diffusion network implemented as DiT as an example, it can include a preset number of denoising modules arranged in series, and a single denoising module can include a normalization layer for normalizing the input according to a plurality of control coefficients received from the generation network. In addition, the denoising module can also include other network layers.

[0097] In one embodiment, a single denoising module can include an attention layer, a feedforward neural network layer, and a normalization layer corresponding to the two layers, respectively.

[0098] The denoising processing can include inputting a feature matrix M1 corresponding to a plurality of regulation coefficients and noise parameter sequences into a first denoising module. In a normalization layer corresponding to an attention layer, each matrix element e1 in the feature matrix M1 can be normalized by row (i.e., each matrix element is subtracted by a row mean and then divided by a row standard deviation) to obtain each normalized matrix element e1'. The feature matrix M1 can be obtained by processing the noise parameter sequences using an embedding layer of a diffusion network, or can be obtained by processing the noise parameter sequences using an external feature extractor. Each row of the feature matrix M1 corresponds to a plurality of noise parameter groups, and each column corresponds to a transformation parameter. Then, for any matrix element in the normalized matrix element e1', an individual scaling coefficient corresponding to the column to which the matrix element belongs is multiplied by the matrix element, and the product is added to an individual offset coefficient corresponding to the column to which the matrix element belongs to obtain a target matrix element e2. Similarly, each target matrix element e2 corresponding to each matrix element e1 can be obtained, and each target matrix element e2 can form a target matrix M2.

[0099] In one example, each matrix element e1 can be normalized according to the following formula:

[0100]

[0101] wherein scale is an individual scaling coefficient corresponding to the column to which the matrix element e1 belongs, shift is an individual offset coefficient corresponding to the column to which the matrix element e1 belongs, μ is a mean of each matrix element in the row to which the matrix element e1 belongs (i.e., row mean), σ is a standard deviation of each matrix element in the row to which the matrix element e1 belongs (i.e., row standard deviation), is the normalized matrix element e1', and e2 is the target matrix element corresponding to the matrix element e1.

[0102] Of course, in practice, each matrix element e1 in the feature matrix M1 can not be normalized by row, but can be scaled and offset based on the corresponding individual scaling coefficient and individual offset coefficient to obtain the corresponding each target matrix element e2.

[0103] In addition, each matrix element e1 can be regulated based on only one individual regulation coefficient, which is not limited in the present specification.

[0104] After that, in the attention layer, the target matrix M2 can be aggregated based on the multi-head attention mechanism to obtain an updated matrix M2'. And in the normalization layer corresponding to the feedforward neural network, each matrix element e2 in the updated matrix M2' can be first normalized by row. Then, for any matrix element in the normalized matrix elements e2', the individual scaling coefficient corresponding to the column to which the matrix element belongs is multiplied by the matrix element, and the product is added to the individual offset coefficient corresponding to the column to which the matrix element belongs to obtain a target matrix element e3. Similarly, each target matrix element e3 corresponding to each matrix element e2 can be obtained, and each target matrix element e3 can form a target matrix M3.

[0105] Finally, in the feedforward neural network, the target matrix M3 is further processed to obtain a final matrix M4 as the output of the first denoising module.

[0106] It should be noted that the above target matrix M2, updated matrix M2' and target matrix M3 all belong to a feature matrix in an intermediate processing process.

[0107] After the first denoising module is processed, it can enter the second denoising module, and so on until the last denoising module is reached. It should be understood that the output of the last denoising module is the above target parameter sequence, which includes a series of predicted parameter values.

[0108] In addition, it should be understood that the above is only one implementation of the denoising module, and in practice, the denoising module can include more network layers, such as, it can also include an embedding layer. It can also include fewer network layers, such as, it does not include the feedforward neural network layer and the normalization layer corresponding thereto.

[0109] Step S206, determining a basic loss according to the difference degree between the target parameter sequence and the real parameter sequence of the sample image layer.

[0110] When the above noise parameter sequence is obtained by adding the noise corresponding to the time step t to each parameter value in the original parameter sequence (i.e., adding noise according to formula 2), the real parameter sequence here can be obtained by adding the noise corresponding to the time step t-1 to each parameter value in the original parameter sequence. That is, for a certain parameter value, replace t in the above formula 2 with t-1 to calculate the corresponding noisy parameter value (i.e., the real parameter value).

[0111] Specifically, the difference degree D1 between each target parameter group in the target parameter sequence and the corresponding real parameter group in the real parameter sequence can be calculated, and then the basic loss L m is determined as the comprehensive loss.

[0112] In one example, the above basic loss L can be calculated according to a mean square error loss function. m Of course, in practice, the mean square error loss function can also be replaced by an L1 norm, etc.

[0113] Additionally, the difference D2 can also be calculated for each adjacent two target parameter groups in the target parameter sequence, and the penalty term Ω1 can be determined through comparison of each difference D2. Then the basic loss L m can be superimposed with the penalty term Ω1 to obtain the comprehensive loss.

[0114] In one example, the above penalty term can be calculated according to a mean square error loss function.

[0115] It should be noted that since each noise parameter group in the noise parameter sequence corresponds to a key frame in the sample image layer, and the target parameter sequence is obtained by denoising the noise parameter sequence, each target parameter group in the target parameter sequence corresponds to a key frame in the sample image layer. Therefore, by adding the penalty term Ω1 to the basic loss L m , the smoothness of the motion trajectory predicted by the trajectory prediction model can be improved.

[0116] In addition, the penalty term Ω2 can also be determined according to the difference D3 between the first and last target parameter groups in the target parameter sequence, and the difference D4 between the first and last real parameter groups in the real parameter sequence. Then the basic loss L m can be superimposed with the penalty term Ω1 and the penalty term Ω2 to obtain the comprehensive loss.

[0117] It should be noted that in many cases, the animation is required to be played in a loop, and the present scheme can improve the closeness of the motion trajectory predicted by the trajectory prediction model by adding the penalty term Ω2 to the basic loss L m .

[0118] In summary, after the trajectory prediction model is trained in combination with the above penalty term Ω1 and the penalty term Ω2, the trajectory prediction model can generate a smooth and seamless loop animation.

[0119] In one example, the comprehensive loss can be calculated according to the following formula:

[0120]

[0121] Wherein, L SLL is the comprehensive loss, N is the number of target parameter groups or real parameter groups, y i is the i-th real parameter group, and are the i-th, i+1-th and i-1-th target parameter groups, y1 and y Nrespectively, are a first-end real parameter group and a last-end real parameter group, and respectively, are a first-end target parameter group and a last-end target parameter group. is a basic loss L m , and respectively, are a penalty term Ω1 and a penalty term Ω2, and λ and α are respectively used to control the intensity of the penalty term Ω1 and the penalty term Ω2.

[0122] At this point, the comprehensive loss corresponding to one sample image layer is obtained. Similarly, the comprehensive loss corresponding to each sample image layer in a batch of training sample set can be obtained.

[0123] It should be noted that in the present scheme, the sample text corresponding to the sample image layer can also be masked with a predetermined probability (for example, 0.4) for a batch of training sample set. It should be understood that after masking the sample text corresponding to a certain sample image layer, the above generation network generates the regulation coefficient based only on the sample image layer, and the gating module in the diffusion network determines the target weight of the transformation parameter based on the empty string (i.e. unconditionally), so that on the one hand, the robustness of the trajectory prediction model can be improved, that is, it does not rely too much on the data of a certain modality, and on the other hand, it can ensure that the diffusion network learns conditional and unconditional denoising processing at the same time.

[0124] Step S208, adjusting the parameters of the trajectory prediction model according to the comprehensive loss.

[0125] For example, the comprehensive loss corresponding to each sample image layer in a batch of training sample set can be summed to obtain a final loss. Then, based on the final loss, the update gradient corresponding to the parameters of the trajectory prediction model can be calculated using the back propagation method, and the parameters of the trajectory prediction model can be updated based on the update gradient to obtain the trained trajectory prediction model.

[0126] At this point, the kth round of iterative training of the trajectory prediction model is completed, and then the k+1th round of iteration can be entered, until the iteration end condition is reached. After that, the trajectory prediction model trained in the last round can be used as the final trajectory prediction model.

[0127] The use process of the trained trajectory prediction model is described below.

[0128] Figure 4 A method for predicting the motion trajectory of an image element according to an embodiment of the present specification is shown, which can be executed by any device, equipment, platform, device cluster with computing and processing capabilities. It should be noted that the method can include multiple steps of prediction, Figure 4 The method steps included in the tth step of prediction are shown. As Figure 4 shown, the method can include the following steps:

[0129] Firstly, the target image layer and the target text track describing the animation effect of the target image element in the target image layer are input into the generation network in the trajectory prediction model to obtain a plurality of regulation coefficients c1.

[0130] As described above, the generation network can include a feature extraction module, a linear layer, and a gating module, so that the target image layer and the target text can be processed by the feature extraction module to obtain a target aggregated representation. Then the linear layer is used to process the target aggregated representation to obtain the overall regulation coefficient shared by the plurality of transformation parameters. And the gating module can be used to determine a plurality of target weights based on the target text, and determine a plurality of regulation coefficients c1 corresponding to a plurality of transformation parameters according to the product of a plurality of target weights and the overall regulation coefficient.

[0131] In one example, the regulation coefficient c1 can be arranged in pairs, so that the plurality of regulation coefficients c1 can also be referred to as a plurality of regulation coefficient pairs. A single regulation coefficient pair can include an individual scaling coefficient and an individual offset coefficient. Wherein the individual scaling coefficient and the individual offset coefficient here can be obtained by processing the target aggregated representation using two linear layers respectively.

[0132] Of course, in practice, the regulation coefficient c1 can also be set to one, which is not limited in the present specification.

[0133] In addition, the input of the generation network can also include a predefined initial key frame. The processing process of the initial key frame can refer to the processing process of the initial key frame of the sample image layer in the training stage, which will not be repeated here.

[0134] Secondly, the plurality of regulation coefficients c1 and the noise data sequence are input into the diffusion network of the trajectory prediction model, so that the diffusion network performs denoising processing on the noise data sequence based on the plurality of regulation coefficients c1 to obtain a data sequence s1 used to determine the motion trajectory of the target image element.

[0135] Wherein, the noise data sequence here has the same dimension as the noise parameter sequence, that is, it includes the same number of noise data groups as the number of noise parameter groups. The number of noise data in a single noise data group is the same as the number of noise parameters in a single noise parameter group.

[0136] It should be understood that the single noise data group is used to determine a key frame of the target image layer, and the single noise data in the single noise data group is used to determine a transformation parameter of the corresponding key frame. That is, each noise data in the single noise data group corresponds to each transformation parameter.

[0137] It should be noted that when the t-th step prediction is the first step prediction, each noise data in the single noise data set can be randomly sampled from a predetermined distribution. The predetermined distribution may, for example, be a Gaussian distribution. When the t-th step prediction is not the first step prediction, the noise data sequence can be the output of the (t-1)-th step prediction.

[0138] In addition to the plurality of regulation coefficients c1 and the noise data sequence, the input of the trajectory prediction model can also include a time step t. When the time step t is also input, the time step t and the noise data sequence can be fused first to obtain a fused data sequence. Since the processing process for the fused data sequence is similar to the processing process for the noise data sequence, the corresponding denoising process will be described below by taking the noise data sequence as an example.

[0139] As described above, the diffusion network can include a preset number of denoising modules including a normalization layer arranged in series.

[0140] Taking the case where the regulation coefficients c1 are arranged in pairs as an example, in the normalization layer in any denoising module, each matrix element in a feature matrix corresponding to the noise data sequence is standardized by row to obtain a standardized matrix element. Each row of the feature matrix corresponds to a noise data set, and each column corresponds to a transformation parameter. For any matrix element of the standardized matrix element, the individual scaling coefficient corresponding to the column to which the matrix element belongs is multiplied by the matrix element, and the product is added to the individual offset coefficient corresponding to the column to obtain a target matrix element. Each target matrix element corresponding to each matrix element forms a target matrix, which is used to determine the output of the any denoising module.

[0141] Of course, the diffusion network can also include other network layers, such as an attention layer and a feedforward neural network layer. The processing processes of the other network layers are conventional processing processes, which will not be described herein.

[0142] It should be understood that the output of the last denoising module is the data sequence s1. It should be understood that the data sequence s1 herein is the output of the t-th step prediction.

[0143] In this scheme, the target image layer can also be input into the generation network to obtain a plurality of regulation coefficients c2. Then, the plurality of second regulation coefficients c2 and the noise data sequence are input into the diffusion network, so that the diffusion network performs denoising processing on the noise data sequence based on the plurality of regulation coefficients c2 to obtain a data sequence s2. Finally, the data sequence s1 and the data sequence s2 can be weighted and fused to obtain a final sequence. Here, the weight of the data sequence s1 is greater than that of the data sequence s2. According to the final sequence, the motion trajectory of the target image element is determined.

[0144] In one example, the final sequence can be determined according to the following formula:

[0145]

[0146] wherein z t is a noise data sequence, I is a target image layer, T is a target text, i is an initial key frame, ∈ θ (z t |I,T,i) is a data sequence s1, is a data sequence s2, (1+ω) and ω are fusion weights of the data sequence s1 and the data sequence s2 respectively, is a final sequence.

[0147] It should be understood that, in the case of determining the final sequence, the final sequence can be taken as an output of the t-th step prediction.

[0148] It should be noted that, since in the training phase, the scheme will mask part of the sample text of the sample image layer with a predetermined probability, so that the model can learn an unconditional denoising method. Therefore, in the inference prediction, an unconditional denoising result (i.e. the data sequence s2) can be obtained, and the scheme fuses the conditional denoising result (i.e. the data sequence s1) and the unconditional denoising result to obtain a final denoising result, which can amplify the influence of the important condition (text). In this way, not only can the model ignore the secondary conditions (image layer and initial key frame), but also it is helpful to generate clearer and more realistic animations.

[0149] It should be understood that, after the multi-step prediction process is completed, the data sequence s1 or the final sequence obtained by the last step of prediction processing can be taken as a target sequence, and based on the target sequence, the motion trajectory of the target image element in the target image layer can be obtained.

[0150] The target sequence described above can also be post-processed as follows: interpolation optimization, that is, applying cubic spline interpolation to the generated target sequence to generate a smooth animation curve. Loop correction, that is, if a deviation is detected between the first and last parameter groups, the intermediate parameter groups are forced to close the loop through linear adjustment.

[0151] Finally, for the target sequence described above, it can also be exported as a JSON file. That is, the JSON file contains each data group and the corresponding interpolation method, supports import into an editor to realize re-editing, re-rendering and other functions, and can serve animation editors and designers.

[0152] After generating the corresponding target sequences for multiple image layers, existing software or tools can be used to generate target animations.

[0153] In summary, the method for predicting the motion trajectory of an image element provided by the embodiments of the present disclosure can generate control conditions (i.e., a plurality of regulation coefficients) for a diffusion network based on multi-modal data such as images and texts, thereby improving the accuracy of the generated control conditions and the accuracy of the denoising result. In addition, the trajectory prediction model trained by the present solution can generate editable intermediate products (i.e., a series of transformation parameters describing the motion trajectory of an image element) of an animation, which can greatly improve the accuracy and flexibility of the animation generation. Furthermore, the present solution can also support the generation of seamless loop animations. Finally, the present solution integrates AI algorithms into the animation generation process, which can help users achieve a more intelligent and automated design process, improve production efficiency, reduce labor costs, and help animation effect platform products have better user experience and greater competitiveness in the industry.

[0154] In summary, the present solution breaks through the technical bottlenecks of business products in animation design and generation, especially in terms of intelligence, automation, and personalization.

[0155] Corresponding to the method for training the trajectory prediction model, one embodiment of the present disclosure also provides a device for training a trajectory prediction model, which includes a generation network and a diffusion network. As shown in the above method for training the trajectory prediction model, the device can include: Figure 5

[0156] The generation unit 502 is configured to input the sample image layer extracted from the sample animation and the sample text describing the animation effect of the image element in the sample image layer into the generation network to obtain a plurality of regulation coefficients corresponding to a plurality of transformation parameters describing the motion trajectory of the image element.

[0157] The denoising unit 504 is configured to input the plurality of regulation coefficients and a noise parameter sequence into the diffusion network, so that the diffusion network performs denoising processing on the noise parameter sequence based on the plurality of regulation coefficients to obtain a target parameter sequence. The noise parameter sequence includes a plurality of noise parameter groups of a plurality of key frames of the sample image layer, and a single noise parameter group includes a plurality of parameter values with noise corresponding to a plurality of transformation parameters of a key frame.

[0158] The determination unit 506 is configured to determine a basic loss according to a first difference between the target parameter sequence and a real parameter sequence of the sample image layer.

[0159] The adjustment unit 508 is configured to adjust the parameters of the trajectory prediction model according to a comprehensive loss, wherein the comprehensive loss includes the basic loss.

[0160] In one embodiment, the generation network includes a feature extraction module, a linear layer, and a gating module, and the generation unit 502 includes:

[0161] ​The processing submodule 5022 is configured to process the sample image layer and the sample text by using the feature extraction module to obtain a target aggregated representation.

[0162] The processing submodule 5022 is further configured to process the target aggregated representation by using a linear layer to obtain a plurality of overall regulation coefficients shared by a plurality of transformation parameters.

[0163] The determining submodule 5024 is configured to determine a plurality of target weights for the plurality of transformation parameters based on the sample text by using a gating module, and determine a plurality of regulation coefficients corresponding to the plurality of transformation parameters according to a product of the plurality of target weights and the overall regulation coefficients.

[0164] In one embodiment, the generation network described above further includes a first image encoder and a text encoder, and the apparatus further includes:

[0165] The feature extraction unit 510 is configured to perform feature extraction on the sample image layer by using the first image encoder to obtain a first image representation.

[0166] The text encoder is configured to perform feature extraction on the sample text to obtain a text representation.

[0167] The processing submodule 5022 is specifically configured to:

[0168] The feature extraction module is configured to aggregate the target image representation and the text representation to obtain a target aggregated representation, wherein the target image representation is determined at least according to the first image representation.

[0169] In one embodiment, the generation network described above further includes a second image encoder and a representation fusion network, and the apparatus further includes a processing unit 512.

[0170] The feature extraction unit 510 is further configured to perform feature extraction on the starting key frame of the sample animation by using the second image encoder to obtain a second image representation.

[0171] The processing unit 512 is configured to process a combination of the first image representation and the second image representation by using the representation fusion network to obtain a target image representation.

[0172] In one embodiment, the feature extraction module described above includes a first attention layer, a second attention layer and a third attention layer, and the processing submodule 5022 is specifically configured to:

[0173] In the first attention layer, the text representation is aggregated based on a self-attention mechanism to obtain a corresponding updated representation.

[0174] In the second attention layer, the target vector is updated based on the target image representation by using a cross-attention mechanism to obtain an intermediate vector after fusing image information.

[0175] In the third attention layer, based on the updated representation, the intermediate vector is updated by using a cross-attention mechanism to obtain a target aggregated representation fusing image and text information.

[0176] In an embodiment, the diffusion network includes a preset number of denoising modules including a normalization layer arranged in series, and the regulation coefficient corresponding to a single transformation parameter includes an individual scaling coefficient and an individual offset coefficient.

[0177] The denoising unit 504 includes:

[0178] The normalization submodule 5042 is configured to perform normalization on each matrix element in a feature matrix corresponding to the noise parameter sequence in the normalization layer in any first denoising module, to obtain normalized matrix elements, each row of the feature matrix corresponding to a plurality of noise parameter groups, and each column corresponding to a transformation parameter.

[0179] The calculation submodule 5044 is configured to multiply an individual scaling coefficient corresponding to a column to which a first matrix element in the normalized matrix elements belongs by the first matrix element, and add a product to an individual offset coefficient corresponding to the column to obtain a target matrix element, each target matrix element corresponding to each matrix element forming a target matrix, which is used to determine an output of the first denoising module.

[0180] In an embodiment, the first denoising module is the last denoising module in the preset number of denoising modules, and the output is the target parameter sequence.

[0181] In an embodiment, the apparatus further includes a calculation unit 514.

[0182] The calculation unit 514 is configured to calculate a second difference degree for each adjacent two target parameter groups in the target parameter sequence.

[0183] The determination unit 506 is further configured to determine a first penalty term by comparing the second difference degrees.

[0184] The determination unit 506 is further configured to determine a comprehensive loss according to the basic loss and the first penalty term.

[0185] In an embodiment, the determination unit 506 is further configured to determine a second penalty term according to a third difference degree of a first and a last target parameter group in the target parameter sequence and a fourth difference degree of a first and a last real parameter group in the real parameter sequence.

[0186] The determination unit 506 is further configured to add the basic loss, the first penalty term, and the second penalty term to obtain the comprehensive loss.

[0187] The functions of each functional unit of the above-mentioned embodiment device can be realized through each step of the above-mentioned method embodiment, and therefore, the specific working process of the device provided by one embodiment of the present specification will not be repeated here.

[0188] The device for training the trajectory prediction model provided by one embodiment of the present specification can generate an editable intermediate product of the animation, so that the accuracy and flexibility of the animation generation can be greatly improved.

[0189] Corresponding to the above-mentioned method for predicting the motion trajectory of the image element, one embodiment of the present specification also provides a device for predicting the motion trajectory of the image element, as shown in Figure 6 The device can include:

[0190] The acquisition unit 602 is configured to acquire a trajectory prediction model.

[0191] The generation unit 604 is configured to input the target image layer and the target text describing the animation effect of the target image element in the target image layer into a generation network in the trajectory prediction model, to obtain a plurality of first control coefficients.

[0192] The denoising unit 606 is configured to input the plurality of first control coefficients and a noise data sequence into a diffusion network of the trajectory prediction model, so that the diffusion network performs denoising processing on the noise data sequence based on the plurality of first control coefficients, to obtain a first data sequence used to determine the motion trajectory of the target image element, the noise data sequence includes a plurality of noise data groups, and each noise data in a single noise data group corresponds to each first control coefficient.

[0193] In one embodiment, the device further includes a fusion unit 608 and a determination unit 610.

[0194] The generation unit 604 is further configured to input only the target image layer into the generation network to obtain a plurality of second control coefficients, each second control coefficient corresponding to each noise data in a single noise data group.

[0195] The denoising unit 606 is further configured to input the plurality of second control coefficients and the noise data sequence into the diffusion network, so that the diffusion network performs denoising processing on the noise data sequence based on the plurality of second control coefficients, to obtain a second data sequence.

[0196] The fusion unit 608 is configured to perform weighted fusion on the first data sequence and the second data sequence to obtain a final data sequence, the weight of the first data sequence being greater than that of the second data sequence.

[0197] The determination unit 610 is configured to determine the motion trajectory of the target image element according to the final data sequence.

[0198] The functions of each functional unit of the device in the above embodiments of the present specification can be realized by each step of the above method embodiments, and thus the specific working process of the device provided by one embodiment of the present specification is not repeated here.

[0199] The device for predicting the motion trajectory of the image element provided by one embodiment of the present specification can predict the editable intermediate product of the animation, and thus the accuracy and flexibility of the animation generation can be greatly improved.

[0200] According to another aspect of the embodiments, a computer readable storage medium is also provided, which stores a computer program. When the computer program is executed in a computer, the computer is caused to perform the method described above. Figure 2

[0201] According to another aspect of the embodiments, a computer readable storage medium is also provided, which stores a computer program. When the computer program is executed in a computer, the computer is caused to perform the method described above. Figure 2

[0202] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the medium or device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0203] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in which they are recited and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.

[0204] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present specification. It should be understood that the above description is only for specific embodiments of the present specification and is not used to limit the protection scope of the present specification. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present specification shall be included in the protection scope of the present specification.​​

Claims

1. A method for training a trajectory prediction model, the trajectory prediction model comprising a generation network and a diffusion network; the method comprising: inputting a sample image layer extracted from a sample animation and sample text describing animation effects of image elements in the sample image layer into the generation network to obtain a plurality of control coefficients corresponding to a plurality of transformation parameters describing motion trajectories of the image elements; inputting the plurality of control coefficients and a noise parameter sequence into the diffusion network to perform denoising processing on the noise parameter sequence based on the plurality of control coefficients to obtain a target parameter sequence; the noise parameter sequence comprises a plurality of noise parameter groups of a plurality of key frames of the sample image layer, and a single noise parameter group comprises noisy parameter values of the plurality of transformation parameters corresponding to a key frame; determining a basic loss according to a first difference between the target parameter sequence and a real parameter sequence of the sample image layer; adjusting parameters of the trajectory prediction model according to a comprehensive loss; the comprehensive loss comprises the basic loss.

2. The method of claim 1, wherein, the generation network comprises a feature extraction module, a linear layer and a gating module; the obtaining of the plurality of control coefficients corresponding to the plurality of transformation parameters comprises: processing the sample image layer and the sample text by using the feature extraction module to obtain a target aggregated representation; processing the target aggregated representation by using the linear layer to obtain an overall control coefficient shared by the plurality of transformation parameters; determining a plurality of target weights for the plurality of transformation parameters based on the sample text by using the gating module, and determining the plurality of control coefficients corresponding to the plurality of transformation parameters according to products of the plurality of target weights and the overall control coefficient respectively.

3. The method of claim 2, wherein, the generation network further comprises a first image encoder and a text encoder; the method further comprises: extracting features of the sample image layer by using the first image encoder to obtain a first image representation; extracting features of the sample text by using the text encoder to obtain a text representation; the obtaining of the target aggregated representation comprises: aggregating the target image representation and the text representation by using the feature extraction module to obtain the target aggregated representation; the target image representation is determined at least according to the first image representation.

4. The method of claim 3, wherein, the generation network further comprises a second image encoder and a representation fusion network; the method further comprises: extracting features of a starting key frame of the sample animation by using the second image encoder to obtain a second image representation; processing a combination of the first image representation and the second image representation by using the representation fusion network to obtain the target image representation.

5. The method of claim 3, wherein, the feature extraction module comprises a first attention layer, a second attention layer and a third attention layer; the aggregating of the target image representation and the text representation comprises: performing aggregation processing on the text representation based on a self-attention mechanism in the first attention layer to obtain a corresponding updated representation; in the second attention layer, updating a target vector based on the target image representation by using a cross-attention mechanism to obtain an intermediate vector after fusing image information; In the third attention layer, based on the updated representation, the intermediate vector is updated by using a cross-attention mechanism to obtain a target aggregated representation fusing image and text information.

6. The method of claim 1, wherein, The diffusion network comprises a preset number of denoising modules comprising a normalization layer arranged in series. The regulation coefficient corresponding to a single transformation parameter comprises an individual scaling coefficient and an individual offset coefficient. The denoising processing of the noise parameter sequence comprises: In the normalization layer in any first denoising module, each matrix element in a feature matrix corresponding to the noise parameter sequence is normalized by row to obtain a normalized matrix element; each row of the feature matrix corresponds to the plurality of noise parameter groups, and each column corresponds to each transformation parameter; For any first matrix element in the normalized matrix elements, an individual scaling coefficient corresponding to the column to which the first matrix element belongs is multiplied by the first matrix element, and a product is added to an individual offset coefficient corresponding to the column to obtain a target matrix element; each target matrix element corresponding to each matrix element forms a target matrix, which is used to determine the output of the first denoising module.

7. The method of claim 6, wherein, The first denoising module is the last denoising module in the preset number of denoising modules, and the output is the target parameter sequence.

8. The method of claim 1, further comprising: calculating a second difference degree for each adjacent two target parameter groups in the target parameter sequence; determining a first penalty term by comparing the second difference degrees; determining the comprehensive loss according to the basic loss and the first penalty term.

9. The method of claim 8, further comprising: determining a second penalty term according to a third difference degree of a first and a last target parameter group in the target parameter sequence and a fourth difference degree of a first and a last real parameter group in the real parameter sequence; the determination of the comprehensive loss comprises: adding the basic loss, the first penalty term and the second penalty term to obtain the comprehensive loss.

10. A method for predicting a motion trajectory of an image element, comprising: obtaining a trajectory prediction model trained according to the method of claim 1; inputting a target image layer and a target text describing an animation effect of a target image element in the target image layer into a generation network in the trajectory prediction model to obtain a plurality of first regulation coefficients; inputting the plurality of first regulation coefficients and a noise data sequence into a diffusion network of the trajectory prediction model, so that the diffusion network performs denoising processing on the noise data sequence based on the plurality of first regulation coefficients to obtain a first data sequence used to determine a motion trajectory of the target image element, the noise data sequence comprising a plurality of noise data groups, each noise data in a single noise data group corresponding to each first regulation coefficient.

11. The method of claim 10, further comprising: inputting only the target image layer into the generation network to obtain a plurality of second regulation coefficients; each second regulation coefficient corresponding to each noise data in a single noise data group. input the second plurality of regulation coefficients and the noise data sequence into the diffusion network, so that the diffusion network denoises the noise data sequence based on the second plurality of regulation coefficients to obtain a second data sequence; perform weighted fusion on the first data sequence and the second data sequence to obtain a final data sequence; the weight of the first data sequence is greater than that of the second data sequence; determine the motion trajectory of the target image element according to the final data sequence.

12. An apparatus for training a trajectory prediction model, the trajectory prediction model comprising, a generative network and a diffusion network; The device comprises: a generation unit configured to input a sample image layer extracted from a sample animation and sample text describing animation effects of image elements in the sample image layer into the generation network to obtain a plurality of regulation coefficients corresponding to a plurality of transformation parameters describing the motion trajectory of the image elements; a denoising unit configured to input the plurality of regulation coefficients and a noise parameter sequence into the diffusion network, so that the diffusion network denoises the noise parameter sequence based on the plurality of regulation coefficients to obtain a target parameter sequence; the noise parameter sequence comprises a plurality of noise parameter groups of a plurality of key frames of the sample image layer, and a single noise parameter group comprises noise parameter values of the plurality of transformation parameters corresponding to a key frame; a determination unit configured to determine a basic loss according to a first difference between the target parameter sequence and a true parameter sequence of the sample image layer; an adjustment unit configured to adjust parameters of the trajectory prediction model according to a comprehensive loss; the comprehensive loss comprises the basic loss.

13. A device for predicting a motion trajectory of an image element, comprising: an acquisition unit configured to acquire a trajectory prediction model trained according to the method of claim 1; a generation unit configured to input a target image layer and target text describing animation effects of target image elements in the target image layer into a generation network in the trajectory prediction model to obtain a plurality of first regulation coefficients; a denoising unit configured to input the plurality of first regulation coefficients and a noise data sequence into a diffusion network of the trajectory prediction model, so that the diffusion network denoises the noise data sequence based on the plurality of first regulation coefficients to obtain a first data sequence used to determine the motion trajectory of the target image elements; the noise data sequence comprises a plurality of noise data groups, and each noise data in a single noise data group corresponds to a first regulation coefficient.

14. A computer readable storage medium having stored thereon a computer program, wherein, When the computer program is executed in the computer, the computer is caused to perform the method of any one of claims 1-11.

15. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and the processor executes the executable code to implement the method of any one of claims 1-11. The memory stores executable code, and the processor executes the executable code to implement the method of any one of claims 1-11.