Action Generation Method and Related Devices, Electronic Devices, and Storage Media

By acquiring and modeling individual feature representations and performing action mapping, the problem of poor compatibility of action generation methods in the prior art in single and multiple individual application scenarios is solved, and efficient action generation is achieved.

CN114494543BActive Publication Date: 2025-07-22SHANGHAI SENSETIME TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210089863.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-07-22
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The existing action generation method is inefficient when compatible with application scenarios of single individuals and multiple individuals, and cannot be effectively compatible.

Method used

By obtaining the first feature representation of several individuals in several action frames and the second feature representation of the target action category, relationship modeling is performed to obtain the fusion feature representation of each individual in each action frame, and action mapping is performed based on the fusion feature representation. The type of relationship modeling is related to the total number of individuals, including modeling timing relationships and interaction relationships.

Benefits of technology

It realizes application scenarios that are compatible with single individuals and multiple individuals under the premise of improving the efficiency of action generation, improves the authenticity and rationality of action sequences, and reduces the dependence on manual modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494543B_ABST
    Figure CN114494543B_ABST
Patent Text Reader

Abstract

The present application discloses an action generation method and related devices, an electronic device, and a storage medium. Among them, the action generation method includes: obtaining first feature representations respectively characterizing a plurality of individuals in a plurality of action frames, and obtaining second feature representations respectively characterizing the plurality of individuals with respect to a target action category; performing relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each individual in each action frame; wherein the type of relationship modeling is related to the first total number of the plurality of individuals; performing action mapping based on the fused feature representations to obtain an action sequence of the plurality of individuals with respect to the target action category; wherein the action sequence includes a plurality of action frames, and each action frame includes action representations of each individual. The above solution can be compatible with two application scenarios of a single individual and multiple individuals on the premise of improving the action generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a method for generating actions and related devices, electronic devices, and storage media. Background Art

[0002] Action generation is crucial for many computer vision tasks such as animation creation and humanoid robot interaction. Currently, the existing action generation methods mainly include two types. One is the modeling-rendering method based on computer graphics, which requires designers to spend a lot of time and effort on modeling, skinning, and motion capture, etc., with low efficiency. The other is the method based on machine learning, especially deep learning.

[0003] Thanks to the rapid development of machine learning technology in recent years, using deep neural networks to perform action generation tasks can greatly improve the efficiency of action generation. However, the existing methods usually only consider the action generation of a single individual and cannot be compatible with application scenarios of multiple individuals. In view of this, how to be compatible with both single-individual and multiple-individual application scenarios while improving the efficiency of action generation has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides a method for generating actions and related devices, electronic devices, and storage media.

[0005] In a first aspect of this application, a method for generating actions is provided, including: obtaining first feature representations respectively characterizing a plurality of individuals in a plurality of action frames, and obtaining second feature representations respectively characterizing the plurality of individuals with respect to a target action category; performing relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each individual in each action frame; wherein, the type of relationship modeling is related to the first total number of the plurality of individuals; performing action mapping based on the fused feature representations to obtain action sequences of the plurality of individuals with respect to the target action category; wherein, the action sequences include a plurality of action frames, and the action frames include action representations of each individual.

[0006] Therefore, first feature representations respectively characterizing a number of individuals in a number of action frames are obtained, and second feature representations respectively characterizing the number of individuals with respect to a target action category are obtained. On this basis, relationship modeling is performed based on the first feature representations and the second feature representations to obtain fused feature representations of each individual in each action frame, and the type of relationship modeling is related to the first total number of the number of individuals. Then, action mapping is performed based on the fused feature representations, and an action sequence of the number of individuals with respect to the target action category can be obtained. The action sequence includes a number of action frames, and each action frame includes action representations of each individual. Therefore, on the one hand, actions can be automatically generated without relying on manual labor, and on the other hand, by specifically performing relationship modeling according to the first total number of the number of individuals, two application scenarios of a single individual and multiple individuals can be compatible. Therefore, two application scenarios of a single individual and multiple individuals can be compatible on the premise of improving the action generation efficiency.

[0007] Among them, when the first total number of the number of individuals is single, the relationship modeling includes modeling the temporal relationship between each action frame; and / or, when the first total number of the number of individuals is multiple, the relationship modeling includes modeling the interaction relationship between a number of individuals in each action frame and modeling the temporal relationship between each action frame.

[0008] Therefore, when the first total number of the number of individuals is single, the relationship modeling includes modeling the temporal relationship between each action frame, so the temporal coherence between action frames can be improved by modeling the temporal relationship, which is beneficial to improving the authenticity of the action sequence. When the first total number of the number of individuals is multiple, the relationship modeling includes modeling the interaction relationship between a number of individuals in each action frame and modeling the temporal relationship between each action frame. Therefore, the interaction rationality between individuals can be improved by modeling the interaction relationship, and the temporal coherence between action frames can be improved by modeling the temporal relationship, which is beneficial to improving the authenticity of the action sequence.

[0009] Among them, when the relationship modeling includes modeling the temporal relationship, performing relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each individual in each action frame includes: selecting an individual as the target individual, and using the first feature representation and the second feature representation corresponding to the target individual as the temporal feature representations of the target individual at different times; respectively selecting each time as the first current time, and selecting the temporal feature representation of the first current time as the first current time representation; obtaining the fused feature representation corresponding to the first current time representation based on the correlation between each first reference time representation and the first current time representation; where the first reference time representation includes the temporal feature representations of the target individual at each time.

[0010] Among them, when the relationship modeling includes modeling interaction relationships, relationship modeling is performed based on the first feature representation and the second feature representation to obtain the fused feature representation of each entity in each action frame, including: selecting an entity as the target entity, and using the first feature representation and the second feature representation corresponding to the target entity as the temporal feature representations of the target entity at different time sequences; respectively selecting each time sequence as the second current time sequence, and selecting the temporal feature representation of the second current time sequence as the second current time sequence representation; obtaining the fused feature representation corresponding to the second current time sequence representation based on the relevance between each second reference time sequence representation and the second current time sequence representation; where the second reference time sequence representation includes the temporal feature representations of each entity at the second current time sequence respectively.

[0011] Therefore, an entity is selected as the target entity, and the first feature representation and the second feature representation corresponding to the target entity are used as the temporal feature representations of the target entity at different time sequences. Based on this, the temporal feature representations at different time sequences are respectively used as the current time sequence representations, and then the fused feature representation corresponding to the current time sequence representation is obtained based on the relevance between each reference time sequence representation and the current time sequence representation. And when modeling the temporal relationship, the reference time sequence representation includes the temporal feature representations of the target entity at each time sequence. When modeling the interaction relationship, the reference time sequence representation includes the temporal feature representations of each entity at the reference time sequence respectively, and the reference time sequence is the time sequence corresponding to the current time sequence representation. Therefore, the temporal relationship and the interaction relationship can be modeled through a similar modeling process, which can further improve the compatibility of the two application scenarios of a single entity and multiple entities.

[0012] Among them, when the relationship modeling includes modeling interaction relationships and temporal relationships, relationship modeling is performed based on the first feature representation and the second feature representation to obtain the fused feature representation of each entity in each action frame, including: modeling the prior relationship based on the first feature representation and the second feature representation to obtain the output feature representation of the prior relationship, and modeling the subsequent relationship based on the output feature representation to obtain the fused feature representation; where the prior relationship is an interaction relationship and the subsequent relationship is a temporal relationship, or the prior relationship is a temporal relationship and the subsequent relationship is an interaction relationship.

[0013] Therefore, when the relationship modeling includes modeling interaction relationships and temporal relationships, the output feature representation of the prior modeled interaction relationship is used as the input feature representation of the subsequent modeled temporal relationship. Therefore, in the application scenario of multiple entities, by modeling the interaction relationship and the temporal relationship successively, each fused feature representation is respectively incorporated into the interaction relationship and the temporal relationship, which is beneficial to improving the fusion effect of the interaction relationship and the temporal relationship.

[0014] Among them, the action sequence is obtained by an action generation model, the action generation model includes a relationship modeling network, and the relationship modeling network includes a temporal modeling sub-network and an interaction modeling sub-network. The temporal modeling sub-network is used to model temporal relationships, and the interaction modeling sub-network is used to model interaction relationships.

[0015] Therefore, the action sequence is obtained by an action generation model. The action generation model includes a relationship modeling network, and the relationship modeling network includes a temporal modeling sub-network and an interaction modeling sub-network. The temporal modeling sub-network is used to model temporal relationships, and the interaction modeling sub-network is used to model interaction relationships. Therefore, the action generation task can be completed through the network model, which is beneficial to further improving the action generation efficiency.

[0016] Among them, the first feature representation is obtained based on sampling of a Gaussian process.

[0017] Therefore, obtaining the first feature representation based on sampling of a Gaussian process is beneficial to greatly reducing the acquisition complexity of the first feature representation, and can also improve the generation quality on action data with rich categories.

[0018] Among them, obtaining the first feature representations respectively characterizing several individuals in several action frames includes: in several Gaussian processes, sampling the second total number of times respectively to obtain first original representations respectively characterizing the second total number of action frames; where the length of the first original representation is the same as the number of Gaussian processes, and the characteristic length scales of each Gaussian process are different; based on the first total number and the first original representation, obtaining the third total number of first feature representations; where the third total number is the product of the first total number and the second total number.

[0019] Therefore, in several Gaussian processes, sampling the second total number of times respectively to obtain first original representations respectively characterizing the second total number of action frames, and the length of the first original representation is the same as the number of Gaussian processes, and the characteristic length scales of each Gaussian process are different. Based on this, and then based on the first total number and the first original representation, obtaining the third total number of first feature representations, and the third total number is the product of the first total number and the second total number. Since the characteristic length scales of each Gaussian process are different, and each sampling of the Gaussian process can obtain the feature information of each action frame, the accuracy of each first feature representation can be improved.

[0020] Among them, the second feature representation is obtained based on mapping of the target action category.

[0021] Therefore, obtaining the second feature representation based on mapping the target action category, so only simple processing such as mapping of text information is required to obtain the second feature representation, which is beneficial to greatly reducing the complexity of driving action generation.

[0022] Among them, obtaining second feature representations respectively characterizing a number of individuals with respect to a target action category includes: performing an embedding representation on the target action category to obtain a second original representation; and obtaining a first total number of second feature representations based on the first total number and the second original representation.

[0023] Therefore, performing an embedding representation on the target action category to obtain a second original representation, and obtaining a first total number of second feature representations based on the first total number and the second original representation, that is, by performing an embedding representation on the text information and combining with the first total number for relevant processing, a first total number of second feature representations can be obtained, which is beneficial to greatly reducing the complexity of obtaining the second feature representation.

[0024] Among them, both the first feature representation and the second feature representation are fused with position encodings; where, in the case where the number of individuals is a single individual, the position encoding includes a temporal position encoding, and in the case where the number of individuals is multiple individuals, the position encoding includes an individual position encoding and a temporal position encoding.

[0025] Therefore, both the first feature representation and the second feature representation are fused with position encodings. In the case where the number of individuals is a single individual, the position encoding includes a temporal position encoding, and in the case where the number of individuals is multiple individuals, the position encoding includes an individual position encoding and a temporal position encoding. Therefore, different position encoding strategies can be adopted to distinguish different feature representations in two application scenarios of a single individual and multiple individuals, so that the position encodings of the feature representations are different from each other, which is beneficial to improving the accuracy of the feature representation.

[0026] Among them, the action sequence is obtained by an action generation model, and the position encoding is adjusted together with the network parameters of the action generation model during the training process of the action generation model until the training of the action generation model converges.

[0027] Therefore, the action sequence is obtained by an action generation model, and the position encoding is adjusted together with the network parameters of the action generation model during the training process of the action generation model until the training of the action generation model converges. Since the position encoding is trained together with the network model, the representation ability of the position encoding can be improved, and after the training converges, the position encoding is no longer adjusted, that is, it remains fixed, so that a strong prior constraint can be added, so that a balance can be achieved between the prior constraint and the representation ability, and further the accuracy of the feature representation can be improved, which is beneficial to improving the generation effect of the action sequence.

[0028] Among them, the action representation of an individual in an action frame includes: in the action frame, the first position information of the key points of the individual and the pose information of the individual, and the pose information includes the second position information of a number of joint points of the individual.

[0029] Therefore, the action representation of an individual in an action frame includes: the first position information of the key points of the individual in the action frame and the pose information of the individual, and the pose information includes the second position information of several joint points of the individual. Therefore, the individual action can be expressed by the position information of both the key points and the joint points, which is beneficial to improving the accuracy of action representation.

[0030] Among them, the action sequence is obtained by an action generation model, and the action generation model and the discriminative model are obtained through generative adversarial training.

[0031] Therefore, by using generative adversarial training to co-train the action generation model and the discriminative model, the action generation model and the discriminative model can promote each other during the co-training process, complementing each other, and ultimately being beneficial to improving the model performance of the action generation model.

[0032] Among them, the steps of generative adversarial training include: obtaining the sample action sequences of several sample individuals regarding the sample action categories; among them, the sample action sequence includes a preset number of sample action frames, and the sample action sequence is labeled with a sample label, and the sample label indicates whether the sample action sequence is actually generated by the action generation model; decomposing each sample action frame in the sample action sequence to obtain sample graph data; among them, the sample graph data includes a preset number of node graphs, the node graphs are formed by connecting nodes, the nodes include key points and joint points, the node graphs include the node feature representations of each node, and the position feature representation of the node is obtained by splicing the position feature representations of several sample individuals at the corresponding nodes respectively; based on the discriminative model, discriminating the sample graph data and the sample action categories to obtain a prediction result; among them, the prediction result includes the first prediction label of the sample action sequence and the second prediction label, the first prediction label indicates the possibility that the sample action sequence is predicted to be generated by the action generation model, and the second prediction label indicates the possibility that the sample action sequence belongs to the sample action category; based on the sample label, the first prediction label and the second prediction label, adjusting the network parameters of either the action generation model or the discriminative model.

[0033] Therefore, by decomposing the sample action representation into sample graph data, the discrimination of the action sequence can be cleverly resolved into the discrimination of graph data, which is beneficial to greatly reducing the training complexity and the construction difficulty of the discriminative model.

[0034] Among them, when the sample action sequence is collected from a real scene, the position feature representation of the node is obtained by splicing the position feature representations of several sample individuals at the corresponding nodes in a random order of the several sample individuals.

[0035] Therefore, the position feature representation of the nodes is obtained by splicing the position feature representations of several sample individuals at the corresponding nodes in a random order of several sample individuals. Thus, during the training process, the action generation model treats the cases where different orderings actually belong to the same sample action sequence as different samples and models them, thereby enabling data augmentation and further facilitating the improvement of the model robustness.

[0036] The second aspect of this application provides an action generation device, including: a feature acquisition module, a relationship modeling module, and an action mapping module. The feature acquisition module is used to acquire the first feature representations respectively characterizing several individuals in several action frames, and acquire the second feature representations respectively characterizing several individuals regarding the target action category; the relationship modeling module is used to perform relationship modeling based on the first feature representations and the second feature representations to obtain the fused feature representations of each individual in each action frame; wherein, the type of relationship modeling is related to the first total number of several individuals; the action mapping module is used to perform action mapping based on the fused feature representations to obtain the action sequences of several individuals regarding the target action category; wherein, the action sequences include several action frames, and the action frames include the action representations of each individual.

[0037] The third aspect of this application provides an electronic device, including a memory and a processor coupled to each other. The processor is used to execute the program instructions stored in the memory to implement the action generation method in the first aspect above.

[0038] The fourth aspect of this application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by the processor, the action generation method in the first aspect above is implemented.

[0039] In the above solution, the first feature representations respectively characterizing several individuals in several action frames are acquired, and the second feature representations respectively characterizing several individuals regarding the target action category are acquired. On this basis, relationship modeling is performed based on the first feature representations and the second feature representations to obtain the fused feature representations of each individual in each action frame, and the type of relationship modeling is related to the first total number of several individuals. Then, action mapping is performed based on the fused feature representations, and the action sequences of several individuals regarding the target action category can be obtained. The action sequences include several action frames, and the action frames include the action representations of each individual. Therefore, on the one hand, actions can be automatically generated without relying on manual work, and on the other hand, by specifically performing relationship modeling according to the first total number of several individuals, the two application scenarios of a single individual and multiple individuals can be compatible. Thus, the two application scenarios of a single individual and multiple individuals can be compatible on the premise of improving the action generation efficiency. Description of the Drawings

[0040] Figure 1 is a schematic flowchart of an embodiment of the action generation method of this application;

[0041] Figure 2 It is a schematic process diagram of an embodiment of the method for generating actions in this application;

[0042] Figure 3a It is a schematic diagram of an embodiment of an action sequence;

[0043] Figure 3b It is a schematic diagram of an embodiment of an action sequence;

[0044] Figure 3c It is a schematic diagram of an embodiment of an action sequence;

[0045] Figure 3d It is a schematic diagram of an embodiment of an action sequence;

[0046] Figure 3e It is a schematic diagram of an embodiment of an action sequence;

[0047] Figure 3f It is a schematic diagram of an embodiment of an action sequence;

[0048] Figure 4 It is a schematic flowchart of an embodiment of the training method of an action generation model;

[0049] Figure 5 It is a schematic diagram of obtaining an embodiment of a sample action frame;

[0050] Figure 6 It is a schematic diagram of an embodiment of sample graph data;

[0051] Figure 7 It is a schematic framework diagram of an embodiment of the action generation device in this application;

[0052] Figure 8 It is a schematic framework diagram of an embodiment of the electronic device in this application;

[0053] Figure 9 It is a schematic framework diagram of an embodiment of the computer-readable storage medium in this application. Detailed implementation manners

[0054] The solutions of the embodiments of this application will be described in detail below with reference to the accompanying drawings of the specification.

[0055] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to understand this application thoroughly.

[0056] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship. In addition, "plurality" in this document means two or more than two.

[0057] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the method for generating actions in this application. Specifically, it may include the following steps:

[0058] Step S11: Obtain first feature representations respectively characterizing a plurality of individuals in a plurality of action frames, and obtain second feature representations respectively characterizing the plurality of individuals with respect to the target action category.

[0059] In one implementation scenario, the first total number of a plurality of individuals and the target action category can be specified by the user before formally implementing action generation. Exemplarily, the user can specify the target action category as "hug" and specify the first total number of a plurality of individuals as two; or, the user can specify the target action category as "dance" and specify the first total number of a plurality of individuals as one; or, the user can specify the target action category as "fight" and specify the first total number of a plurality of individuals as three. It should be noted that the above examples are only several possible implementation manners in the actual application process, and do not limit the target action category and the first total number of a plurality of individuals in the actual application process.

[0060] In another implementation scenario, the target action category can be specified by the user before formally implementing action generation, and the first total number of a plurality of individuals can be automatically analyzed based on the target action category. Exemplarily, the user can specify the target action category as "high-five", then based on this target action category, the first total number of a plurality of individuals can be automatically analyzed as two; or, the user can specify the target action category as "exchange items", then based on this target action category, the first total number of a plurality of individuals can be automatically analyzed as two; or, the user can specify the target action category as "carry items", then based on this target action category, the first total number of a plurality of individuals can be automatically analyzed as one. It should be noted that the above examples are only several possible implementation manners in the actual application process, and do not limit the target action category and the first total number of a plurality of individuals in the actual application process.

[0061] In yet another implementation scenario, the target action category can be specified by the user before the formal generation of the action, and the first total number of several individuals can be automatically analyzed based on the target action category, and the user's modification instruction for the first total number obtained by automatic analysis can be accepted to correct the first total number obtained by automatic analysis. Exemplarily, the user can specify the target action category as "fighting", then based on this target action category, the first total number of several individuals can be automatically analyzed as two, and the user's modification instruction for the first total number obtained by automatic analysis can be accepted to correct it to four; or, the user can specify the target action category as "taking a walk", then based on this target action category, the first total number of several individuals can be automatically analyzed as one, and the user's modification instruction for the first total number obtained by automatic analysis can be accepted to correct it to two. It should be noted that the above examples are only several possible implementation manners in the actual application process, and do not limit the target action category and the first total number of several individuals in the actual application process accordingly.

[0062] It should be noted that the above-mentioned several individuals can all be people. Of course, it is not excluded that several individuals include both people and animals at the same time. Exemplarily, the target action category can be specified as "walking the dog", then several individuals can include a person and a dog.

[0063] In one implementation scenario, the second total number of several action frames can be specified in advance. Exemplarily, the second total number can be 10, 15, 20, 30, etc., which is not limited here.

[0064] In one implementation scenario, the first feature representation of each individual in each action frame can be obtained. For example, for the case where the first total number of several individuals is one (i.e., for the action generation scenario of a single individual), the first feature representation of this single individual in each action frame can be obtained; or, for the case where the first total number of several individuals is two (i.e., for the action generation scenario of two individuals), the first feature representation of each individual in each action frame can be obtained. For the convenience of description, these two individuals can be respectively called "A" and "B", then the first feature representation of "A" in each action frame can be obtained, and the first feature representation of "B" in each action frame can be obtained. Other cases can be deduced by analogy, and no more examples will be given here.

[0065] In an implementation scenario, it should be noted that the action frame is included in the action sequence ultimately expected to be generated in the disclosed implementation example of the action generation method of the present application. That is, when obtaining the first feature representation, the action frame is not actually generated, and the first feature representation can be regarded as the feature representation initialized by each individual in each action frame respectively. Specifically, the first feature representation can be obtained by sampling based on a Gaussian process. It should be noted that the Gaussian process is a type of stochastic process in probability theory and mathematical statistics, which is a combination of a series of random variables obeying the normal distribution within an exponential set. For the specific meaning of the Gaussian process, the technical details of the Gaussian process can be referred to and will not be elaborated here.

[0066] In a specific implementation scenario, it is possible to sample the second total number of times respectively in several Gaussian processes to obtain the first original representations respectively representing the second total number of action frames, and the length of the first original representation is the same as the number of Gaussian processes, and the characteristic length scales of each Gaussian process are different. On this basis, based on the first total number and the first original representation, the third total number of first feature representations is obtained, and the third total number is the product of the first total number and the second total number. Exemplarily, for the sake of convenience of description, the second total number of several action frames can be denoted as T, and the characteristic length scales σ c of several Gaussian processes can take values of 1, 10, 100, and 1000 respectively. Then, sample T times in the Gaussian process with the characteristic length scale σ c being 1 to obtain a one-dimensional vector with a length of T. And so on, on the Gaussian processes with the characteristic length scales σ c being 10, 100, and 1000, one-dimensional vectors with a length of T can be sampled. By combining the elements at the same positions on the one-dimensional feature vectors with a length of T sampled from the above 4 Gaussian processes respectively, T first original representations with a length of 4 can be obtained, and these T first original representations correspond to the T action frames one by one, that is, the first first original representation corresponds to the first action frame, the second first original representation corresponds to the second action frame,..., and the Tth first original representation corresponds to the Tth action frame. In addition, please refer to Figure 2 , Figure 2 which is a schematic process diagram of an implementation example of the action generation method of the present application. As Figure 2As shown, for the convenience of description, the length of the first original representation obtained by the above sampling can be recorded as C0, so the above first original representations that respectively represent several action frames can be recorded as (T, C0). On this basis, the above first original representation (T, C0) can be input mapped (for example, the first original representation can be sampled by a multi-layer perceptron to map) to change the dimension of the original first original representation (T, C0). In addition, the number of the first original representations after mapping is still T. In the above method, since the characteristic length scales of each Gaussian process are different, and each time the Gaussian process is sampled, the characteristic information of each action frame can be obtained, the accuracy of each first characteristic representation can be improved.

[0067] In a specific implementation scenario, after obtaining the first original representations representing the second total number of action frames, it can be determined whether to copy the first original representation representing each action frame based on whether the first total number is equal to one or greater than one, so as to obtain the first original representations of several individuals in each action frame. For example, when the first total number is equal to one, it can be determined that the action is generated as a scene of a single individual, and the first original representation representing each action frame obtained by the aforementioned sampling can be directly used as the first original representation of the single individual in each action frame; or, when the first total number is greater than one, it can be determined that the action is generated as a scene of multiple individuals, and the first original representation representing each action obtained by the aforementioned sampling can be copied the first total number of times to obtain the first original representations of multiple individuals in each action frame. For example, when the first total number is 2, the first original representation representing the first action frame can be copied into two first original representations, and the two first original representations respectively represent the first original representations of the two individuals in the first action frame. Other situations can be deduced by analogy, and examples are not given one by one here.

[0068] In a specific implementation scenario, please continue to refer to Figure 2, in order to distinguish different first original representations in the case of a single individual and multiple individuals, the position information of each first original representation can be encoded on the basis of the first original representation to obtain corresponding first feature representations. That is to say, the first feature representations are fused with position encodings, and the position encodings are different from each other. Specifically, in the case where the number of individuals is a single individual, the position encoding includes a temporal position encoding. That is to say, in the case where the number of individuals is a single individual, different first original representations are mainly distinguished by encoding action frames at different times, so as to obtain the first feature representations. Exemplarily, still taking T action frames as an example, in the case of a single individual, the temporal position encodings (such as 1, 2,......, T) can be respectively incorporated into the first original representations of these T action frames, so as to obtain the first feature representations respectively representing the single individual in these T action frames. Similarly, in the case where the number of individuals is multiple individuals, the position encoding can include a temporal position encoding and an individual position encoding. That is to say, in the case where the number of individuals is multiple individuals, not only the action frames at different times need to be encoded, but also the multiple individuals in each action frame need to be encoded, so as to distinguish different first original representations, so as to obtain the first feature representations (as shown by the dashed box after position encoding). Exemplarily, still taking T action frames as an example, in the case of multiple individuals, the temporal position encoding (such as 1) can be incorporated into the first original representation of the first action frame, and the individual position encoding (such as 1, 2,......) can be further incorporated into the multiple individuals in the first action frame, so as to combine the temporal position encoding and the individual position encoding as the position encoding, so that the first feature representations representing the multiple individuals in the first action frame are respectively fused with different position encodings (such as 1-1, 1-2,......); similarly, the temporal position encoding (such as 2) can be incorporated into the first original representation of the second action frame, and the individual position encoding (such as 1, 2,......) can be further incorporated into the multiple individuals in the first action frame, so as to combine the temporal position encoding and the individual position encoding as the position encoding, so that the first feature representations representing the multiple individuals in the second action frame are respectively fused with different position encodings (such as 2-1, 2-2, ……), and so on for other action frames, which will not be listed one by one here. In addition, the above position encoding is only an example. In the actual application process, an action generation model can be pre-trained, and the position encoding can be adjusted together with the network parameters of the action generation model during the training process of the action generation model until the training of the action generation model converges. Since then, in the subsequent application process, the adjusted position encoding can be used. The above method can adopt different position encoding strategies to distinguish different feature representations in two application scenarios of a single individual and multiple individuals, so that the position encodings of the feature representations are different from each other, which is beneficial to improving the accuracy of the feature representations. Figure 2 as shown by the dashed box after position encoding

[0069] In an implementation scenario, similar to the first feature representation, for the second feature representation, the second feature representation of each individual with respect to the target action category can be obtained. For example, in the case where the first total number of several individuals is one (i.e., for the action generation scenario of a single individual), the second feature representation of the single individual with respect to the target action category can be obtained; or, in the case where the first total number of several individuals is two (i.e., for the action generation scenario of two individuals), the second feature representation of each individual with respect to the target action category can be obtained. For the sake of easy description, these two individuals can be respectively called "A" and "B", then the second feature representation of "A" with respect to the target action category can be obtained, and the second feature representation of "B" with respect to the target action category can be obtained. Other cases can be deduced by analogy, and no further examples will be given here.

[0070] In an implementation scenario, as mentioned above, the target action category can be specified by the user. After determining the target action category, the second feature representation can be obtained through mapping based on this target action category.

[0071] In a specific implementation scenario, the target action category can be embedded to obtain the second original representation. On this basis, based on the first total number and the second original representation, the first total number of second feature representations can be obtained. It should be noted that the role of the above embedding is to convert the target action category into a vector. Exemplarily, the category vectors of different action categories can be preset. For example, when there are a total of 26 different action categories, 26 category vectors of action categories can be preset (for example, the length of each category vector can be 200). After determining the target action category, the category vector of the action category that is the same as the target action category can be used as the second original representation of the target action category; or, the target action category can also be one-hot encoded first, and then a linear transformation can be performed using a fully connected layer to obtain the second original representation of the target action category. For example, when there are a total of 26 different action categories, the target action category can be one-hot encoded into a 26-dimensional vector, and the linear transformation of the above fully connected layer can be regarded as a transformation matrix of N (such as 200) * 26. Then, multiplying this matrix by the 26-dimensional one-hot encoding can obtain the second original representation of the target action category.

[0072] In a specific implementation scenario, similar to obtaining the first feature representation, after obtaining the second original representation characterizing the target action category, it is also possible to determine whether to copy the movement mechanism of the first original representation based on whether the first total number is equal to one or greater than one, so as to obtain the second original representations of several individuals respectively regarding the target action category. For example, in the case where the first total number is equal to one, it can be determined that the action is generated for a single individual, and then the second original representation obtained by the foregoing sampling and characterizing the target action category can be directly used as the second original representation of this single individual regarding the target action category; or, in the case where the first total number is greater than one, it can be determined that the action is generated for multiple individuals, and then the second original representation characterizing the target action category can be copied the first total number of times to obtain the second original representations of multiple individuals respectively regarding the target action category. For example, when the first total number is 2, the second original representation characterizing the target action category can be copied into two second original representations, and these two second original representations respectively represent the second original representations of these two individuals regarding the target action category. Other cases can be deduced by analogy and will not be exemplified one by one here.

[0073] In a specific implementation scenario, please continue to refer to Figure 2 , similar to obtaining the first feature representation, in order to distinguish different second original representations in the case of a single individual and multiple individuals, the position information of each second original representation can be encoded on the basis of the second original representation to obtain the corresponding second feature representation. That is to say, similar to the first feature representation, the second feature representation also incorporates position encoding, and each position encoding is different. It should be noted that not only are the position encodings incorporated in each second feature representation different from each other, but also the position encodings incorporated in the second feature representation are different from the position encodings incorporated in the first feature representation. Specifically, in the case where the number of individuals is a single individual, the position encoding includes temporal position encoding. That is to say, in the case where the number of individuals is a single individual, the second feature representation can be distinguished from the first feature representation of different action frames in the temporal dimension. Exemplarily, still taking T action frames as an example, in the case of a single individual, the temporal position encoding (such as 1, 2,..., T) can be incorporated into the first original representations of these T action frames respectively, so as to obtain the first feature representations respectively characterizing the single individual in these T action frames, and then the temporal position encoding (such as T + 1) can be incorporated into the second original representation of the target action category, so as to obtain the second feature representation of the single individual regarding the target action category. Similarly, in the case where the number of individuals is multiple individuals, the position encoding can include temporal position encoding and individual position encoding. That is to say, in the case where the number of individuals is multiple individuals, it is necessary to make distinctions in both the temporal dimension and the individual dimension (such as Figure 2as shown by the dashed box after position encoding. Exemplarily, still taking T action frames as an example, in the case of multiple individuals, the second original representations of multiple individuals regarding the target action category can first incorporate temporal position encoding (e.g., T + 1), and further incorporate individual position encoding (e.g., 1) into the second original representation of the first individual regarding the target action category, and further incorporate individual position encoding (e.g., 2) into the second original representation of the second individual regarding the target action category, and so on, thereby combining temporal position encoding and individual position encoding, such that the representations of multiple individuals regarding the target action category are respectively fused with different position encodings (e.g., T + 1 - 1, T + 1 - 2, ……). In addition, the above position encoding is only an example. In the actual application process, an action generation model can be pre-trained, and the position encoding can be adjusted together with the network parameters of the action generation model during the training process of the action generation model until the training of the action generation model converges. After that, in the subsequent application process, the adjusted position encoding can be used. The above method can adopt different position encoding strategies to distinguish different feature representations in both single-individual and multi-individual application scenarios, making the position encodings of the feature representations different from each other, which is beneficial to improving the accuracy of the feature representations.

[0074] In one implementation scenario, as described above, both the first feature representation and the second feature representation incorporate position encoding. And in the case where the number of individuals is a single individual, the position encoding includes temporal position encoding. In the case where the number of individuals is multiple individuals, the position encoding includes individual position encoding and temporal position encoding. Specifically, reference can be made to Figure 2 the above description and will not be elaborated here. Further, for the sake of distinction, each position encoding can be different from each other. Taking T action frames and P (P equals 1, or P is greater than 1) individuals as an example, after the above operations, a feature representation of (T + 1) * P can be finally obtained, where it includes T * P first feature representations respectively representing each individual in each action frame, and P second feature representations respectively representing each individual regarding the target action category.

[0075] It should be noted that for each individual, the first original representation of the individual in a number of action frames and the second original representation regarding the target action category can be used as the original temporal representations of the individual at different time series. Still taking T action frames as an example, for the p-th individual, its first original representation in T action frames and the second original representation regarding the target action category can be regarded as its original temporal representations from the 1st time series to the (T + 1)-th time series. On this basis, in the action generation scenario of a single individual, when the range of time series t is from 1 to T, the temporal position encoding TPE of the t-th time series can be tAdd it to the t-th original temporal representation to obtain the first feature representation of the t-th time series. In the case where the time series t is T + 1, the temporal position encoding TPE of the t-th time series can be used t Add it to the t-th original temporal representation to obtain the second feature representation of the t-th time series. Similarly, in the action generation scenario of multiple individuals, the individual position encoding PPE of the p-th individual can be obtained first p Concatenate it with the temporal position encoding of the t-th time series to obtain the position encoding PE(t, p) of the p-th individual at the t-th time series = concat(TPE t , PPE p ). Among them, concat represents the concatenation operation. Then, in the case where the range of the time series t is from 1 to T, the position encoding PE(t, p) of the p-th individual at the t-th time series can be added to the original temporal representation of the p-th individual at the t-th time series to obtain the first feature representation of the p-th individual at the t-th time series. In the case where the time series t is T + 1, the position encoding PE(t, p) of the p-th individual at the t-th time series can be added to the original temporal representation of the p-th individual at the t-th time series to obtain the second feature representation of the p-th individual at the t-th time series. In addition, in addition to the combined encoding of the temporal position encoding and the individual position encoding described above, it is also possible not to distinguish between the temporal position encoding and the individual position encoding, but to use a completely independent fixed encoding. That is, for the action generation scenario of T action frames and P individuals, (T + 1) × P independent position encodings can be preset

[0076] Step S12: Perform relationship modeling based on the first feature representation and the second feature representation to obtain the fused feature representation of each individual in each action frame

[0077] In the embodiments of the present disclosure, the type of relationship modeling is related to the first total number of several individuals. Specifically, in the case where the first total number of several individuals is single, the relationship modeling includes modeling the temporal relationship between each action frame, so as to improve the temporal coherence between action frames by modeling the temporal relationship, which is beneficial to improving the authenticity of the action sequence. In the case where the first total number of several individuals is multiple, the relationship modeling includes modeling the interaction relationship between several individuals in each action frame and modeling the temporal relationship between each action frame, so as to improve the interaction rationality between individuals by modeling the interaction relationship, and can improve the temporal coherence between action frames by modeling the temporal relationship, which is beneficial to improving the authenticity of the action sequence

[0078] In an implementation scenario, when the first total number of several individuals is single, only the temporal relationship needs to be modeled. In this case, this single individual can be directly selected as the target individual, and the first feature representation and the second feature representation corresponding to the target individual are used as the temporal feature representations of the target individual at different time series. Exemplarily, still taking T action frames as an example, the first feature representation of the target individual in the first action frame can be used as the first temporal feature representation, the first feature representation of the target individual in the second action frame can be used as the second temporal feature representation,..., the first feature representation of the target individual in the T-th action frame can be used as the T-th temporal feature representation, and the second feature representation of the target individual regarding the target action category can be used as the (T + 1)-th temporal feature representation. On this basis, each time series can be respectively selected as the current time series, and the temporal feature representation of the current time series is selected as the current temporal feature representation, and based on the relevance between each reference temporal representation and the current temporal representation, the fused feature representation corresponding to the current temporal representation is obtained. That is to say, when the i-th temporal feature representation is used as the current temporal representation, the temporal feature representations of the target individual at each time series (i.e., 1 to T + 1) can be used as the reference temporal representations, and based on the relevance between these reference temporal representations and the i-th temporal feature representation respectively, the fused feature representation corresponding to the i-th temporal feature representation is obtained. Thus, in the action generation scenario of a single individual, finally T + 1 fused feature representations can be obtained. These T + 1 fused feature representations include: the feature representations of the single individual in T action frames after fusing the temporal relationship, and the feature representation of the single individual regarding the target action category after fusing the temporal relationship. It should be noted that for the convenience of distinguishing from the subsequent interaction relationship modeling steps, in temporal modeling, the current time series can be named the first current time series, the temporal feature representation of the current time series can be named the first current temporal representation, and the reference temporal representation can be named the first reference temporal representation.

[0079] In a specific implementation scenario, as described above, to improve the action generation efficiency, an action generation model can be pre-trained, and the action generation model can include a relationship modeling network, and the relationship modeling network can further include a temporal modeling sub-network. Exemplarily, the temporal modeling sub-network can be constructed based on Transformer. For the convenience of description, the Transformer included in the temporal modeling sub-network can be called T-Former. Then, for the foregoing T + 1 temporal feature representations, they can first be respectively linearly transformed to obtain the {query, key, value} feature representations corresponding to each temporal feature representation. Taking the t-th temporal feature representation F t as an example, through linear transformation, the corresponding {query, key, value} feature representations q t , k t , v t :

[0080] q t = W q F t ,k t = W k F t ,v t = W v F t ......(1)

[0081] In the above formula (1), W q ,W k , respectively represent linear transformation parameters and can be adjusted during the training process of the action generation model. On this basis, when selecting the t-th temporal feature representation as the current temporal representation, the correlation degree w between the query feature representation corresponding to the t-th temporal feature representation and the key feature representations of the t'-th (where the value range is from 1 to T + 1) temporal feature representations can be obtained t,t′ :

[0082] w t,t′ = q t ·k t ,......(2)

[0083] After obtaining the correlation degree w t,t ', based on this correlation degree w t,t , the value feature representations of the t'-th (where the value range is from 1 to T + 1) temporal feature representations are weighted to obtain the fused feature representation H after the t-th temporal feature representation fuses the temporal relationship t :

[0084]

[0085] In a specific implementation scenario, the temporal modeling sub-network can be formed by stacking L (L is greater than or equal to 1) layers of Transformer. On this basis, after obtaining the fused feature representation output by the l-th layer of Transformer, it can be used as the input of the (l + 1)-th layer of Transformer, and the aforementioned temporal modeling process is re-executed to obtain the fused feature representation output by the (l + 1)-th layer of Transformer And so on, finally, the fused feature representation Afterwards, since the first to the T-th final fused feature representations have been fully incorporated into the target action category, before generating the action in the subsequent step S13, the (T + 1)-th final fused feature representation related to the target action category can be discarded.

[0086] In an implementation scenario, when several individuals are multiple individuals, it is necessary to model the temporal relationship and the interaction relationship, and the interaction relationship and the temporal relationship can be modeled successively. Exemplarily, the interaction relationship can be modeled first, and then the temporal relationship; or, the temporal relationship can be modeled first, and then the interaction relationship. In addition, the output feature representation of the relationship modeled first is the input feature representation of the relationship modeled later. That is to say, when the relationship modeling includes modeling the interaction relationship and the temporal relationship, the relationship modeled first can be modeled based on the first feature representation and the second feature representation to obtain the output feature representation of the relationship modeled first, and then the relationship modeled later can be modeled based on the output feature representation to obtain the fused feature representation. It should be noted that the relationship modeled first is the interaction relationship, and the relationship modeled later is the temporal relationship, or the relationship modeled first is the temporal relationship, and the relationship modeled later is the temporal relationship.

[0087] In a specific implementation scenario, as described above, in order to improve the action generation efficiency, an action generation model can be pre-trained, and the action generation model can include a relationship modeling network, and the relationship modeling network can include a temporal modeling sub-network and an interaction modeling sub-network. Exemplarily, both the temporal modeling sub-network and the interaction modeling sub-network can be constructed based on Transformer. For the convenience of description, the Transformer included in the temporal modeling sub-network can be called T-Former, and the Transformer included in the interaction modeling sub-network can be called I-Former. Similar to the action generation scenario of the aforementioned single individual, in the action generation scenario of multiple individuals, one of the individuals can also be selected as the target individual. Exemplarily, the p-th individual among the P individuals can be selected as the target individual. On this basis, the first feature representation and the second feature representation corresponding to the target individual can be used as the temporal feature representations of the target individual at different time sequences. For the convenience of distinction, the first feature representation of the target individual at T action frames and the second feature representation regarding the target action category are respectively regarded as the temporal feature representations at time sequence 1 to time sequence T + 1. Then, for the aforementioned T + 1 temporal feature representations, they can first be linearly transformed respectively to obtain the {query, key, value} feature representations corresponding to each temporal feature representation. Taking the p-th individual selected as the target individual as an example, its t-th temporal feature representation can be linearly transformed to obtain the corresponding {query, key, value} feature representations

[0088]

[0089] When constructing the interaction relationship first, similar to the construction of the temporal relationship described above, after obtaining the temporal feature representations of the target individual at different time sequences, each time sequence can be selected as the current time sequence respectively, and the temporal feature representation of the current time sequence can be selected as the current time sequence representation. Then, based on the relevance between each reference time sequence representation and the current time sequence representation, the fused feature representation corresponding to the current time sequence representation can be obtained. Different from the construction of the temporal relationship, in the case of modeling the interaction relationship, the reference time sequence representation includes the temporal feature representations of each individual at the current time sequence. It should be noted that in order to distinguish from the modeling steps of the aforementioned temporal relationship, the current time sequence can be named the second current time sequence, the temporal feature representation of the second current time sequence can be named the second current time sequence representation, and the reference time sequence representation can be named the second reference time sequence representation. Specifically, the t-th time sequence can be used as the reference time sequence, and the temporal feature representations of each individual at the reference time sequence are the key feature representations of each individual at the reference time sequence respectively. Among them, the value range of p′ is from 1 to P. In this case, the relevance can be expressed as:

[0090]

[0091] Furthermore, based on this relevance the value feature representation of the p′-th (p′ ranges from 1 to P) individual at the t-th time sequence can be weighted to obtain the fused feature representation after fusing the interaction relationship of the temporal feature representation of the p-th individual at the t-th time sequence.

[0092]

[0093] In a specific implementation scenario, after obtaining the fused feature representations of each individual after fusing the interaction relationship at each time sequence then, as described above, these fused feature representations can be used as the input feature representations for constructing the temporal relationship to continue constructing the temporal relationship. The construction process of the temporal relationship can refer to the aforementioned relevant descriptions and will not be elaborated here.

[0094] In a specific implementation scenario, please refer to Figure 2 for reference. The I-Former for constructing the interaction relationship and the T-Former for constructing the temporal relationship can be combined into a group of Transformers to jointly construct the interaction relationship and the temporal relationship. Then the relationship construction network can include L groups of Transformers. On this basis, for the p-th individual in the action frame of the t-th time sequence, the fused feature representation output by the l-th group of Transformers After that, it can be used as the input of the (l + 1)-th layer of Transformer, and the aforementioned temporal modeling process is re-executed to obtain the fused feature representation output by the (l + 1)-th layer of Transformer. And so on, finally, the fused feature representation output by the last layer of Transformer can be used as the final fused feature representation. In addition, after obtaining the final fused feature representation since the first to the T-th final fused feature representations have fully incorporated the target action category, before generating the action in the subsequent step S13, the (T + 1)-th final fused feature representation related to the target action category can be discarded. In addition, please refer to Table 1 in combination. Table 1 is a schematic structural table of an embodiment of the action generation model. As shown in Table 1, the action generation model can exemplarily include 2 groups of Transformers. Of course, 3 groups of Transformers, 4 groups of Transformers, or 5 groups of Transformers, etc. can also be set, which is not limited herein. It should be noted that for the specific meanings of the input mapping layer and the category embedding layer, please refer to the specific acquisition processes of the aforementioned first feature representation and second feature representation respectively, which will not be elaborated herein. In addition, the action generation model shown in Table 1 is only a possible implementation manner in the actual application process, and the specific structure of the action generation model is not limited herein. For example, the number of input / output channels of each network layer shown in Table 1 can also be adaptively adjusted according to actual application needs.

[0095] Table 1 Schematic Structural Table of an Embodiment of the Action Generation Model

[0096]

[0097] It should be noted that, as can be seen from the foregoing embodiments, whether modeling the temporal relationship or the interaction relationship, the modeling processes of both tend to be similar, that is, an individual can be first selected as the target individual, and the first feature representation and the second feature representation corresponding to the target individual are used as the temporal feature representations of the target individual at different time sequences. Then, each time sequence is respectively selected as the current time sequence, and the temporal feature representation of the current time sequence is selected as the current time sequence representation. Then, based on the relevance between each reference time sequence representation and the current time sequence representation, the fused feature representation corresponding to the current time sequence representation is obtained. The difference between the two is that in the case of modeling the temporal relationship, the reference time sequence representation includes the temporal feature representations of the target individual at each time sequence, and in the case of modeling the interaction relationship, the reference time sequence representation includes the temporal feature representations of each individual at the current time sequence. Therefore, the temporal relationship and the interaction relationship can be modeled through a similar modeling process, so as to further improve the compatibility of the two application scenarios of a single individual and multiple individuals.

[0098] Step S13: Perform action mapping based on the fused feature representation to obtain action sequences of several individuals for the target action category.

[0099] In the embodiments of the present disclosure, the action sequence includes several action frames, and each action frame contains the action representations of each individual. Exemplarily, the action sequence may include T action frames, and several individuals are P individuals. Then each action frame contains the action representations of P individuals. Therefore, a temporally continuous three-dimensional action can be generated.

[0100] In an implementation scenario, as described above, to improve the action generation efficiency, an action generation model can be pre-trained, and the action generation model can include an action mapping network. As shown in Table 1, the action mapping network can specifically include linear layers such as fully connected layers. The specific structure of the action mapping network is not limited herein. On this basis, the fused feature representations of each individual in each action frame can be input into the action mapping network to obtain the action sequences of several individuals for the target action category. Taking T action frames and P individuals as an example, T*P fused feature representations can be obtained. Then the above T*P fused feature representations can be input into the action mapping network to obtain T action frames, and each action frame contains the action representations of P individuals. Thus, the T action frames can be combined in chronological order to obtain the action sequence. For ease of description, the action sequence can be represented as {M t |t ∈ [1,..., T]}, M t represents the t-th action frame, and each action frame M t contains the action representations of P individuals, that is

[0101] In an implementation scenario, the action representation of an individual in an action frame may include: in the action frame, the first position information of the key point (such as the hip) of the individual and the pose information of the individual. And the pose information can specifically include the second position information of several joint points (such as the left shoulder, right shoulder, left elbow, right elbow, left knee, right knee, left foot, right foot, etc.) of the individual. Exemplarily, taking the p-th individual in the t-th action frame as an example, the first position information can be denoted as which can specifically be the absolute position of the key point in the local coordinate system, and the pose information can be denoted as Specifically, it may include the position coordinates of each joint point in the local coordinate system. Exemplarily, each action frame in the action sequence can be represented as a tensor of size (P, C), that is, the action representation of each individual in the action frame can be represented by a C-dimensional vector. Based on this, the action sequence can be represented as a tensor of size (P, T, C). Of course, the above pose information can be expressed as the pose representation in SMPL (i.e., Skinned Multi Person Model), which is a widely used parametric human model. For its specific meaning, refer to the technical details of SMPL and will not be elaborated here.

[0102] In one implementation scenario, please refer to Figures 3a to 3f , Figures 3a to 3f which are all schematic diagrams of an embodiment of the action sequence. As Figures 3a to 3f shown, Figure 3a is the action sequence generated when the target action category is "toast", Figure 3b is the action sequence generated when the target action category is "take a photo", Figure 3c is the action sequence generated when the target action category is "support", Figure 3d is the action sequence generated when the target action category is "raid", Figure 3e is the action sequence generated when the target action category is "stretch", Figure 3f is the action sequence generated when the target action category is "dance".

[0103] In one implementation scenario, as Figures 3a to 3f shown, the action sequence generated by the action generation model only contains the action representations of each individual in each action frame, without including the appearance of each individual and the action scene. Therefore, after obtaining the action sequence, the appearance of each individual (such as hairstyle, clothing, hair color, etc.) can be freely designed according to needs, and the action scene (such as streets, shopping malls, parks, etc.) can also be freely designed according to needs. Exemplarily, after determining that the target action category is "take a photo" and the first total number of several individuals is 2, the action sequence as Figure 3b shown can be generated through the foregoing process. On this basis, the appearance of the left individual in Figure 3b (such as short hair, shirt, shorts, black hair, etc.) and the appearance of the right individual (long hair, dress, black hair, etc.) can be designed, and the action scene can be designed as "park", so as to further enrich the obtained animation, which can on the one hand improve the design flexibility and on the other hand greatly reduce the creation workload.

[0104] In the above solution, first obtain first feature representations respectively characterizing a plurality of individuals in a plurality of action frames, and obtain second feature representations respectively characterizing the plurality of individuals with respect to a target action category. On this basis, perform relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each individual in each action frame, and the type of relationship modeling is related to the first total number of the plurality of individuals. Then, perform action mapping based on the fused feature representations to obtain an action sequence of the plurality of individuals with respect to the target action category. The action sequence includes a plurality of action frames, and each action frame includes action representations of each individual. Therefore, on the one hand, actions can be automatically generated without relying on manual labor, and on the other hand, by specifically performing relationship modeling according to the first total number of the plurality of individuals, two application scenarios of a single individual and multiple individuals can be compatible. Therefore, it is possible to be compatible with two application scenarios of a single individual and multiple individuals while improving the action generation efficiency.

[0105] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an embodiment of a training method for an action generation model. As described above, the action sequence is obtained by the action generation model. To improve the training effect, the action generation model and the discriminative model can be obtained through generative adversarial training. The specific training process may include the following steps:

[0106] Step S41: Obtain a sample action sequence of a plurality of sample individuals with respect to a sample action category.

[0107] In the embodiments of the present disclosure, the sample action sequence includes a preset number of sample action frames, and the sample action sequence is labeled with a sample label, and the sample label indicates whether the sample action sequence is actually generated by the action generation model. Specifically, the sample action sequence may be generated by the action generation model or collected in a real scenario.

[0108] In one implementation scenario, please refer to Figure 5 , Figure 5 which is a schematic diagram for obtaining an embodiment of a sample action frame. As Figure 5 shown, a plurality of sample captured images of the sample individuals with respect to the sample action category can be obtained, that is, the real individuals can be photographed while demonstrating the sample action category. On this basis, the sample action representations of each sample individual in the sample captured images can be extracted. For example, the sample action representation of each sample individual may include the key points of the sample individual and the position information of a plurality of joint points. On this basis, each sample captured image can be represented as a sample action frame, and the sample action representation of each sample individual in each sample action frame, similar to the action representation in the foregoing disclosed embodiments, can be represented by a C-dimensional vector.

[0109] Step S42: Decompose each sample action frame in the sample action sequence to obtain sample graph data.

[0110] In the embodiments of the present disclosure, the sample graph data includes a preset number of node graphs. The node graphs are formed by connecting nodes. The nodes include key points and joint points. The node graphs include the node feature representations of each node, and the position feature representation of the node is obtained by splicing the position feature representations of several sample individuals at the corresponding node. Still taking the example that the sample action representation of each sample individual includes the position information of the key points and several joint points of the sample individual, the C-dimensional vector of the sample action representation can be decomposed into K D-dimensional vectors (as described above, the vector represents position information, such as position coordinates, etc.), and C = K × D, where K is the total number of key points and several joint points of the sample individual. For example, the total number of key points and several joint points of each sample individual is 18.

[0111] In one implementation scenario, please refer to Figure 6 , Figure 6 which is a schematic diagram of an embodiment of the sample graph data. As Figure 6 shown, for the scenario of a single sample individual, each node graph only needs to represent a single sample individual. Therefore, each node graph is formed by connecting K nodes, and each node on the node graph is expressed by the D-dimensional vector of the node. Therefore, each node graph can be represented as a tensor of size (K, D). Based on this, the sample graph data can be represented as a tensor of size (T, K, D).

[0112] In another implementation scenario, different from the scenario of a single sample individual, in the scenario of multiple sample individuals, each node graph needs to represent multiple sample individuals. At this time, each node graph is still formed by connecting K nodes, but each node on the node graph is obtained by splicing the D-dimensional vectors of multiple sample individuals at the node, that is, each node graph can be represented as a tensor of size (K, P·D). Based on this, the sample graph data can be represented as a tensor of size (T, K, P·D). In addition, for multiple sample individuals in the sample action sequence, if the sorting is different, it may lead to different prediction results of the subsequent discrimination model, thus bringing uncertainty to model training. To make up for this deficiency, when the sample action sequence is collected from the real scenario, the position feature representation of the node is obtained by splicing the position feature representations of several sample individuals at the corresponding node in a random order of several sample individuals, so that during the training process, the action generation model regards the situations of different sorting but actually belonging to the same sample action sequence as different samples and models them, so as to achieve data augmentation and thus be beneficial to improving the robustness of the model.

[0113] Step S43: Based on the discrimination model, discriminate the sample graph data and the sample action category to obtain a prediction result.

[0114] In an implementation scenario, the discrimination model can be constructed based on a spatio-temporal graph convolutional network. Exemplarily, please refer to Table 2, which is a structural schematic table of an embodiment of the discrimination model. It should be noted that Table 2 is only a possible implementation manner of the discrimination model in the actual application process, and does not thereby limit the specific structure of the discrimination model. In addition, for the specific meaning of the spatio-temporal convolution in Table 2, the relevant technical details of the spatio-temporal convolution can be referred to and will not be elaborated here.

[0115] Table 2 Structural Schematic Table of an Embodiment of the Discrimination Model

[0116]

[0117]

[0118] In the embodiments of the present disclosure, the prediction result includes a first prediction label and a second prediction label of the sample action sequence. The first prediction label represents the possibility that the sample action sequence is generated by the action generation model through prediction, and the second prediction label represents the possibility that the sample action sequence belongs to the sample action category. It should be noted that the first prediction label and the second prediction label can be represented by numerical values, and the larger the numerical value, the higher the corresponding possibility. Taking the discrimination model adopting the network structure shown in Table 2 as an example, the sample graph data can be denoted as x. After being processed by each spatio-temporal graph convolutional layer, a 512-dimensional vector φ(x) can be obtained. After the sample action category is represented by category embedding, a 512-dimensional vector y can also be obtained. The inner product of the two is φ(x)·y. Further, the vector φ(x) can be input into the output mapping layer to obtain Combined with the foregoing inner product φ(x)·y, the scores given by the discrimination model to the input sample action category and sample action sequence can be obtained, that is, the foregoing first prediction label and second prediction label.

[0119] Step S44: Adjust the network parameters of any one of the action generation model and the discrimination model based on the sample label, the first prediction label, and the second prediction label.

[0120] Specifically, the discrimination loss of the discrimination model can be measured by the first prediction tag and the sample tag, while the generation loss of the action generation model can be measured by the second prediction tag and the sample tag. During the training process, every time the discrimination model is trained M times (at this time, the network parameters of the discrimination model are adjusted), the action generation model is trained N times (at this time, the network parameters of the action generation model are adjusted). For example, every time the discrimination model is trained 4 times, the action generation model is trained 1 time, which is not limited here. On this basis, by training the discrimination model, the discrimination ability of the discrimination model for the sample action sequence can be improved (that is, the ability to distinguish the sample action sequence generated by the model and the sample action sequence collected in reality), which can promote the action generation model to improve the authenticity of the generated action sequence. And by training the action training model, the authenticity of the action sequence generated by the action generation model can be improved (that is, the action sequence generated by the model is as close as possible to the sample action sequence collected in reality), which in turn prompts the discrimination model to improve its discrimination ability. Furthermore, the discrimination model and the action generation model promote each other and complement each other. After several rounds of training, the model performance of the action generation model becomes increasingly excellent, and the discrimination model can no longer distinguish the action sequence generated by the action generation model and the sample action sequence collected in reality. At this point, the training can be ended. It should be noted that for the specific process of generative adversarial training, the specific technical details of generative adversarial training can be referred to and will not be elaborated here. In addition, as described in the foregoing disclosed embodiments, during the action generation process, position encoding can be performed, and during the training process of the action generation model, the position encoding can be adjusted together with the network parameters of the action generation model.

[0121] In the above solution, by using generative adversarial training to co-train the action generation model and the discrimination model, the action generation model and the discrimination model can promote each other during the co-training process and complement each other, which is ultimately beneficial to improving the model performance of the action generation model. In addition, by decomposing the sample action representation into sample graph data, the discrimination of the action sequence can be cleverly resolved into the discrimination of graph data, which is beneficial to greatly reducing the training complexity and the construction difficulty of the discrimination model.

[0122] Please refer to Figure 7 , Figure 7It is a framework schematic diagram of an embodiment of the action generation device 70 of the present application. The action generation device 70 includes: a feature acquisition module 71, a relationship modeling module 72, and an action mapping module 73. The feature acquisition module 71 is configured to acquire first feature representations respectively characterizing a plurality of individuals in a plurality of action frames, and acquire second feature representations respectively characterizing the plurality of individuals with respect to a target action category. The relationship modeling module 72 is configured to perform relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each individual in each action frame. Wherein, the type of relationship modeling is related to the first total number of the plurality of individuals. The action mapping module 73 is configured to perform action mapping based on the fused feature representations to obtain an action sequence of the plurality of individuals with respect to the target action category. Wherein, the action sequence includes a plurality of action frames, and each action frame includes action representations of each individual.

[0123] In the above solution, on the one hand, actions can be automatically generated without relying on manual labor. On the other hand, by specifically performing relationship modeling according to the first total number of a plurality of individuals, it can be compatible with two application scenarios of a single individual and multiple individuals. Therefore, it can be compatible with two application scenarios of a single individual and multiple individuals on the premise of improving the action generation efficiency.

[0124] In some disclosed embodiments, when the first total number of the plurality of individuals is single, the relationship modeling includes modeling the temporal relationship between each action frame; and / or, when the first total number of the plurality of individuals is multiple, the relationship modeling includes modeling the interaction relationship between a plurality of individuals in each action frame and modeling the temporal relationship between each action frame.

[0125] In some disclosed embodiments, the relationship modeling module 72 includes a temporal modeling sub-module. The temporal modeling sub-module includes a first selection unit configured to select an individual as a target individual, and use the first feature representation and the second feature representation corresponding to the target individual as the temporal feature representations of the target individual at different times, and use different times as the first current time respectively, and use the temporal feature representation of the first current time as the first current time representation. The temporal modeling sub-module includes a first representation fusion unit configured to obtain a fused feature representation corresponding to the first current time representation based on the relevance between each first reference time representation and the first current time representation. Wherein, the first reference time representation includes the temporal feature representations of the target individual at each time.

[0126] In some disclosed embodiments, the relationship modeling module 72 includes an interaction modeling sub-module, and the interaction modeling sub-module includes a second selection unit for selecting an individual as a target individual, and using the first feature representation and the second feature representation corresponding to the target individual as the temporal feature representations of the target individual at different time series, and using different time series as the second current time series respectively, and using the temporal feature representation of the second current time series as the second current time series representation; the interaction modeling sub-module includes a second fusion unit for obtaining a fused feature representation corresponding to the second current time series representation based on the correlation degrees between each second reference time series representation and the second current time series representation; wherein, the second reference time series representation includes the temporal feature representations of each individual at the second current time series respectively.

[0127] In some disclosed embodiments, when the relationship modeling includes modeling interaction relationships and temporal relationships, the relationship modeling module 72 includes a prior modeling sub-module for modeling a prior relationship based on the first feature representation and the second feature representation to obtain an output feature representation of the prior relationship, and the relationship modeling module 72 includes a subsequent modeling sub-module for modeling a subsequent relationship based on the output feature representation to obtain a fused feature representation; wherein, the prior relationship is an interaction relationship and the subsequent relationship is a temporal relationship, or the prior relationship is a temporal relationship and the subsequent relationship is an interaction relationship.

[0128] In some disclosed embodiments, the action sequence is obtained by an action generation model, the action generation model includes a relationship modeling network, and the relationship modeling network includes a temporal modeling sub-network for modeling temporal relationships and an interaction modeling sub-network for modeling interaction relationships.

[0129] In some disclosed embodiments, the first feature representation is obtained based on sampling of a Gaussian process.

[0130] In some disclosed embodiments, the feature acquisition module 71 includes a first acquisition sub-module, and the first acquisition sub-module includes a process sampling unit for respectively sampling a second total number of times in a plurality of Gaussian processes to obtain first original representations respectively characterizing a second total number of action frames; wherein, the length of the first original representation is the same as the number of Gaussian processes, and the characteristic length scales of each Gaussian process are different; the first acquisition sub-module includes a first acquisition unit for obtaining a third total number of first feature representations based on the first total number and the first original representation; wherein, the third total number is the product of the first total number and the second total number.

[0131] In some disclosed embodiments, the second feature representation is obtained based on mapping of a target action category.

[0132] In some disclosed embodiments, the feature acquisition module 71 includes a second acquisition sub-module, and the second acquisition sub-module includes an embedding representation unit for performing an embedding representation on the target action category to obtain a second original representation; the second acquisition sub-module includes a second acquisition unit for obtaining a first total number of second feature representations based on the first total number and the second original representation.

[0133] In some disclosed embodiments, both the first feature representation and the second feature representation are fused with position encoding; wherein, in the case where the number of individuals is a single individual, the position encoding includes a temporal position encoding, and in the case where the number of individuals is multiple individuals, the position encoding includes an individual position encoding and a temporal position encoding.

[0134] In some disclosed embodiments, the action sequence is obtained by an action generation model, and during the training process of the action generation model, the position encoding is adjusted together with the network parameters of the action generation model until the training of the action generation model converges.

[0135] In some disclosed embodiments, the action representation of an individual in an action frame includes: in the action frame, the first position information of the key points of the individual and the pose information of the individual, and the pose information includes the second position information of several joint points of the individual.

[0136] In some disclosed embodiments, the action sequence is obtained by an action generation model, and the action generation model and the discriminative model are obtained through generative adversarial training.

[0137] In some disclosed embodiments, the action generation device 70 includes a sample sequence acquisition module for acquiring sample action sequences of a number of sample individuals with respect to sample action categories; wherein, the sample action sequence includes a preset number of sample action frames, and the sample action sequence is labeled with a sample label, and the sample label indicates whether the sample action sequence is actually generated by an action generation model; the action generation device 70 includes a sample sequence decomposition module for decomposing each sample action frame in the sample action sequence to obtain sample graph data; wherein, the sample graph data includes a preset number of node graphs, the node graphs are formed by connecting nodes, the nodes include key points and joint points of the sample individuals, the node graphs include node feature representations of each node, and the position feature representation of the nodes is obtained by splicing the position feature representations of a number of sample individuals at the corresponding nodes respectively; the action generation device 70 includes a sample sequence discrimination module for discriminating the sample graph data and the sample action categories based on a discrimination model to obtain a prediction result; wherein, the prediction result includes a first prediction label and a second prediction label of the sample action sequence, the first prediction label indicates the possibility that the sample action sequence is predicted to be generated by an action generation model, and the second prediction label indicates the possibility that the sample action sequence belongs to the sample action category; the action generation device 70 includes a network parameter adjustment module for adjusting the network parameters of either the action generation model or the discrimination model based on the sample label, the first prediction label, and the second prediction label.

[0138] In some disclosed embodiments, when the sample action sequence is collected from a real scene, the position feature representation of the nodes is obtained by splicing the position feature representations of a number of sample individuals at the corresponding nodes in a random order of the number of sample individuals.

[0139] Please refer to Figure 8 , Figure 8 is a schematic framework diagram of an embodiment of the electronic device 80 of the present application. The electronic device 80 includes a mutually coupled memory 81 and a processor 82. The processor 82 is configured to execute program instructions stored in the memory 81 to implement the steps of any of the above-mentioned action generation method embodiments. In a specific implementation scenario, the electronic device 80 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 80 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.

[0140] Specifically, the processor 82 is used to control itself and the memory 81 to implement the steps of any of the above method embodiments for action generation. The processor 82 may also be referred to as a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with the ability to process signals. The processor 82 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 82 may be implemented jointly by integrated circuit chips.

[0141] In the above solution, on the one hand, actions can be automatically generated without relying on manual labor. On the other hand, by specifically performing relationship modeling based on the first total number of several individuals, it can be compatible with two application scenarios of a single individual and multiple individuals. Therefore, it can be compatible with two application scenarios of a single individual and multiple individuals on the premise of improving the action generation efficiency.

[0142] Please refer to Figure 9 , Figure 9 which is a schematic framework diagram of an embodiment of the computer-readable storage medium 90 of the present application. The computer-readable storage medium 90 stores program instructions 901 that can be run by a processor, and the program instructions 901 are used to implement the steps of any of the above method embodiments for action generation.

[0143] In the above solution, on the one hand, actions can be automatically generated without relying on manual labor. On the other hand, by specifically performing relationship modeling based on the first total number of several individuals, it can be compatible with two application scenarios of a single individual and multiple individuals. Therefore, it can be compatible with two application scenarios of a single individual and multiple individuals on the premise of improving the action generation efficiency.

[0144] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in electrical, mechanical or other forms.

[0145] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0147] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

Claims

1. A method for generating an action, characterized in that, Including: Obtaining first feature representations respectively characterizing a plurality of individuals in a plurality of action frames, and obtaining second feature representations respectively characterizing the plurality of individuals with respect to a target action category; wherein, the first feature representations are feature representations initialized by each of the individuals in each action frame, the second feature representations are mapped based on the target action category, and both the first feature representations and the second feature representations are fused with position encodings; Performing relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each of the individuals in each of the action frames; wherein, when the first total number of the plurality of individuals is one, the position encoding includes a temporal position encoding, and the relationship modeling includes modeling the temporal relationships between the action frames; or, when the first total number of the plurality of individuals is multiple, the position encoding includes an individual position encoding and the temporal position encoding, and the relationship modeling includes modeling the interaction relationships between the plurality of individuals in each of the action frames and modeling the temporal relationships between the action frames; Performing action mapping based on the fused feature representations to obtain an action sequence of the plurality of individuals with respect to the target action category; wherein, the action sequence includes the plurality of action frames, and the action frame includes action representations of each of the individuals.

2. The method according to claim 1, wherein When modeling the temporal relationships, the performing relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each of the individuals in each of the action frames includes: Selecting the individual as a target individual, and using the first feature representation and the second feature representation corresponding to the target individual as the temporal feature representations of the target individual at different times; Respectively selecting each of the times as a first current time, and using the temporal feature representation of the first current time as a first current time representation; Obtaining a fused feature representation corresponding to the first current time representation based on the correlations between each of the first reference time representations and the first current time representation; Wherein, the first reference time representation includes the temporal feature representations of the target individual at each of the times.

3. The method according to claim 1, characterized in that When modeling the interaction relationships, the performing relationship modeling based on the first feature representations and the second feature representations to obtain fused feature representations of each of the individuals in each of the action frames includes: Selecting the individual as a target individual, and using the first feature representation and the second feature representation corresponding to the target individual as the temporal feature representations of the target individual at different times; Respectively selecting each of the times as a second current time, and using the temporal feature representations of the second current time as second current time representations respectively; Obtaining a fused feature representation corresponding to the second current time representation based on the correlations between each of the second reference time representations and the second current time representation; Wherein, the second reference time representation includes the temporal feature representations of each of the individuals at the second current time.

4. The method according to claim 1, characterized in that, When the relationship modeling includes modeling the interaction relationship and the temporal relationship, the relationship modeling based on the first feature representation and the second feature representation to obtain the fused feature representation of each individual in each action frame includes: Modeling the prior relationship based on the first feature representation and the second feature representation to obtain the output feature representation of the prior relationship; Modeling the subsequent relationship based on the output feature representation to obtain the fused feature representation; Wherein, the prior relationship is the interaction relationship, the subsequent relationship is the temporal relationship, or the prior relationship is the temporal relationship, and the subsequent relationship is the interaction relationship.

5. The method according to claim 1, wherein The action sequence is obtained by an action generation model, the action generation model includes a relationship modeling network, and the relationship modeling network includes a temporal modeling sub-network and an interaction modeling sub-network. The temporal modeling sub-network is used to model the temporal relationship, and the interaction modeling sub-network is used to model the interaction relationship.

6. The method according to claim 1, wherein The first feature representation is obtained based on sampling of a Gaussian process.

7. The method according to claim 6, wherein The obtaining of the first feature representations respectively characterizing a plurality of individuals in a plurality of action frames includes: In a plurality of the Gaussian processes, sampling a second total number of times respectively to obtain first original representations respectively characterizing a second total number of the action frames; wherein, the length of the first original representation is the same as the number of the Gaussian processes, and the characteristic length scales of the Gaussian processes are different from each other; Based on the first total number and the first original representations, obtaining a third total number of the first feature representations; wherein, the third total number is the product of the first total number and the second total number.

8. The method according to claim 1, characterized in that, The obtaining of the second feature representations respectively characterizing the plurality of individuals with respect to a target action category includes: Performing an embedding representation on the target action category to obtain a second original representation; Based on the first total number and the second original representation, obtaining the first total number of the second feature representations.

9. The method according to claim 1, wherein The action sequence is obtained by an action generation model, and during the training process of the action generation model, the positional encoding is adjusted together with the network parameters of the action generation model until the training of the action generation model converges.

10. The method according to claim 1, characterized in that The action representation of the individual in the action frame includes: in the action frame, the first position information of the key points of the individual and the pose information of the individual, and the pose information includes the second position information of a plurality of joint points of the individual.

11. The method according to claim 1, wherein The action sequence is obtained by an action generation model, and the action generation model and the discriminative model are obtained through generative adversarial training.

12. The method according to claim 11, wherein The steps of the generative adversarial training include: Obtaining sample action sequences of a plurality of sample individuals with respect to a sample action category; wherein, the sample action sequence includes a preset number of sample action frames, and the sample action sequence is labeled with a sample label, and the sample label indicates whether the sample action sequence is actually generated by the action generation model. Decompose each of the sample action frames in the sample action sequence to obtain sample graph data; wherein, the sample graph data includes the preset number of node graphs, the node graphs are formed by connecting nodes, the nodes include the key points and joint points of the sample individuals, the node graphs include the node feature representations of each of the nodes, and the position feature representation of the nodes is obtained by splicing the position feature representations of the several sample individuals at the corresponding nodes respectively; Based on the discrimination model, discriminate the sample graph data and the sample action category to obtain a prediction result; wherein, the prediction result includes a first prediction label and a second prediction label of the sample action sequence, the first prediction label represents the possibility that the sample action sequence is generated by the action generation model through prediction, and the second prediction label represents the possibility that the sample action sequence belongs to the sample action category; Based on the sample label, the first prediction label and the second prediction label, adjust the network parameters of any one of the action generation model and the discrimination model.

13. The method according to claim 12, wherein When the sample action sequence is collected from a real scene, the position feature representation of the nodes is obtained by splicing the position feature representations of the several sample individuals at the corresponding nodes in a random order of the several sample individuals.

14. An action generation device, characterized in that, Comprising: A feature acquisition module, configured to acquire first feature representations respectively characterizing several individuals in several action frames, and acquire second feature representations respectively characterizing the several individuals with respect to a target action category; wherein, the first feature representations are feature representations initialized by each of the individuals in each action frame, the second feature representations are mapped based on the target action category, and both the first feature representations and the second feature representations are fused with position encodings; A relationship modeling module, configured to perform relationship modeling based on the first feature representations and the second feature representations to obtain the fused feature representations of each of the individuals in each of the action frames; wherein, when the first total number of the several individuals is single, the position encoding includes a temporal position encoding, and the relationship modeling includes modeling the temporal relationships between each of the action frames; or, when the first total number of the several individuals is multiple, the position encoding includes an individual position encoding and the temporal position encoding, and the relationship modeling includes modeling the interaction relationships between the several individuals in each of the action frames and modeling the temporal relationships between each of the action frames; An action mapping module, configured to perform action mapping based on the fused feature representations to obtain an action sequence of the several individuals with respect to the target action category; wherein, the action sequence includes the several action frames, and the action frames include the action representations of each of the individuals.

15. An electronic device, characterized in that, Comprising a memory and a processor coupled to each other, the processor is configured to execute program instructions stored in the memory to implement the action generation method according to any one of claims 1 to 13.

16. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the action generation method according to any one of claims 1 to 13 is implemented.