Animation generation method, device, electronic device and storage medium

By identifying and encoding virtual character actions and virtual scenes in video animations, combined with the method of generating three-dimensional animations, the problem of low video animation quality in the existing technology is solved, and a more natural role-scene interaction is achieved.

CN114037781BActive Publication Date: 2025-05-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111340369.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-05-09
Estimated Expiration
2041-11-12

Smart Images

  • Figure CN114037781B_ABST
    Figure CN114037781B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an animation generation method, device, electronic device and storage medium, the method comprising: obtaining input information, extracting input features from the input information; identifying the input features, determining the action information of a virtual character, and determining the action coding features corresponding to the action information; identifying the input features, determining a virtual object, generating a virtual scene according to the virtual object, and determining the scene coding features corresponding to the virtual scene; generating a three-dimensional animation according to the action coding features and the scene coding features. According to the scheme of the present disclosure, by processing the character animation and the virtual scene at the same time, the interaction between the virtual character and the virtual object can be made more natural and reasonable, thereby improving the quality of the video animation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an animation generation method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] In recent years, influenced by the development of mobile devices and social media technology, short videos have been rapidly promoted and popularized. Compared with traditional text or picture introductions, video animations with rich content have become a more popular information dissemination medium for users due to their realistic features.

[0003] In the related art, video animation includes character animation and virtual scene. Character animation can be obtained by using phase function neural network, neural state machine, local action phase and other processing. Virtual scene can be generated by using RGB-D (depth image), text and other information as input. However, in the related art, since character animation and virtual scene are usually processed separately to generate video animation, the problem of low video animation quality is prone to occur. Summary of the invention

[0004] The present disclosure provides an animation generation method, device, electronic device, computer-readable storage medium, and computer program product to at least solve the problem of low quality of video animation in the related art. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an animation generation method, comprising:

[0006] Acquire input information, and extract input features from the input information;

[0007] Identifying the input features, determining action information of the virtual character, and determining action coding features corresponding to the action information;

[0008] Identify the input features, determine the virtual objects, generate a virtual scene according to the virtual objects, and determine the scene coding features corresponding to the virtual scene;

[0009] A three-dimensional animation is generated according to the action coding features and the scene coding features.

[0010] In one of the embodiments, the motion information includes displacement motion information and non-displacement motion information;

[0011] The step of identifying the input feature, determining the action information of the virtual character, and determining the action coding feature corresponding to the action information includes:

[0012] Classifying and identifying the input features to obtain displacement action categories, and acquiring displacement action information corresponding to the displacement action categories, wherein the displacement action information is generated by using an action generation network;

[0013] Determine an action identifier that matches the input feature from a non-displacement action database, and generate non-displacement action information according to the matched action identifier;

[0014] The action coding feature is generated according to the displacement action information and the non-displacement action information.

[0015] In one embodiment, the generating non-displacement action information according to the matched action identifier includes:

[0016] Searching for a motion capture file corresponding to the motion identifier from a non-displacement motion database;

[0017] The motion capture file is fully connected and feature extracted from the motion capture file after the full connection processing to obtain the non-displacement motion information.

[0018] In one embodiment, the step of identifying the input features, determining the virtual objects, and generating the virtual scenes according to the virtual objects includes:

[0019] Classifying and identifying the input features to obtain the virtual object;

[0020] A graph network structure of the virtual objects is constructed according to the virtual objects and virtual object relationships, wherein the virtual object relationships are used to represent association relationships between different virtual objects;

[0021] A virtual scene matching the graph network structure is determined from existing virtual scenes.

[0022] In one embodiment, the determining of the scene coding feature corresponding to the virtual scene includes:

[0023] Segmenting the virtual object in the virtual scene to obtain a plurality of segmented blocks of preset sizes;

[0024] The segmented blocks corresponding to the virtual objects in the virtual scene are input into a coding network to obtain the scene coding features.

[0025] In one embodiment, generating a three-dimensional animation according to the action coding feature and the scene coding feature includes:

[0026] Decoding the action coding features and the scene coding features through a pre-trained action generation and decoding network to generate an initial three-dimensional animation, wherein the initial three-dimensional animation includes an initial character action of the virtual character and an initial virtual scene, wherein the initial virtual scene includes a virtual object;

[0027] Controlling the virtual character to perform the initial character action in the initial virtual scene, obtaining the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generating a loss result according to the distance;

[0028] The initial character action and the initial virtual scene are adjusted according to the loss result to obtain the three-dimensional target animation.

[0029] In one embodiment, the action generation decoding network includes a weight learning network and an action prediction network;

[0030] The step of decoding the action coding features and the scene coding features through a pre-trained action generation and decoding network to generate an initial three-dimensional animation includes:

[0031] Mapping the action coding features and the scene coding features to a preset number of weight parameters through the weight learning network, wherein the weight parameters are used to distinguish the action categories of the virtual character;

[0032] The weight parameters, the action coding features and the scene coding features are input into the action prediction network, and the skeletal motion trajectory of the virtual character in the next frame of the current frame is obtained through the action prediction network to generate the initial three-dimensional animation.

[0033] In one embodiment, obtaining the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generating a loss result according to the distance includes:

[0034] When a first virtual object collides with the virtual character, determining a penetration distance between the virtual character and the first virtual object, and generating a penetration loss result according to the penetration distance;

[0035] The loss result is generated according to the mode penetration loss result.

[0036] In one of the embodiments, the virtual object is divided into a plurality of division blocks of preset sizes;

[0037] The determining the penetration distance between the virtual character and the first virtual object includes:

[0038] The number of segments that the virtual character passes through when colliding with the first virtual object is obtained, and the penetration distance is determined according to the number.

[0039] In one embodiment, generating a penetration loss result according to the penetration distance includes:

[0040] The ratio between the penetration distance corresponding to each of the first virtual objects and the total number of segmented blocks is obtained, and the sum of the ratios corresponding to the first virtual objects is obtained; the sum of the ratios is normalized to generate the penetration loss result.

[0041] In one embodiment, the obtaining of the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generating a loss result according to the distance, further includes:

[0042] When there is a contact relationship between the second virtual object and the virtual character, determining a contact distance between a target joint point of the virtual character and the second virtual object, and generating a contact loss result according to the contact distance;

[0043] The step of generating the loss result according to the mode penetration loss result includes:

[0044] The loss result is generated according to the mode penetration loss result and the contact loss result.

[0045] In one of the embodiments, a contact point corresponding to the target joint point is marked on the second virtual object, and the contact point is a point that the target joint point should contact when the virtual character interacts with the second virtual object;

[0046] The determining the contact distance between the target joint point of the virtual character and the second virtual object includes:

[0047] Acquire the position where the target joint point appears in the initial virtual scene during the complete process of executing the initial character action;

[0048] Obtaining an average joint point position of the target joint point according to the positions that appeared in the initial virtual scene;

[0049] The contact distance is determined according to the distance between the average position of the joint points and the contact point.

[0050] In one embodiment, generating a contact loss result according to the contact distance includes:

[0051] Obtaining a second norm of a contact distance of each of the second virtual objects;

[0052] The sum of the second norms of the second virtual object is obtained, and the sum of the second norms is normalized to generate the contact loss result.

[0053] In one embodiment, generating the loss result according to the mode penetration loss result and the contact loss result includes:

[0054] Obtaining a penetration coefficient corresponding to the penetration loss result and a contact coefficient corresponding to the contact loss result;

[0055] The loss result is generated according to the penetration loss result and the penetration coefficient, and the contact loss result and the contact coefficient.

[0056] In one embodiment, the adjusting the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional animation includes:

[0057] Adjusting the initial character action and the initial virtual scene according to the loss result, and controlling the virtual character to perform the adjusted initial character action in the adjusted initial virtual scene;

[0058] Acquire a current distance between the virtual character and the virtual object when the virtual character performs the adjusted initial character action, and generate a current loss result according to the current distance;

[0059] The above steps are iterated until a preset stop condition is met to obtain the three-dimensional animation.

[0060] In one embodiment, the adjusting the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional animation includes:

[0061] Fine-tuning the initial character action to obtain the fine-tuned initial character action;

[0062] Fine-tuning the initial virtual scene to obtain the fine-tuned initial virtual scene;

[0063] Generate a reference loss result according to the fine-tuned initial character action and the fine-tuned initial virtual scene;

[0064] When the relationship between the loss result and the reference loss result satisfies a preset condition, adjusting the initial character action and the initial virtual scene respectively according to the loss result to obtain the three-dimensional animation;

[0065] When the relationship between the loss result and the reference loss result does not satisfy the preset condition, the initial three-dimensional animation is used as the three-dimensional animation.

[0066] According to a second aspect of an embodiment of the present disclosure, there is provided an animation generating device, comprising:

[0067] An information acquisition module is configured to acquire input information and extract input features from the input information;

[0068] An action feature generation module is configured to perform recognition of the input feature, determine action information of the virtual character, and determine an action coding feature corresponding to the action information;

[0069] A scene feature generation module is configured to perform recognition of the input features, determine virtual objects, generate a virtual scene according to the virtual objects, and determine a scene coding feature corresponding to the virtual scene;

[0070] The animation generation module is configured to generate a three-dimensional animation according to the action coding feature and the scene coding feature.

[0071] In one embodiment, the motion information includes displacement motion information and non-displacement motion information; the motion feature generation module includes:

[0072] A first classification unit is configured to perform classification and recognition on the input feature to obtain a displacement action category and obtain displacement action information corresponding to the displacement action category, wherein the displacement action information is generated by using an action generation network;

[0073] A first matching unit is configured to determine an action identifier matching the input feature from a non-displacement action database, and generate non-displacement action information according to the matched action identifier;

[0074] The feature generation unit is configured to generate the action coding feature according to the displacement action information and the non-displacement action information.

[0075] In one of the embodiments, the first matching unit is configured to search for a motion capture file corresponding to the motion identifier from a non-displacement motion database; perform full-connection processing on the motion capture file, and perform feature extraction on the fully-connected motion capture file to obtain the non-displacement motion information.

[0076] In one embodiment, the scene feature generation module includes:

[0077] A second classification unit is configured to perform classification and recognition on the input feature to obtain the virtual object;

[0078] a structure construction unit configured to execute construction according to the virtual objects and virtual object relationships to obtain a graph network structure of the virtual objects, wherein the virtual object relationships are used to represent association relationships between different virtual objects;

[0079] The second matching unit is configured to determine a virtual scene that matches the graph network structure from existing virtual scenes.

[0080] In one embodiment, the scene feature generation module includes:

[0081] A segmentation unit, configured to segment the virtual object in the virtual scene to obtain a plurality of segmented blocks of preset sizes;

[0082] The scene feature encoding unit is configured to input the segmented blocks corresponding to the virtual objects in the virtual scene into an encoding network to obtain the scene encoding features.

[0083] In one embodiment, the animation generation module includes:

[0084] A decoding unit is configured to perform decoding processing on the action coding features and the scene coding features through a pre-trained action generation decoding network to generate an initial three-dimensional animation, wherein the initial three-dimensional animation includes an initial character action of the virtual character and an initial virtual scene, wherein the initial virtual scene includes a virtual object;

[0085] an action execution unit, configured to execute and control the virtual character to execute the initial character action in the initial virtual scene, obtain the distance between the virtual character and the virtual object when the virtual character executes the initial character action, and generate a loss result according to the distance;

[0086] The adjustment unit is configured to adjust the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional target animation.

[0087] In one embodiment, the action generation decoding network includes a weight learning network and an action prediction network; the decoding unit includes:

[0088] A weight learning subunit, configured to execute mapping of the action coding features and the scene coding features to a preset number of weight parameters through the weight learning network, wherein the weight parameters are used to distinguish the action categories of the virtual character;

[0089] The action prediction subunit is configured to input the weight parameters, the action coding features and the scene coding features into the action prediction network, obtain the skeleton motion trajectory of the virtual character in the next frame of the current frame through the action prediction network, and generate the initial three-dimensional animation.

[0090] In one embodiment, the action execution unit includes:

[0091] A first action execution subunit is configured to determine a penetration distance between the virtual character and the first virtual object when a first virtual object collides with the virtual character, and generate a penetration loss result according to the penetration distance;

[0092] The loss result determination subunit is configured to generate the loss result according to the penetration loss result.

[0093] In one of the embodiments, the virtual object is divided into a plurality of blocks of preset sizes; the first action execution subunit is configured to obtain the number of blocks through which the virtual character collides with the first virtual object, and determine the penetration distance according to the number.

[0094] In one of the embodiments, the first action execution subunit is configured to execute obtaining the ratio between the penetration distance corresponding to each of the first virtual objects and the total number of segmented blocks, and obtaining the sum of the ratios corresponding to the first virtual objects; normalizing the sum of the ratios to generate the penetration loss result.

[0095] In one embodiment, the action execution unit further includes:

[0096] A second action execution subunit is configured to determine a contact distance between a target joint point of the virtual character and the second virtual object when there is a contact relationship between the second virtual object and the virtual character, and generate a contact loss result according to the contact distance;

[0097] The loss result determination subunit is configured to generate the loss result according to the penetration loss result and the contact loss result.

[0098] In one of the embodiments, a contact point corresponding to the target joint point is pre-marked on the second virtual object, and the contact point is a point that the target joint point should contact when the virtual character interacts with the second virtual object;

[0099] The second action execution sub-unit is configured to obtain the position of the target joint point where it appeared in the initial virtual scene during the complete process of executing the initial character action; obtain the average joint point position of the target joint point based on the position where it appeared in the initial virtual scene; and determine the contact distance based on the distance between the average joint point position and the contact point.

[0100] In one of the embodiments, the second action execution subunit is configured to execute obtaining the second norm of the contact distance of each second virtual object; obtaining the sum of the second norms of the second virtual objects, normalizing the sum of the second norms, and generating the contact loss result.

[0101] In one embodiment, the loss result determination subunit is configured to execute acquisition of a penetration coefficient corresponding to the penetration loss result and a contact coefficient corresponding to the contact loss result; and generate the loss result based on the penetration loss result and the penetration coefficient, and the contact loss result and the contact coefficient.

[0102] In one embodiment, the adjustment unit comprises:

[0103] A first adjustment subunit is configured to adjust the initial character action and the initial virtual scene according to the loss result, and control the virtual character to perform the adjusted initial character action in the adjusted initial virtual scene;

[0104] A loss result generating subunit is configured to obtain a current distance between the virtual character and the virtual object when the virtual character performs the adjusted initial character action, and generate a current loss result according to the current distance;

[0105] The iterative updating unit is configured to iteratively execute the above steps until a preset stop condition is met to obtain the three-dimensional animation.

[0106] In one embodiment, the adjustment unit comprises:

[0107] A first fine-tuning subunit is configured to perform fine-tuning on the initial character action to obtain the fine-tuned initial character action;

[0108] A second fine-tuning subunit is configured to perform fine-tuning on the initial virtual scene to obtain the fine-tuned initial virtual scene;

[0109] A reference result generating subunit is configured to generate a reference loss result according to the fine-tuned initial character action and the fine-tuned initial virtual scene;

[0110] The second adjustment subunit is configured to execute, when the relationship between the loss result and the reference loss result meets a preset condition, adjusting the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional animation; when the relationship between the loss result and the reference loss result does not meet the preset condition, using the initial three-dimensional animation as the three-dimensional animation.

[0111] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0112] processor;

[0113] a memory for storing instructions executable by the processor;

[0114] Wherein, the processor is configured to execute the instructions to implement the animation generation method as described in any one of the first aspects above.

[0115] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the animation generation method as described in any one of the above-mentioned first aspects.

[0116] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes instructions, and is characterized in that when the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the animation generation method as described in any one of the above-mentioned first aspects.

[0117] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0118] Extract input features from the acquired input information; identify the input features, determine action information of the virtual character, and determine action coding features corresponding to the action information; identify the input features, determine the virtual object, generate a virtual scene based on the virtual object, and determine the scene coding features corresponding to the virtual scene; generate a three-dimensional animation based on the action coding features and the scene coding features, and by simultaneously processing the action coding features representing the character animation and the virtual scene representing the virtual scene, the interaction between the virtual character and the virtual object can be made more natural and reasonable, thereby improving the quality of the video animation.

[0119] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0120] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0121] Figure 1 The figure is a flowchart of an animation generating method according to an exemplary embodiment.

[0122] Figure 2 The figure is a flowchart showing steps of generating action coding features according to an exemplary embodiment.

[0123] Figure 3 The figure is a flowchart showing steps of generating a virtual scene according to an exemplary embodiment.

[0124] Figure 4 The figure is a flowchart showing steps of generating a three-dimensional animation according to an exemplary embodiment.

[0125] Figure 5 The figure is a flowchart showing steps of generating an initial 3D animation according to an exemplary embodiment.

[0126] Figure 6 The figure is a flowchart showing the steps of determining loss results according to an exemplary embodiment.

[0127] Figure 7 The figure is a flowchart showing the steps of adjusting the initial character action and the initial virtual scene according to an exemplary embodiment.

[0128] Figure 8 The figure is a flowchart of an animation generating method according to an exemplary embodiment.

[0129] Fig. 9 It is a schematic diagram showing an animation generating method according to an exemplary embodiment.

[0130] Fig.10 The invention is a block diagram of an animation generating device according to an exemplary embodiment.

[0131] Fig.11 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0132] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0133] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0134] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0135] The animation generation method provided by the present disclosure can be applied to a terminal. The terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, portable wearable devices. An application capable of producing three-dimensional animations is deployed in the terminal. The application can be an application dedicated to generating three-dimensional animations, or it can be a plug-in program, a small program, etc. mounted on the main application. Specifically, the terminal obtains input information and extracts input features from the input information. The input features are identified to determine the action information of the virtual character, and the action coding features corresponding to the action information are determined. The input features are identified to determine the virtual objects, a virtual scene is generated based on the virtual objects, and the scene coding features corresponding to the virtual scene are determined. The terminal generates a three-dimensional animation based on the action coding features and the scene coding features.

[0136] Furthermore, the animation generation method provided by the present disclosure can be implemented offline or online.

[0137] Furthermore, the animation generation method provided by the present disclosure can also be applied to a server. Compared with the application to a terminal, the difference between the two is mainly the difference in the execution subject, and the implementation principle and logic are similar.

[0138] Figure 1 is a flowchart of an animation generation method according to an exemplary embodiment. Figure 1 As shown, the animation generation method is used in a terminal and includes the following steps.

[0139] In step S110, input information is acquired, and input features are extracted from the input information.

[0140] The input information may include, but is not limited to, instruction information for character animation and virtual scene in 3D animation, and may be any one of text, voice, image, video, etc. For example, the input voice is “make an animation of running indoors”. The input information may be information collected in real time, or may be information obtained from a database of a client or server.

[0141] Specifically, a pre-trained feature extraction network is pre-deployed in the client. The feature extraction network may correspond to the information type of the input information, so that the input information can be feature extracted to obtain input features of the input information. In an example, the input information is an image, and the feature extraction network may use a residual network.

[0142] In one embodiment, the input information is text or speech. For example, the maximum sequence length of the text or speech can be pre-configured, and after obtaining the input information, the input information is sequence encoded and filled. The filled input information is input into the bidirectional encoder representation network (Bidirectional Encoder Representations from Transformers, BERT). The intermediate coding features are obtained. The intermediate coding features are input into the long short-term memory network unit to model the relationship between multiple sentences in the input information, obtain a high-dimensional coding feature of a fixed length, and use the high-dimensional coding feature of the fixed length as the input feature. Among them, BERT is based on the Transformer (a network based on an attention mechanism) structure, which can better model the relationship between long input sequences compared to traditional recurrent neural network units and long short-term memory network units.

[0143] In step S120, the input features are identified, the action information of the virtual character is determined, and the action coding features corresponding to the action information are determined.

[0144] The virtual character can be a pre-configured 3D character model, and the user can select a suitable 3D character model from a database for use according to actual needs. The virtual character can also be obtained by recognizing and converting the user's real portrait, for example, the corresponding virtual character is obtained by recognizing the user's real portrait based on deep learning theory.

[0145] Action information may include, but is not limited to, high-dimensional coding features generated based on the character actions of the virtual character and / or the environmental information that generates the character actions. Among them, the character actions may include, but are not limited to, displacement actions that can generate displacement, such as running, walking, crawling, etc., and / or non-displacement actions that do not generate displacement, such as clapping, waving, etc. The character actions may include, but are not limited to, the rotation angles that drive the various skeletal joints of the virtual character to move. The skeletal joints can be predefined, such as wrist joints, spine joints, clavicle joints, shoulder joints, elbow joints, ankle joints, hip joints, etc. The environmental information may include, but is not limited to, objects in the environment when the action is collected. For example, if the action is sitting, the environmental information may be the size, shape, and position of the chair.

[0146] Specifically, after obtaining the input features, the client determines the action information that matches the input features. In one example, an action classification network can be deployed in the client. After obtaining the input features of the input information, the input features are input into the action classification network to obtain the action category, and then the action information corresponding to the action category is obtained from the pre-stored action database. In another example, the client can calculate the similarity between the input features and multiple existing action identifiers, obtain the action identifier with the highest similarity, and then obtain the action information corresponding to the action identifier with the highest similarity from the action database.

[0147] In step S130, the input features are identified, the virtual objects are determined, a virtual scene is generated according to the virtual objects, and a scene coding feature corresponding to the virtual scene is determined.

[0148] The virtual scene may be an initial scene for the virtual character to perform role actions, and the display mode may be static or dynamic, such as a snow mountain scene, a desert scene, a sea scene, a castle scene, an indoor scene, etc. The virtual scene includes at least one virtual object, such as a building, a sofa, a prop, a tree, a road, a musical instrument, a bed, a table and a chair, etc.

[0149] Specifically, after obtaining the input features, the client determines the virtual object that matches the input features. In one example, an object classification network can be deployed in the client. After obtaining the input features of the input information, the input features are input into the object classification network to obtain a virtual object (i.e., an object category). In another example, the client can calculate the similarity between the input features and multiple existing virtual objects to obtain the virtual object with the highest similarity. The client generates a virtual scene corresponding to the obtained virtual object based on the association relationship between the pre-configured virtual objects and virtual scenes. The virtual scene is encoded to obtain a scene encoding feature.

[0150] In step S140, a three-dimensional animation is generated according to the action coding features and the scene coding features.

[0151] Specifically, the client decodes the action coding features and the scene coding features together, predicts the skeletal motion trajectory of the virtual character in the virtual scene, and then generates a three-dimensional animation according to the virtual scene and the skeletal motion trajectory of the virtual character.

[0152] In the above animation generation method, input features are extracted from the acquired input information; the input features are identified to determine the action information of the virtual character, and the action coding features corresponding to the action information are determined; the input features are identified to determine the virtual object, a virtual scene is generated based on the virtual object, and the scene coding features corresponding to the virtual scene are determined; three-dimensional animation is generated based on the action coding features and the scene coding features. By simultaneously processing the action coding features representing the character animation and the virtual scene representing the virtual scene, the interaction between the virtual character and the virtual object can be made more natural and reasonable, thereby improving the quality of the video animation.

[0153] In an exemplary embodiment, when the motion information includes displacement motion information and non-displacement motion information, such as Figure 2 As shown, step S120, identifying the input features, determining the action information of the virtual character, and determining the action coding features corresponding to the action information, includes:

[0154] In step S210, the input features are classified and identified to obtain displacement action categories, and displacement action information corresponding to the displacement action categories is obtained, wherein the displacement action information is generated by using an action generation network.

[0155] Among them, the action generation network can include multiple layers of linear layers (i.e., fully connected layers) connected in sequence. The action generation network is pre-trained using a number of displacement actions, and it has the ability to generate actions based on text instructions and binding poses (T-POSE). The action generation network can be pre-adopted to generate displacement action encoding features (a high-dimensional encoding feature) based on text instructions and binding poses as displacement action information, and the displacement action information and the corresponding displacement action categories are stored in the displacement action database. Among them, the binding pose refers to the preset pose used to bind the skeleton of a three-dimensional model in the animation.

[0156] Specifically, since the number of displacement actions is usually relatively fixed, after obtaining the input features, the client can use the action classification network to classify and identify the input features to obtain the displacement action category. Then, the displacement action information corresponding to the displacement action category is obtained from the displacement action data.

[0157] In step S220, an action identifier matching the input feature is determined from a non-displacement action database, and non-displacement action information is generated according to the matching action identifier.

[0158] The non-displacement action database stores a plurality of non-displacement actions (motion capture files) and action identifiers (such as action names, description information, etc.) corresponding to each non-displacement action information. The non-displacement action database can be generated based on existing open source data sets, such as Mixamo 3D (a character action library), AMASS (a character action library), etc., and a custom data set can also be used. The actions in the non-displacement action database can be saved as BVH files.

[0159] Specifically, after the client obtains the input feature, it calculates the similarity between the input feature and each action identifier in the database, and determines the action identifier with a similarity greater than a threshold (e.g., 0.85) as the matched action identifier. The motion capture file corresponding to the matched action identifier is obtained. The similarity can be characterized by cosine distance, Euclidean distance, etc.

[0160] The human skeleton mapping network is pre-deployed in the client. The input data format of the human skeleton mapping network is consistent with that of the action generation network. The human skeleton mapping network includes at least one linear layer and a feature extraction unit connected in sequence. Among them, the feature extraction unit can be, but is not limited to, a gated recurrent unit (GRU), a long short-term memory network (LSTM), etc.

[0161] After obtaining the motion capture file, the client inputs the motion capture file into the human skeleton mapping network, and performs full connection processing on the rotation angles and hierarchical relationships of all bone nodes in the motion capture file based on the spatial relationship of the human skeleton joints in the motion capture file through the linear layer. Then, the extracted features are input into the linear layer again for full connection processing based on the bone hierarchy. Finally, the features obtained after two full connection processes are input into the feature extraction unit to obtain the displacement coding features (a high-dimensional coding feature) of the motion capture file as non-displacement motion information.

[0162] In step S230, an action coding feature is generated according to the displacement action information and the non-displacement action information.

[0163] Specifically, the client uses the displacement action coding features and the non-displacement action coding features as action coding features.

[0164] In this embodiment, the actions of the virtual character are divided into displacement actions and non-displacement actions. On the one hand, the acquired actions have diversified characteristics; on the other hand, the character actions of the virtual character can be quickly acquired, which helps to improve the efficiency of animation generation.

[0165] In an exemplary embodiment, if Figure 3 As shown, in step S130, the input features are identified, virtual objects are determined, and a virtual scene is generated according to the virtual objects, including:

[0166] In step S310, the input features are classified and identified to obtain a virtual object.

[0167] Specifically, after obtaining the input features, the client inputs the input features into the object classification network to obtain the virtual object (ie, the virtual object category) output by the object classification network.

[0168] In step S320, a graph network structure of virtual objects is constructed based on the virtual objects and the virtual object relationships.

[0169] Among them, the virtual object relationship is used to characterize the association relationship between different virtual objects. A scene database including several virtual scenes can be pre-deployed. The virtual objects in the scene database include corresponding attribute descriptions, which may include but are not limited to object category, material information, object purpose, object name, etc. At the same time, each virtual object in the scene database does not exist alone, but will be placed in a specific virtual scene. The association relationship between each virtual object and other virtual objects (i.e., virtual object relationship) can be constructed in advance according to the virtual scene to which the virtual object belongs. When loading a virtual scene from the scene database, the position information, size, object annotation information, etc. of the virtual object in the virtual scene, as well as the virtual object relationship with other objects in the virtual scene can be loaded synchronously.

[0170] Specifically, after the client obtains the virtual object, it obtains the pre-built virtual object relationship corresponding to the virtual object, and constructs a graph network structure of the virtual object according to the virtual object and the virtual object relationship.

[0171] In step S330, a virtual scene matching the graph network structure is determined from existing virtual scenes.

[0172] Specifically, each virtual scene in the scene database is pre-generated and processed to generate a semantic scene graph for each virtual scene. The client matches the graph network structure with each semantic scene graph to determine the virtual scene corresponding to the graph network structure. At the same time, the position and size of the virtual object in the virtual scene are determined according to the object parameters (e.g., position information, size, etc.) of the virtual object output by the object classification network.

[0173] In this embodiment, by constructing a graph network structure according to virtual objects and searching for a virtual scene that matches the graph network structure, the scene indication information in the input information can be fully utilized, thereby obtaining an accurate virtual scene.

[0174] In an exemplary embodiment, in step S130, determining a scene coding feature corresponding to a virtual scene includes: segmenting a virtual object in the virtual scene to obtain a plurality of segmented blocks of a preset size; and inputting the segmented blocks corresponding to the virtual object in the virtual scene into a coding network to obtain a scene coding feature.

[0175] Specifically, for each virtual object in the virtual scene, the client divides each virtual object in the virtual scene into multiple segments according to a cube of a preset size. After all virtual objects in the virtual scene are segmented, the client inputs the segmented virtual scene into the encoding network to obtain a high-dimensional encoding feature of the virtual scene as a scene encoding feature.

[0176] In this embodiment, by dividing each virtual object in the virtual scene into multiple blocks according to fixed-size cubes, compared to using only one collision box to describe the virtual object, the collision situation between the virtual character and the virtual object, especially the virtual object with concave surfaces, can be accurately estimated, which helps to improve the collision situation between the virtual character and the virtual object.

[0177] In an exemplary embodiment, if Figure 4 As shown, step S140, generating a three-dimensional animation according to the action coding features and the scene coding features, includes:

[0178] In step S410, the action coding features and the scene coding features are decoded by a pre-trained action generation and decoding network to generate an initial three-dimensional animation.

[0179] The initial three-dimensional animation includes the initial character action of the virtual character and the initial virtual scene. The initial character action can be the initial action of the virtual character, including the joint point action that drives each skeletal joint point of the virtual character to move. The initial virtual scene includes the virtual object obtained according to the above embodiment.

[0180] Specifically, the client inputs the action coding features and scene coding features obtained through the above embodiments into the action generation and decoding network, and decodes the action coding features and scene coding features through the action generation and decoding network to obtain the skeletal motion trajectory of each frame of the virtual character in the virtual scene, and then obtains the initial three-dimensional animation.

[0181] In step S420, the virtual character is controlled to perform an initial character action in the initial virtual scene, the distance between the virtual character and the virtual object when the virtual character performs the initial character action is obtained, and a loss result is generated according to the distance.

[0182] Among them, the distance between a virtual character and a virtual object may refer to the distance between the virtual character and each virtual object existing in the virtual scene during the execution of an action, for example, the closest distance, the farthest distance, etc. between the virtual character and each virtual object; it may also refer to the distance between a virtual character and a specified object, for example, the distance between a virtual character and a virtual object with which an interactive relationship exists.

[0183] The loss result (i.e., loss value) is used to reflect the difference between the actual distance and the ideal distance. The loss result can be obtained through a pre-deployed loss function.

[0184] Specifically, the specific manner in which the client controls the virtual character to perform the initial character action may be determined according to the type of the initial character action. For example, when the initial character action includes an action that generates displacement, the starting point and the end point may be determined from the initial virtual scene, and a path may be planned between the starting point and the end point so that the virtual character performs the action of moving from the starting point to the end point along the path. When the initial character action does not include an action that generates displacement, for example, the initial character action is to sit, the initial virtual character may perform the action of sitting down to the virtual object. After controlling the virtual character to perform the initial character action in the initial virtual scene, the client obtains the distance between the virtual character and the virtual object. Substitute the distance into the pre-configured loss function to obtain the loss result.

[0185] In step S430, the initial character action and the initial virtual scene are adjusted according to the loss result to obtain a three-dimensional animation.

[0186] Specifically, the client adjusts the initial character action and the layout of virtual objects in the initial virtual scene according to the loss result, and obtains the target character action and the target virtual scene, so that the distance obtained after the virtual character performs the target character action in the target virtual scene can be close to the ideal distance. Among them, the layout of the virtual object can be but not limited to the size, position, direction, etc. of the virtual object. The client generates the target character animation of the virtual character according to the virtual character and the target character action. Combined with the target character animation and the target virtual scene, the target character animation can be played in the target virtual scene, thereby generating a three-dimensional animation.

[0187] In one embodiment, generating a target animation may include but is not limited to requiring a motion route of a virtual character in a virtual scene, a target character action, and a target virtual scene. The above motion route may be selected through various path decision methods, such as A* algorithm, RRT-connect algorithm (random sampling algorithm), etc.

[0188] In this embodiment, after obtaining the initial three-dimensional animation, by simultaneously processing the character animation and the virtual scene, the character movements of the virtual character and the layout of the virtual scene are optimized according to the interaction between the virtual character in the character animation and the virtual objects in the virtual scene. This can make the interaction between the virtual character and the virtual objects more natural and reasonable, thereby improving the quality of the video animation.

[0189] In an exemplary embodiment, the action generation decoding network includes a weight learning network and an action prediction network; Figure 5 As shown, in step S410, the action coding features and the scene coding features are decoded by the action generation and decoding network to generate an initial three-dimensional animation, including:

[0190] In step S510, the action coding features and the scene coding features are mapped to a preset number of weight parameters through a weight learning network.

[0191] The weight parameters are used to distinguish the displacement action categories of the virtual characters. The weight learning network includes multiple linear layers. Specifically, the client inputs the action coding features and the scene coding features into the weight learning network, and the action coding features and the scene coding features are mapped to a preset number of weight parameters through the weight learning network.

[0192] In step S520, the weight parameters, action coding features and scene coding features are input into the action prediction network, and the skeleton motion trajectory of the virtual character in the next frame of the current frame is obtained through the action prediction network to generate an initial three-dimensional animation.

[0193] Among them, the action prediction network can adopt a generative adversarial network. The generative adversarial network includes: a generator, which is used to generate data based on input data; a discriminator, which is used to determine whether the generated data is true data. Specifically, the client inputs the weight parameters, action coding features and scene coding features output by the weight learning network into the action prediction network, and obtains the skeleton motion trajectory of the virtual character in the next frame of the current frame through the generator. The discriminator determines whether the obtained skeleton motion trajectory is true, and when it is true, the initial three-dimensional animation is obtained.

[0194] In this embodiment, the initial three-dimensional animation is directly predicted based on the action coding features and the scene coding features, which can improve the generation speed of the initial three-dimensional animation and thus speed up the production efficiency of the three-dimensional animation.

[0195] In an exemplary embodiment, the distance includes a penetration distance. In step S420, the distance between the virtual character and the virtual object when the virtual character performs the initial character action is obtained, and a loss result is generated according to the distance, including: when there is a collision between the first virtual object and the virtual character, determining the penetration distance between the virtual character and the first virtual object, and generating a penetration loss result according to the penetration distance; generating a loss result according to the penetration loss result.

[0196] The penetration distance is used to reflect the actual collision degree between the virtual character and the first virtual object. The penetration distance can be obtained based on collision box detection. For example, the first virtual object is surrounded by a cube, and the length of the virtual character passing through the cube is calculated as the penetration distance.

[0197] Specifically, after the virtual character has finished moving in the initial virtual scene, when there is a collision between the first virtual object and the virtual character, the client obtains the penetration distance when the virtual character collides with the first virtual object. In order to ensure that the virtual character does not collide with the virtual object during the movement, the penetration distance is constrained by a pre-configured collision loss function to obtain a penetration loss result.

[0198] In one embodiment, the client may use the mold penetration loss result as the loss result for final use. In another embodiment, the client may generate a loss result based on the mold penetration loss result, for example, predicting the loss result based on the mold penetration loss result; or calculating the mold penetration loss result according to a preset calculation formula to obtain the loss result.

[0199] In this embodiment, after the virtual character has completed its movement in the initial virtual scene, a penetration loss function is used to constrain the penetration distance between the virtual character and the virtual object, which can improve the collision between the virtual character in the animation and the objects in the virtual scene.

[0200] In an exemplary embodiment, the virtual object is divided into a plurality of blocks of preset sizes; determining the penetration distance between the virtual character and the first virtual object includes: obtaining the number of blocks through which the virtual character collides with the first virtual object, and determining the penetration distance according to the number.

[0201] Specifically, the client divides each virtual object existing in the virtual scene into a plurality of blocks according to a cube of a preset size. When obtaining a collision between a virtual character and a first virtual object, the number of blocks through which the virtual character and the first virtual object collide is obtained. In one embodiment, the number can be used as a penetration distance. In another embodiment, the penetration distance can be generated by predicting the number, for example, by calculating the number through a preset function to obtain the penetration distance.

[0202] In this embodiment, by dividing each virtual object in the virtual scene into multiple blocks according to fixed-size cubes, compared to using only one collision box to describe the virtual object, the collision situation between the virtual character and the virtual object, especially the virtual object with concave surfaces, can be accurately estimated, which helps to improve the collision situation between the virtual character and the virtual object.

[0203] In an exemplary embodiment, generating a penetration loss result according to the penetration distance includes: obtaining the ratio between the penetration distance corresponding to each first virtual object and the total number of segmented blocks, and obtaining the sum of the ratios corresponding to the first virtual object; normalizing the sum of the ratios to generate a penetration loss result. In this embodiment, by setting a penetration loss function applicable to the current iteration mode, the obtained penetration distance can be better constrained, so that the difference between the current collision situation and the ideal collision situation can be accurately characterized, which helps to improve the collision situation between the virtual character and the virtual object.

[0204] In an exemplary embodiment, the distance includes a contact distance. In step S420, the distance between the virtual character and the virtual object when the virtual character performs the initial character action is obtained, and a loss result is generated according to the distance, including: when there is a contact relationship between the second virtual object and the virtual character, determining the contact distance between the target joint point of the virtual character and the second virtual object, and generating a contact loss result according to the contact distance.

[0205] The second virtual object and the first virtual object may be the same object or different objects.

[0206] The contact relationship between the virtual character and the virtual object can be determined according to the character action performed by the virtual character. The contact relationship between the character action and the virtual object is pre-configured, and a second virtual object having a contact relationship with the initial character action is determined from the initial virtual scene according to the contact relationship. Exemplarily, the initial character action includes sitting, and the virtual object having a contact relationship with sitting includes a sofa. If there is a sofa in the initial virtual scene, the sofa is used as the second virtual object having a contact relationship with the virtual character.

[0207] The target joint point may be one or more fixed skeletal joint points in the virtual character; may also be one or more skeletal joint points randomly selected from the virtual character; or may also be one or more skeletal joint points pre-set to correspond to the second virtual object.

[0208] The contact distance is used to reflect the actual contact degree between the virtual character and the second virtual object. The contact distance can be determined based on the relative distance between the position of the target joint point of the virtual character and the position of the second virtual object, for example, the farthest or closest distance between the target joint point and the second virtual object is used as the contact distance.

[0209] Specifically, if there is a contact relationship between the second virtual object and the virtual character, then when the virtual character completes the movement in the virtual scene, the client obtains the relative distance between the target joint point and the second virtual object as the contact distance. In order to ensure that the virtual character can maintain a good interaction with the virtual object in contact during the movement, the contact distance is constrained by a pre-configured contact loss function to obtain a contact loss result. Then, a loss result is generated based on the contact loss result, for example, the loss result is obtained based on the contact loss result prediction; the contact loss result is calculated according to a preset calculation formula to obtain the loss result.

[0210] In this embodiment, after the virtual character has completed its movement in the initial virtual scene, a contact loss function is used to constrain the penetration distance between the virtual character and the virtual object, thereby ensuring that the virtual character and the virtual object in the final animation maintain a reasonable contact range.

[0211] In an exemplary embodiment, contact points corresponding to target joints are pre-marked on the second virtual object; determining the contact distance between the target joints of the virtual character and the second virtual object includes: obtaining the positions of the target joints in the initial virtual scene during the complete process of executing the initial character action; obtaining the average joint position of the target joint based on the positions in the initial virtual scene; and determining the contact distance based on the distance between the average joint position and the corresponding contact point.

[0212] Among them, the contact point is the point that the target joint point should touch when the virtual character interacts with the virtual object. The contact point corresponding to the target joint point is pre-marked on each virtual object that has a contact relationship with the virtual character. For example, the virtual object is a sofa and the virtual action is sitting. Then the target joint point may include the wrist joint. The virtual object with a contact relationship can be determined according to the actions supported by the system. For example, it can be a chair, table, sofa, keyboard, piano and other objects. When the number of target joint points is one, the contact point corresponding to the target joint point is marked on the virtual object; when the number of target joint points is multiple, the contact point corresponding to each target joint point is marked on the virtual object.

[0213] Specifically, during the entire process of the virtual object performing the initial character action, the client obtains the positions of the target joints of the virtual character in the initial virtual scene. The average of the positions of each target joint is calculated as the average joint position of each target joint. The distance between the average joint position of each target joint and the contact point corresponding to the target joint is obtained, and the distance is used as the contact distance.

[0214] In this embodiment, by marking the contact points corresponding to the target joint points on the virtual object, the contact situation between the virtual character and the virtual object is accurately judged based on the contact state between the two points, which helps to improve the improvement accuracy of the contact situation.

[0215] In an exemplary embodiment, generating a contact loss result according to the contact distance includes: obtaining the second norm of the contact distance of each second virtual object; obtaining the sum of the second norms of the second virtual objects, normalizing the sum of the second norms, and generating the contact loss result. In this embodiment, by setting a contact loss function suitable for the current iteration mode, the obtained contact distance can be better constrained, so that the difference between the current contact situation and the ideal contact situation can be accurately characterized, so that the virtual character and the virtual object can have a better contact state.

[0216] In an exemplary embodiment, the distance includes a penetration distance and a contact distance, in which case, Figure 6 As shown, step S420, obtaining the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generating a loss result according to the distance, can be achieved through the following steps.

[0217] In step S610, when a first virtual object collides with a virtual character, a penetration distance between the virtual character and the first virtual object is determined, and a penetration loss result is generated according to the penetration distance.

[0218] In step S620, when there is a contact relationship between the second virtual object and the virtual character, the contact distance between the target joint point of the virtual character and the second virtual object is determined, and a contact loss result is generated according to the contact distance.

[0219] In step S630, a loss result is generated according to the mold penetration loss result and the contact loss result.

[0220] Specifically, the generation method of the mold penetration loss result and the contact loss result can refer to the above embodiment, and will not be described in detail here. After obtaining the mold penetration loss result and the contact loss result, the mold penetration loss result and the contact loss result can be substituted into the preset calculation formula to obtain the loss result. The preset calculation formula depends on the actual iteration requirements, for example, an addition formula, a weighted sum formula, etc. can be used.

[0221] In this embodiment, the penetration loss function and the contact loss function are used simultaneously to constrain the distance between the virtual character and the virtual object, which can not only ensure that the virtual character in the final animation does not collide with the objects in the virtual scene, but also ensure that the virtual character in the final animation and the virtual object maintain a reasonable contact range, thereby achieving the best interaction state between the virtual character and the virtual scene.

[0222] In an exemplary embodiment, step S530 generates a loss result based on the penetration loss result and the contact loss result, including: obtaining a penetration coefficient corresponding to the penetration loss result, and a contact coefficient corresponding to the contact loss result; generating a loss result based on the penetration loss result and the penetration coefficient, and the contact loss result and the contact coefficient.

[0223] Among them, the penetration coefficient is mainly used to enhance the impact of the penetration loss result. The higher the penetration loss coefficient, the more collisions can be avoided. It can be a constant obtained by summarizing several historical iterative trainings; it can also be a coefficient determined according to the current iteration process, for example, a weight coefficient determined according to the current collision and contact conditions.

[0224] The contact coefficient is mainly used to improve the impact of contact loss results. The higher the contact loss coefficient, the more collisions can be ignored, ensuring that the contact position is closer. The contact coefficient can be a constant obtained by summarizing several historical iterative trainings; it can also be a coefficient determined according to the current iterative process, for example, a weight coefficient determined according to the current collision and contact conditions. The penetration coefficient and number of contacts corresponding to different character actions can be adjusted separately or at the same time.

[0225] In one embodiment, after obtaining the penetration coefficient and the contact coefficient, the client obtains a first product of the penetration coefficient and the penetration distance, and a second product of the contact coefficient and the contact distance. Then the sum of the first product and the second product is calculated as the loss result. In this embodiment, by setting a loss function suitable for the current iteration mode, the difference between the current overall situation and the ideal situation can be accurately characterized, so that the virtual character and the virtual object can have a better interaction state.

[0226] In an exemplary embodiment, step S430, adjusting the initial character action and the initial virtual scene according to the loss result to obtain a three-dimensional animation, includes: adjusting the initial character action and the initial virtual scene according to the loss result, controlling the virtual character to perform the adjusted initial character action in the adjusted initial virtual scene; obtaining the current distance between the virtual character and the virtual object when the virtual character performs the adjusted initial character action, and generating a current loss result according to the current distance; iterating the above steps until a preset stop condition is met to obtain a three-dimensional animation.

[0227] Specifically, after the client obtains the loss result corresponding to the initial character action and the initial virtual scene through step S420, the initial character action and the initial virtual scene are iteratively trained for multiple times according to the loss result. In each iterative training process, the client controls the virtual character to perform the current character action in the current virtual scene, obtains the current distance between the virtual character and the virtual object when performing the current character action, and generates the current loss result according to the current distance. The current character action and the current virtual scene are adjusted according to the current loss result until the current iteration meets the preset stop condition. The client can select the character action and virtual scene that meet the preset conditions from the obtained character actions and virtual scenes as the target character action and target virtual scene, and then generate a three-dimensional animation according to the target character action and target virtual scene. Among them, the preset stop condition can be, but is not limited to, any one of the loss result reaching a preset threshold, the number of iterations reaching a preset number, etc.

[0228] In one embodiment, the client performs iterative updates according to the above process until the number of iterations reaches a preset number. The minimum loss result is obtained from the loss results generated for the preset number of times. The character action and virtual scene corresponding to the minimum loss result are used as the target character action and target virtual scene, and then a three-dimensional animation is generated according to the target character action and target virtual scene.

[0229] In this embodiment, by performing multiple iterations of updating the character actions and the virtual scene, an approximate global optimal solution of the character actions and the virtual scene can be searched, thereby generating an animation with the best quality.

[0230] In an exemplary embodiment, the initial character action and the initial virtual scene can be adjusted based on the simulated annealing algorithm. The simulated annealing algorithm is derived from the solid annealing principle and is a probability-based algorithm. The simulated annealing algorithm is an optimization algorithm of a serial structure that effectively avoids falling into a local minimum and eventually tends to the global optimum by giving the search process a time-varying probability jump that eventually tends to zero. Figure 7 As shown, step S430, adjusting the initial character action and the initial virtual scene according to the loss result to obtain a three-dimensional animation, can be achieved through the following steps.

[0231] In step S710, the initial character action is fine-tuned to obtain a fine-tuned initial character action.

[0232] In step S720, the initial virtual scene is fine-tuned to obtain a fine-tuned initial virtual scene.

[0233] In step S730, a reference loss result is generated according to the fine-tuned initial character action and the fine-tuned initial virtual scene.

[0234] Specifically, the client determines whether to accept the current new solution based on the Metropolis criterion. The Metropolis criterion determines whether to accept a new solution under a certain probability. This probability is to obtain a random number of 0-1, and if it is greater than 0.5, the new solution is accepted. After obtaining the loss result, the client randomly generates a first fine-tuning parameter, adds the first fine-tuning parameter to the initial character action, and obtains the fine-tuned initial character action. Randomly generate a second fine-tuning parameter, add the second fine-tuning parameter to the initial virtual scene, and obtain the fine-tuned initial virtual scene. The client generates a reference loss result based on the fine-tuned initial character action and the fine-tuned initial virtual scene. The specific generation process of the reference loss result can refer to the above embodiment and will not be elaborated here.

[0235] In step S740, when the relationship between the loss result and the reference loss result meets a preset condition, the initial character action and the initial virtual scene are adjusted according to the loss result to obtain a three-dimensional animation.

[0236] In step S750, when the relationship between the loss result and the reference loss result does not satisfy a preset condition, the initial character action is used as the target character action, and the initial three-dimensional animation is used as the three-dimensional animation.

[0237] Specifically, if the client determines that the reference loss result is greater than the loss result, it is determined to accept the current new solution, and the initial character action and the initial virtual scene are adjusted according to the reference loss result to obtain the target character action and the target virtual scene, and then a three-dimensional animation is generated according to the target character action and the target virtual scene. If the client determines that the reference loss result is less than or equal to the loss result, it is determined not to accept the current new solution, and the initial character action and the initial virtual scene may not be adjusted, and the initial three-dimensional animation may be used as the three-dimensional animation.

[0238] In one embodiment, when the initial character action and the initial virtual scene are updated for multiple iterations, the current character action and the current virtual scene of the current iteration may be adjusted based on the simulated annealing algorithm for each iteration.

[0239] In this embodiment, the character actions and virtual scenes are updated based on the simulated annealing algorithm, which can not only search for the approximate global optimal solution of the character actions and virtual scenes to generate the best quality animation; but also speed up the update speed and improve the efficiency of animation generation.

[0240] In one embodiment, referring to Figure 8 The flowchart shown and Fig. 9 The schematic diagram shown provides an animation generation method, such as Figure 8 As shown, the following steps are included:

[0241] In step S802, input information is obtained, and input features are extracted from the input information.

[0242] In step S804, the motion coding feature (including displacement motion information and non-displacement motion information) is obtained according to the input feature. The specific generation method of the displacement motion information and the non-displacement motion information can refer to the above embodiment, which will not be elaborated here.

[0243] In step S806, a virtual scene is generated according to the input features, and the virtual scene is encoded to generate scene encoding features.

[0244] Specifically, the client classifies and identifies the input features to obtain the virtual object. A graph network structure of the virtual object is constructed based on the virtual object and the virtual object relationship, and the virtual object relationship is used to characterize the association relationship between different virtual objects. A virtual scene matching the graph network structure is determined from existing virtual scenes. The virtual object in the virtual scene is segmented to obtain a plurality of segmented blocks of preset sizes. The segmented blocks corresponding to the virtual object in the virtual scene are input into the encoding network to obtain the scene encoding features.

[0245] In step S808, the scene coding features and the action coding features are input into the action generation and decoding network to obtain an initial three-dimensional animation. The initial three-dimensional animation includes the initial character action of the virtual character and the initial virtual scene.

[0246] In step S810, the initial path from the virtual character to the virtual object is searched in the virtual scene by using the A-star algorithm, and the initial loss result is obtained based on the initial path, the initial character action, and the initial virtual scene.

[0247] Among them, the following is the specific implementation process of the simulated annealing algorithm:

[0248] Input:Virtual Scene Layout$E$,Blended Motion$M_o$ / / Input: Scene layout E, blended motion Mo

[0249] Output:Optimized Scene$\hat{E}$,Optimized Motion$\hat{M}$ / / Output: adjusted virtual scene, adjusted character motion, optimized motion

[0250] $R_s\gets SearchRoute(E,M_o)$ / / Search path

[0251] $R_g,M_g,E_g\gets NSMGenerate(R_s,E,M_o)$ / / Generate preliminary actions

[0252] $S_{cur}\gets S_{old}\gets GetRunStatus(R_g,M_g,E_g)$ / / Get the current status, including scene layout, virtual character position, motion trajectory, collision information, etc.

[0253] $OldError\gets CalculateError()$ / / Calculate the loss result

[0254] \For{$each\step$}{ / / loop iteration

[0255] $S_{cur}\gets AddDisturbance(S_{old})$ / / Add random values ​​to get a new virtual scene

[0256] $R_g,M_g,E_g\gets NSMGenerate(S_{cur})$ / / Generate new animation based on the new virtual scene

[0257] $S_{cur}\gets GetRunStatus(R_g,M_g,E_g)$ / / Get the current status

[0258] $error\gets CalculateError()$ / / Calculate reference loss value

[0259] $error\gets AcceptanceProbability(error)$ / / Determine whether to accept the new solution based on the criteria

[0260] $error,OldError,S_{old}\gets UpdateStatus(S_{cur},S_{old},error,OldError)$ / / Update loss value, status

[0261] }

[0262] Specifically, the virtual character is controlled to perform the initial character action according to the initial path in the initial virtual scene. After the execution is completed, when the first virtual object collides with the virtual character, the number of segments that the virtual character passes through when colliding with the first virtual object is obtained as the penetration distance. When the second virtual object is in contact with the virtual character, the average position of the joints of each target joint point during the entire execution process is determined; the distance between the average position of the joints of each target joint point and the corresponding contact point is obtained as the contact distance. Then, the initial loss value is generated according to the penetration distance and the contact distance.

[0263] In step S812, it is determined whether to accept the current new solution according to the Metropolis criterion. If the new solution is accepted, the current path, the current character action, and the current virtual scene are adjusted according to the initial loss result. If the new solution is not accepted, the current path, the current character action, and the current virtual scene are not adjusted and continue to be used.

[0264] In step S814, when the current number of iterations reaches a preset number, a three-dimensional animation is generated.

[0265] Specifically, the client determines whether the current number of iterations reaches a preset number. If the preset number is reached, the character action and virtual scene corresponding to the minimum loss value are obtained as the target character action and target virtual scene. The target character animation is generated according to the target character action of the virtual character, and the target three-dimensional animation is generated by combining the target character animation and the target virtual scene. If the preset number has not been reached, steps S810 to S814 are repeated.

[0266] It should be understood that, although the various steps in the above flowchart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0267] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can refer to each other, and each embodiment focuses on the differences from other embodiments. For related points, please refer to the description of other method embodiments.

[0268] Fig.10FIG. 1 is a block diagram of an animation generating device 1000 according to an exemplary embodiment. Fig.10 The device includes: an information acquisition module 1002, an action feature generation module 1004, a scene feature generation module 1006, and an animation generation module 1008.

[0269] The information acquisition module 1002 is configured to acquire input information and extract input features from the input information; the action feature generation module 1004 is configured to identify the input features, determine the action information of the virtual character, and determine the action coding features corresponding to the action information; the scene feature generation module 1006 is configured to identify the input features, determine the virtual objects, generate the virtual scene according to the virtual objects, and determine the scene coding features corresponding to the virtual scene; the animation generation module 1008 is configured to generate three-dimensional animation according to the action coding features and the scene coding features.

[0270] In an exemplary embodiment, the action information includes displacement action information and non-displacement action information; the action feature generation module 1004 includes: a first classification unit, configured to perform classification and identification of the input feature, obtain a displacement action category, and obtain displacement action information corresponding to the displacement action category, wherein the displacement action information is generated using an action generation network; a first matching unit, configured to perform determination of an action identifier that matches the input feature from a non-displacement action database, and generate non-displacement action information based on the matched action identifier; a feature generation unit, configured to generate the action coding feature based on the displacement action information and the non-displacement action information.

[0271] In an exemplary embodiment, the first matching unit is configured to execute a search for a motion capture file corresponding to the action identifier from a non-displacement action database; perform full-connection processing on the motion capture file, and perform feature extraction on the fully-connected motion capture file to obtain the non-displacement action information.

[0272] In an exemplary embodiment, the scene feature generation module 1006 includes: a second classification unit, configured to perform classification and recognition of input features to obtain virtual objects; a structure construction unit, configured to execute a graph network structure of virtual objects based on virtual objects and virtual object relationships, and the virtual object relationships are used to characterize the association relationship between different virtual objects; a second matching unit, configured to determine a virtual scene that matches the graph network structure from an existing virtual scene.

[0273] In an exemplary embodiment, the scene feature generation module 1006 includes: a segmentation unit, configured to execute segmentation of virtual objects in the virtual scene to obtain multiple segments of preset sizes; a scene feature encoding unit, configured to execute input of the segments corresponding to the virtual objects in the virtual scene into the encoding network to obtain scene coding features.

[0274] In an exemplary embodiment, the animation generation module 1008 includes: a decoding unit, configured to perform decoding processing on action coding features and scene coding features through a pre-trained action generation decoding network to generate an initial three-dimensional animation, wherein the initial three-dimensional animation includes an initial character action of the virtual character and an initial virtual scene, and the initial virtual scene includes a virtual object; an action execution unit, configured to control the virtual character to perform the initial character action in the initial virtual scene, obtain the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generate a loss result according to the distance; an adjustment unit, configured to adjust the initial character action and the initial virtual scene according to the loss result to obtain a three-dimensional animation.

[0275] In an exemplary embodiment, the action generation decoding network includes a weight learning network and an action prediction network; the decoding unit includes: a weight learning subunit, configured to execute mapping of action coding features and scene coding features to a preset number of weight parameters through the weight learning network, and the weight parameters are used to distinguish the action categories of the virtual character; an action prediction subunit, configured to execute input of the weight parameters, action coding features and scene coding features into the action prediction network, and obtain the skeletal motion trajectory of the virtual character in the next frame of the current frame through the action prediction network to generate an initial three-dimensional animation.

[0276] In an exemplary embodiment, an action execution unit includes: a first action execution sub-unit, configured to determine a penetration distance between a virtual character and the first virtual object when a collision occurs between the first virtual object and the virtual character, and generate a penetration loss result according to the penetration distance; a loss result determination sub-unit, configured to generate a loss result according to the penetration loss result.

[0277] In an exemplary embodiment, the virtual object is divided into a plurality of blocks of preset sizes; the first action execution subunit is configured to execute obtaining the number of blocks through which the virtual character collides with the first virtual object, and determine the penetration distance according to the number.

[0278] In an exemplary embodiment, the first action execution subunit is configured to execute the acquisition of the ratio between the penetration distance corresponding to each first virtual object and the total number of segmented blocks, and to obtain the sum of the ratios corresponding to the first virtual objects; the sum of the ratios is normalized to generate a penetration loss result.

[0279] In an exemplary embodiment, the action execution unit also includes: a second action execution sub-unit, which is configured to determine the contact distance between the target joint point of the virtual character and the second virtual object when there is a contact relationship between the second virtual object and the virtual character, and generate a contact loss result based on the contact distance; in this embodiment, the loss result determination sub-unit is configured to generate a loss result based on the penetration loss result and the contact loss result.

[0280] In an exemplary embodiment, contact points corresponding to target joints are pre-marked on the second virtual object, and the contact points are the points that the target joints should touch when the virtual character interacts with the second virtual object; the second action execution subunit is configured to execute the acquisition of the positions of the target joints that appeared in the initial virtual scene during the complete process of executing the initial character action; based on the positions that appeared in the initial virtual scene, the average joint position of the target joints is obtained; and the contact distance is determined based on the distance between the average joint position and the contact point.

[0281] In an exemplary embodiment, the second action execution subunit is configured to execute obtaining the second norm of the contact distance of each second virtual object; obtaining the sum of the second norms of the second virtual objects, normalizing the sum of the second norms, and generating a contact loss result.

[0282] In an exemplary embodiment, the loss result determination subunit is configured to execute acquisition of a penetration coefficient corresponding to the penetration loss result and a contact coefficient corresponding to the contact loss result; and generate a loss result based on the penetration loss result and the penetration coefficient, and the contact loss result and the contact coefficient.

[0283] In an exemplary embodiment, the adjustment unit includes: a first adjustment sub-unit, configured to adjust the initial character action and the initial virtual scene according to the loss result, and control the virtual character to perform the adjusted initial character action in the adjusted initial virtual scene; a loss result generation sub-unit, configured to obtain the current distance between the virtual character and the virtual object when performing the adjusted initial character action, and generate the current loss result according to the current distance; an iterative update unit, configured to iterate the above steps until a preset stop condition is met to obtain a three-dimensional animation.

[0284] In an exemplary embodiment, the adjustment unit includes: a first fine-tuning subunit, configured to perform fine-tuning on the initial character action to obtain the fine-tuned initial character action; a second fine-tuning subunit, configured to perform fine-tuning on the initial virtual scene to obtain the fine-tuned initial virtual scene; a reference result generating subunit, configured to generate a reference loss result based on the fine-tuned initial character action and the fine-tuned initial virtual scene; a second adjustment subunit, configured to perform, when the relationship between the loss result and the reference loss result meets a preset condition, adjusting the initial character action and the initial virtual scene respectively according to the loss result to obtain a three-dimensional animation; when the relationship between the loss result and the reference loss result does not meet the preset condition, taking the initial character action as the target character action and the initial three-dimensional animation as the three-dimensional animation.

[0285] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0286] Fig.10 1 is a block diagram of an electronic device Z00 for generating animation according to an exemplary embodiment. For example, the electronic device Z00 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0287] Reference Fig.10 The electronic device Z00 may include one or more of the following components: a processing component Z02, a memory Z04, a power component Z06, a multimedia component Z08, an audio component Z10, an input / output (I / O) interface Z12, a sensor component Z14, and a communication component Z16.

[0288] The processing component Z02 generally controls the overall operation of the electronic device Z00, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component Z02 may include one or more processors Z20 to execute instructions to complete all or part of the steps of the above-described method. In addition, the processing component Z02 may include one or more modules to facilitate interaction between the processing component Z02 and other components. For example, the processing component Z02 may include a multimedia module to facilitate interaction between the multimedia component Z08 and the processing component Z02.

[0289] The memory Z04 is configured to store various types of data to support operations on the electronic device Z00. Examples of such data include instructions for any application or method operating on the electronic device Z00, contact data, phone book data, messages, pictures, videos, etc. The memory Z04 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, or graphene memory.

[0290] The power supply assembly Z06 provides power to various components of the electronic device Z00. The power supply assembly Z06 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device Z00.

[0291] The multimedia component Z08 includes a screen that provides an output interface between the electronic device Z00 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component Z08 includes a front camera and / or a rear camera. When the electronic device Z00 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0292] The audio component Z10 is configured to output and / or input audio signals. For example, the audio component Z10 includes a microphone (MIC), and when the electronic device Z00 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory Z04 or sent via the communication component Z16. In some embodiments, the audio component Z10 also includes a speaker for outputting an audio signal.

[0293] I / O interface Z12 provides an interface between processing component Z02 and peripheral interface modules, which may be keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0294] The sensor assembly Z14 includes one or more sensors for providing various aspects of status assessment for the electronic device Z00. For example, the sensor assembly Z14 can detect the open / closed state of the electronic device Z00, the relative positioning of the components, such as the display and keypad of the electronic device Z00, and the sensor assembly Z14 can also detect the position change of the electronic device Z00 or the components of the electronic device Z00, the presence or absence of contact between the user and the electronic device Z00, the orientation or acceleration / deceleration of the device Z00 and the temperature change of the electronic device Z00. The sensor assembly Z14 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly Z14 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly Z14 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0295] The communication component Z16 is configured to facilitate wired or wireless communication between the electronic device Z00 and other devices. The electronic device Z00 can access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component Z16 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component Z16 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0296] In an exemplary embodiment, the electronic device Z00 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0297] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory Z04 including instructions, and the above instructions can be executed by a processor Z20 of the electronic device Z00 to complete the above method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0298] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions. The instructions can be executed by the processor Z20 of the electronic device Z00 to implement the above method.

[0299] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. may also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments, which will not be described one by one here.

[0300] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.

[0301] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An animation generation method, characterized in that: include: Acquire input information, and extract input features from the input information; Classifying and identifying the input features to obtain displacement action categories, and acquiring displacement action information corresponding to the displacement action categories, wherein the displacement action information is generated by using an action generation network; Determine an action identifier that matches the input feature from a non-displacement action database, and generate non-displacement action information according to the matched action identifier; generating an action coding feature according to the displacement action information and the non-displacement action information; Identify the input features, determine the virtual objects, generate a virtual scene according to the virtual objects, and determine the scene coding features corresponding to the virtual scene; Decoding the action coding features and the scene coding features through an action generation and decoding network to generate an initial three-dimensional animation, wherein the initial three-dimensional animation includes an initial character action of a virtual character and an initial virtual scene, wherein the initial virtual scene includes a virtual object; Controlling the virtual character to perform the initial character action in the initial virtual scene, obtaining the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generating a loss result according to the distance; The initial character action and the initial virtual scene are adjusted according to the loss result to obtain the three-dimensional animation.

2. The animation generation method according to claim 1, characterized in that: The generating non-displacement action information according to the matched action identifier includes: Searching for a motion capture file corresponding to the motion identifier from a non-displacement motion database; The motion capture file is fully connected and feature extracted from the motion capture file after the full connection processing to obtain the non-displacement motion information.

3. The animation generation method according to claim 1, characterized in that: The step of identifying the input features, determining the virtual objects, and generating the virtual scenes according to the virtual objects includes: Classifying and identifying the input features to obtain the virtual object; A graph network structure of the virtual objects is constructed according to the virtual objects and virtual object relationships, wherein the virtual object relationships are used to represent association relationships between different virtual objects; A virtual scene matching the graph network structure is determined from existing virtual scenes.

4. The animation generation method according to claim 1, characterized in that: The determining of the scene coding feature corresponding to the virtual scene comprises: Segmenting the virtual object in the virtual scene to obtain a plurality of segmented blocks of preset sizes; The segmented blocks corresponding to the virtual objects in the virtual scene are input into a coding network to obtain the scene coding features.

5. The animation generation method according to claim 4, characterized in that: The action generation decoding network includes a weight learning network and an action prediction network; The step of decoding the action coding features and the scene coding features through an action generation decoding network to generate an initial three-dimensional animation includes: Mapping the action coding features and the scene coding features to a preset number of weight parameters through the weight learning network, wherein the weight parameters are used to distinguish the action categories of the virtual character; The weight parameters, the action coding features and the scene coding features are input into the action prediction network, and the skeletal motion trajectory of the virtual character in the next frame of the current frame is obtained through the action prediction network to generate the initial three-dimensional animation.

6. The animation generation method according to claim 4, characterized in that: The obtaining of the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generating a loss result according to the distance, comprises: When a first virtual object collides with the virtual character, determining a penetration distance between the virtual character and the first virtual object, and generating a penetration loss result according to the penetration distance; The loss result is generated according to the mode penetration loss result.

7. The animation generation method according to claim 6, characterized in that: The virtual object is divided into a plurality of division blocks of preset sizes; The determining the penetration distance between the virtual character and the first virtual object includes: The number of segments that the virtual character passes through when colliding with the first virtual object is obtained, and the penetration distance is determined according to the number.

8. The animation generation method according to claim 6, characterized in that: The generating the mold penetration loss result according to the mold penetration distance includes: The ratio between the penetration distance corresponding to each of the first virtual objects and the total number of segmented blocks is obtained, and the sum of the ratios corresponding to the first virtual objects is obtained; the sum of the ratios is normalized to generate the penetration loss result.

9. The animation generation method according to claim 6, characterized in that: The obtaining of the distance between the virtual character and the virtual object when the virtual character performs the initial character action, and generating a loss result according to the distance, further comprises: When there is a contact relationship between the second virtual object and the virtual character, determining a contact distance between a target joint point of the virtual character and the second virtual object, and generating a contact loss result according to the contact distance; The step of generating the loss result according to the mode penetration loss result includes: The loss result is generated according to the mode penetration loss result and the contact loss result.

10. The animation generation method according to claim 9, characterized in that: The second virtual object is marked with a contact point corresponding to the target joint point, and the contact point is a point that the target joint point should contact when the virtual character interacts with the second virtual object; The determining the contact distance between the target joint point of the virtual character and the second virtual object includes: Acquire the position where the target joint point appears in the initial virtual scene during the complete process of executing the initial character action; Obtaining an average joint point position of the target joint point according to the positions that appeared in the initial virtual scene; The contact distance is determined according to the distance between the average position of the joint points and the contact point.

11. The animation generation method according to claim 9, characterized in that: Generating a contact loss result according to the contact distance includes: Obtaining a second norm of a contact distance of each of the second virtual objects; The sum of the second norms of the second virtual object is obtained, and the sum of the second norms is normalized to generate the contact loss result.

12. The animation generation method according to claim 9, characterized in that: The generating the loss result according to the mode penetration loss result and the contact loss result comprises: Obtaining a penetration coefficient corresponding to the penetration loss result and a contact coefficient corresponding to the contact loss result; The loss result is generated according to the penetration loss result and the penetration coefficient, and the contact loss result and the contact coefficient.

13. The animation generation method according to any one of claims 4 to 12, characterized in that: The step of adjusting the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional animation includes: Adjusting the initial character action and the initial virtual scene according to the loss result, and controlling the virtual character to perform the adjusted initial character action in the adjusted initial virtual scene; Acquire a current distance between the virtual character and the virtual object when the virtual character performs the adjusted initial character action, and generate a current loss result according to the current distance; The above steps are iterated until a preset stop condition is met to obtain the three-dimensional animation.

14. The animation generation method according to any one of claims 4 to 12, characterized in that: The step of adjusting the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional animation includes: Fine-tuning the initial character action to obtain the fine-tuned initial character action; Fine-tuning the initial virtual scene to obtain the fine-tuned initial virtual scene; Generate a reference loss result according to the fine-tuned initial character action and the fine-tuned initial virtual scene; When the relationship between the loss result and the reference loss result satisfies a preset condition, adjusting the initial character action and the initial virtual scene respectively according to the loss result to obtain the three-dimensional animation; When the relationship between the loss result and the reference loss result does not satisfy the preset condition, the initial three-dimensional animation is used as the three-dimensional animation.

15. An animation generating device, characterized in that: include: An information acquisition module is configured to acquire input information and extract input features from the input information; The action feature generation module includes a first classification unit configured to perform classification and recognition on the input feature to obtain a displacement action category and obtain displacement action information corresponding to the displacement action category, wherein the displacement action information is generated by using an action generation network; A first matching unit is configured to determine an action identifier matching the input feature from a non-displacement action database, and generate non-displacement action information according to the matched action identifier; A feature generation unit, configured to generate an action coding feature according to the displacement action information and the non-displacement action information; A scene feature generation module is configured to perform recognition of the input features, determine virtual objects, generate a virtual scene according to the virtual objects, and determine a scene coding feature corresponding to the virtual scene; The animation generation module includes a decoding unit configured to perform decoding processing on the action coding features and the scene coding features through a pre-trained action generation decoding network to generate an initial three-dimensional animation, wherein the initial three-dimensional animation includes an initial character action of a virtual character and an initial virtual scene, wherein the initial virtual scene includes a virtual object; an action execution unit, configured to execute and control the virtual character to execute the initial character action in the initial virtual scene, obtain the distance between the virtual character and the virtual object when the virtual character executes the initial character action, and generate a loss result according to the distance; The adjustment unit is configured to adjust the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional animation.

16. The animation generating device according to claim 15, characterized in that: The first matching unit is configured to search for a motion capture file corresponding to the motion identifier from a non-displacement motion database; perform full-connection processing on the motion capture file, and perform feature extraction on the fully-connected motion capture file to obtain the non-displacement motion information.

17. The animation generating device according to claim 15, characterized in that: The scene feature generation module comprises: A second classification unit is configured to perform classification and recognition on the input feature to obtain the virtual object; a structure construction unit configured to execute construction according to the virtual objects and virtual object relationships to obtain a graph network structure of the virtual objects, wherein the virtual object relationships are used to represent association relationships between different virtual objects; The second matching unit is configured to determine a virtual scene that matches the graph network structure from existing virtual scenes.

18. The animation generating device according to claim 15, characterized in that: The scene feature generation module comprises: A segmentation unit, configured to segment the virtual object in the virtual scene to obtain a plurality of segmented blocks of preset sizes; The scene feature encoding unit is configured to input the segmented blocks corresponding to the virtual objects in the virtual scene into an encoding network to obtain the scene encoding features.

19. The animation generating device according to claim 15, characterized in that: The action generation decoding network includes a weight learning network and an action prediction network; the decoding unit includes: A weight learning subunit, configured to execute mapping of the action coding features and the scene coding features to a preset number of weight parameters through the weight learning network, wherein the weight parameters are used to distinguish the action categories of the virtual character; The action prediction subunit is configured to input the weight parameters, the action coding features and the scene coding features into the action prediction network, obtain the skeleton motion trajectory of the virtual character in the next frame of the current frame through the action prediction network, and generate the initial three-dimensional animation.

20. The animation generating device according to claim 15, characterized in that: The action execution unit includes: A first action execution subunit is configured to determine a penetration distance between the virtual character and the first virtual object when a first virtual object collides with the virtual character, and generate a penetration loss result according to the penetration distance; The loss result determination subunit is configured to generate the loss result according to the penetration loss result.

21. The animation generating device according to claim 20, characterized in that: The virtual object is divided into a plurality of blocks of preset sizes; the first action execution subunit is configured to execute obtaining the number of blocks through which the virtual character collides with the first virtual object, and determine the penetration distance according to the number.

22. The animation generating device according to claim 20, characterized in that: The first action execution subunit is configured to execute the step of obtaining a ratio between a penetration distance corresponding to each of the first virtual objects and a total number of segmented blocks, and obtaining a sum of the ratios corresponding to the first virtual objects; The sum of the ratios is normalized to generate the mode penetration loss result.

23. The animation generating device according to claim 20, characterized in that: The action execution unit further includes: A second action execution subunit is configured to determine a contact distance between a target joint point of the virtual character and the second virtual object when there is a contact relationship between the second virtual object and the virtual character, and generate a contact loss result according to the contact distance; The loss result determination subunit is configured to generate the loss result according to the penetration loss result and the contact loss result.

24. The animation generating device according to claim 23, characterized in that: The second virtual object is pre-marked with a contact point corresponding to the target joint point, and the contact point is a point that the target joint point should contact when the virtual character interacts with the second virtual object; The second action execution subunit is configured to execute acquisition of the position where the target joint point appeared in the initial virtual scene during the complete process of executing the initial character action; The average joint point position of the target joint point is obtained according to the position that appeared in the initial virtual scene; and the contact distance is determined according to the distance between the average joint point position and the contact point.

25. The animation generating device according to claim 23, characterized in that: The second action execution subunit is configured to execute obtaining the second norm of the contact distance of each second virtual object; obtaining the sum of the second norms of the second virtual objects, normalizing the sum of the second norms, and generating the contact loss result.

26. The animation generating device according to claim 23, characterized in that: The loss result determination subunit is configured to execute acquisition of a penetration coefficient corresponding to the penetration loss result and a contact coefficient corresponding to the contact loss result; and to generate the loss result according to the penetration loss result and the penetration coefficient, and the contact loss result and the contact coefficient.

27. The animation generating device according to any one of claims 15 to 26, characterized in that: The adjustment unit comprises: A first adjustment subunit is configured to adjust the initial character action and the initial virtual scene according to the loss result, and control the virtual character to perform the adjusted initial character action in the adjusted initial virtual scene; A loss result generating subunit is configured to obtain a current distance between the virtual character and the virtual object when the virtual character performs the adjusted initial character action, and generate a current loss result according to the current distance; The iterative updating unit is configured to iteratively execute the above steps until a preset stop condition is met to obtain the three-dimensional animation.

28. The animation generating device according to any one of claims 15 to 26, characterized in that: The adjustment unit comprises: A first fine-tuning subunit is configured to perform fine-tuning on the initial character action to obtain the fine-tuned initial character action; A second fine-tuning subunit is configured to perform fine-tuning on the initial virtual scene to obtain the fine-tuned initial virtual scene; A reference result generating subunit is configured to generate a reference loss result according to the fine-tuned initial character action and the fine-tuned initial virtual scene; The second adjustment subunit is configured to execute, when the relationship between the loss result and the reference loss result meets a preset condition, adjusting the initial character action and the initial virtual scene according to the loss result to obtain the three-dimensional animation; when the relationship between the loss result and the reference loss result does not meet the preset condition, using the initial three-dimensional animation as the three-dimensional animation.

29. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the animation generation method according to any one of claims 1 to 14.

30. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the animation generating method according to any one of claims 1 to 14.

31. A computer program product, comprising instructions, characterized in that: When the instruction is executed by a processor of an electronic device, the electronic device is enabled to execute the animation generating method as described in any one of claims 1 to 14.

Citation Information

Patent Citations

  • Virtual image live broadcast method and device, electronic device and storage medium

    CN110689570A

  • Real-time interaction method and device based on virtual reality, equipment and storage medium

    CN112379771A