A data processing method and related device
Patent Information
- Application Number
- CN202110797678.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-14
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2041-07-14
AI Technical Summary
[0004]然而,在上述角色动画自动生成技术中,需要用户输入控制信号,例如实时的速度、方向等
[0074] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: n sets of action states are obtained through the first information and the first action state; and the limb movements of the target object are processed based on the first action state and the n sets of action states to obtain a first image and n images, thereby generating a target video based on the first image and the n images. Since the higher-level semantics of the first action type and the first action attribute in the first script are more efficient, intuitive, and understandable, compared to the prior art which requires user input of lower-level control signals, it can reduce the user's workload and professional requirements, thereby improving the efficiency of subsequent target video generation.
Smart Images

Figure CN115617429B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of animation technology, and in particular to a data processing method and related equipment. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. Research in the field of AI includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, fundamental AI theories, and automatic animation generation.
[0003] Currently, automatic character animation generation technology is commonly used in games, animation video creation, and interactive applications, such as controlling characters to perform specific actions in games using a mouse and keyboard.
[0004] However, the aforementioned automatic character animation generation technology requires user input of control signals, such as real-time speed and direction. Therefore, how to generate character animations while reducing the user's workload is a pressing technical problem that needs to be solved. Summary of the Invention
[0005] This application provides a data processing method and related equipment. Compared to existing technologies that require user input of low-level control signals, this method reduces user workload and professional requirements, thereby improving the efficiency of subsequent target video generation.
[0006] The first aspect of this application provides a data processing method that can be applied to the automatic and rapid production of children's educational animations, short video animations, promotional animations, variety show animations, film and television pre-show animations, etc., or to automatic motion control in games, or to automatically generate character actions in interactive applications. This method can be executed by a data processing device or by a component of the data processing device (e.g., a processor, chip, or chip system). The method includes: acquiring first information, the first information including a first action type and a first action attribute, the first action type describing a first limb action, and the first action attribute describing the process of the first limb action; acquiring a first action state; obtaining n sets of action states based on the first action state and the first information, where n is a positive integer; acquiring a target object; processing the limb actions of the target object based on the first action state to obtain a first image; processing the limb actions of the target object based on the n sets of action states to obtain n images; and generating a target video based on the first image and the n images, the target video being related to the first action type and the first action attribute. For example, the first action type is "walking," and the first action attribute is the parameters corresponding to "walking": "target (or displacement)" and "travel path." In other words, the first action type and the first action attribute belong to high-level semantics, which do not require users to have too many technical requirements, and are applicable to a wide range of users.
[0007] In this embodiment, n sets of action states are obtained through first information and first action states. Based on the first action states and the n sets of action states, the limb movements of the target object are processed to obtain a first image and n images. A target video is then generated based on the first image and the n images. Since the higher-level semantics of the first action type and first action attribute in the first script are more efficient, intuitive, and understandable, compared to the prior art which requires user input of lower-level control signals, it can reduce the user's workload and professional requirements, thereby improving the efficiency of subsequent target video generation.
[0008] Optionally, in one possible implementation of the first aspect, the above step of obtaining the first information includes: obtaining a first script input by the user, the first script including a first action type and a first action attribute; or, displaying a first user interface, the first user interface including action type options and / or action attribute options; responding to a first operation by the user on the first user interface, determining the first action type from the action type options, and / or selecting the first action attribute from the action attribute options. It is understood that in practical applications, other methods may also be used to obtain the first information, and specific methods are not limited here.
[0009] In this possible implementation, the user interface can be used to interact with the user, allowing the user to obtain general character animations (i.e., videos of the target object) by inputting scripts with high-level semantics, thus improving the video production experience for users with limited technical skills.
[0010] Optionally, in one possible implementation of the first aspect, the above step of obtaining the first action state includes: obtaining a second action state, the second action state representing a historical action state prior to the first action state; inputting the second action state and first information into a trained first network to obtain the first action state; or, displaying a second user interface to the user, the second user interface including an action state area; responding to a second operation by the user on the second user interface, determining the first action state from the action state area; or, obtaining a pre-set first action state. It is understood that in practical applications, other methods may also be used to obtain the first action state, and specific methods are not limited here.
[0011] In this possible implementation, the initial action state can be obtained through user interaction with the user interface, thereby enhancing the user's experience in creating videos.
[0012] Optionally, in one possible implementation of the first aspect, the above steps: obtaining n sets of action states based on the first action state and the first information include: inputting the first action state and the first information into a trained first network to obtain n sets of action states, wherein the first network is used to obtain a fourth action state based on the third action state, the first action type and the first action attribute, and the change from the third action state to the fourth action state is related to the first action type.
[0013] In this possible implementation, the fourth action state is obtained by combining the third motion state with the first network. Since the first network is trained, the generated n sets of action states are more consistent with the first action type.
[0014] Optionally, in one possible implementation of the first aspect, the trained first network is obtained by training the first network with the first training data as input and the first loss function value being less than a first threshold. The first training data includes a third action state, a first action type, and attributes related to the first action. The third action state includes at least one of the third coordinate, third angle, and third velocity of the joint corresponding to the first limb action. The first loss function is used to indicate the difference between the fourth action state output by the first network and the first target action state. The fourth action state includes at least one of the fourth coordinate, fourth angle, and fourth velocity of the joint corresponding to the first limb action. The first target action state includes at least one of the first target coordinate, first target angle, and first target velocity. The first target action state and the third action state belong to the action states of two adjacent frames in the same action video.
[0015] In this possible implementation, training the first network makes the action state obtained in the next moment, based on the action state, action type, and action attributes of the previous moment, more consistent with the action type.
[0016] Optionally, in one possible implementation of the first aspect, the target object is a three-dimensional model used to perform the first limb action. Specifically, the first limb action corresponding to the first action type and the first action attribute can be redirected to the target object, thereby generating a general character animation.
[0017] Optionally, in one possible implementation of the first aspect, the above steps further include: obtaining the value of n; obtaining n sets of action states based on the first action state and the first information, including: obtaining the output through the first network based on the first action state and the first information and iterating to obtain n sets of action states.
[0018] In this possible implementation, the flexibility of video production is improved by introducing the number of predicted action states, or the number of action states, thereby enhancing the user experience.
[0019] Optionally, in one possible implementation of the first aspect, the termination condition mentioned above relates to at least one of the following: the completion progress of the first limb movement, the end time of the first action type, and parameters in the first action attribute. The termination condition is either pre-set or input by the user on a third user interface. For example, if the first action type is "walk," the termination condition could be related to the first action attribute ("target"). If the target is another character, the prediction can stop after reaching the other character's position, i.e., the walking action can be stopped. If the termination condition includes the completion progress of the first limb movement, this completion progress can be obtained through prediction processing based on the first information, the first action state, and the first time period (first start time and first end time) of the first action type. Specifically, the first information, the first action state, and the first time period are input into the first network to obtain n sets of action states and completion progress.
[0020] In this possible implementation, by introducing a termination condition, the number of predicted action states, or the total number of action states, can be determined, thereby improving the user experience.
[0021] Optionally, in one possible implementation of the first aspect, the above steps further include: acquiring environmental information, which includes at least one of props, objects, and terrain where the target object interacts with the target object; obtaining n sets of action states based on the first action state and the first information, including: obtaining n sets of action states and n environmental contact information corresponding to the n sets of action states based on the first action state, the first information, and the environmental information, wherein one of the environmental contact information is used to indicate whether a joint in the action state corresponding to the environmental contact information is in contact with the environmental information; processing the limb movements of the target object based on the n sets of action states to obtain n images, including: processing the limb movements of the target object based on the n sets of action states and the environmental contact information to obtain n images. Props may include trees, tools, or stools, vehicles, etc., that may come into contact with the target object, and are not specifically limited here. The method of acquiring environmental information (or scene) may be through user import of scene model files (e.g., user creation and import of a scene, or user selection of a scene from a database), or user building a scene on the user interface displayed on the device, etc., and is not specifically limited here.
[0022] In this possible implementation, environmental information is introduced to predict environmental contact information, thereby generating animations that interact with the environment (props, other characters, etc.), thus increasing the diversity of generated animations.
[0023] Optionally, in one possible implementation of the first aspect, the above steps further include: obtaining second information, the second information including a second action type and a second action attribute, the second action type being used to describe a second limb action, and the second action attribute being used to describe the process of the second limb action occurring; obtaining a fifth action state; obtaining m sets of action states based on the fifth action state and the second information, where m is a positive integer; processing the limb action of the target object based on the fifth action state to obtain a second image; processing the limb action of the target object based on the m sets of action states to obtain m images; and generating a target video based on the first image and the n images, including: generating a target video based on the first image, the n images, the second image, and the m images.
[0024] In this possible implementation, the target video includes a first limb movement corresponding to a first action type and a second limb movement corresponding to a second action type. For example, if the first action type is walking and the second action type is running, the generated target video can be an ordered image that includes walking and running, that is, a continuous action video can be generated.
[0025] Optionally, in one possible implementation of the first aspect, the first information further includes a first time period of a first action type, and the second information further includes a second time period of a second action type; generating a target video based on a first image, n images, a second image, and m images includes: generating a target video based on a first image, n images, a second image, m images, a first time period, and a second time period, wherein the first time period corresponds to the first image and n images, and the second time period corresponds to the second image and m images.
[0026] In this possible implementation, the first image, n images, the second image, and m images are processed based on a first time period of a first action type and a second time period of a second action type to generate the target video. This ensures the continuity of the target video.
[0027] Optionally, in one possible implementation of the first aspect, the above steps: obtaining m sets of action states based on the fifth action state and the second information include: inputting the fifth action state and the second information into a trained second network to obtain m sets of action states, wherein the second network is used to obtain a seventh action state based on the sixth action state, the second action type and the second action attribute, and the change from the sixth action state to the seventh action state is related to the second action type.
[0028] In this possible implementation, m sets of action states can be obtained through the fifth action state and the second network. Since the second network is trained, the generated action states are more consistent with the second action type.
[0029] Optionally, in one possible implementation of the first aspect, the trained second network is obtained by training the second network with the second training data as input and the second loss function value being less than a second threshold as the objective. The second training data includes a sixth action state, a second action type, and second action attributes. The sixth action state includes at least one of the sixth coordinate, sixth angle, and sixth velocity of the joint corresponding to the second limb action. The second loss function is used to indicate the difference between the seventh action state output by the second network and the second target action state. The seventh action state includes at least one of the seventh coordinate, seventh angle, and seventh velocity of the joint corresponding to the second limb action. The second target action state includes at least one of the second target coordinate, second target angle, and second target velocity. The second target action state and the sixth action state belong to the action states corresponding to two adjacent frames in the same action video.
[0030] In this possible implementation, by training the second network, the seventh action state predicted subsequently based on the sixth action state, the second action type, and the second action attribute is made to better match the second action type.
[0031] Optionally, in one possible implementation of the first aspect, the above steps further include: obtaining third information, the third information including a third action type and a third action attribute, the third action type being used to describe a third limb action, the third action attribute being used to describe the process of the third limb action occurring, and the third limb corresponding to the third limb action being a local limb in the first limb corresponding to the first limb action; obtaining an eighth action state; obtaining p sets of action states based on the eighth action state and the third information, where p is a positive integer; processing the limb actions of the target object based on the first action state to obtain a first image, including: using the limb actions of the target object corresponding to the eighth action state to cover the limb actions of the target object corresponding to the first action state based on the coverage relationship between the third limb and the first limb to obtain a first image; processing the limb actions of the target object based on n sets of action states to obtain n images, including: using the limb actions of the target object corresponding to p sets of action states to cover the limb actions of the target object corresponding to n sets of action states based on the coverage relationship between the third limb and the first limb to obtain n images.
[0032] In this possible implementation, by overlaying the local limb (i.e., the third limb) with the first limb, the limb action corresponding to the local limb is overlaid onto the limb action corresponding to the first limb. This allows adjustment of a certain local action in the first action sequence. In other words, by introducing multiple local limb actions, complex animations can be generated.
[0033] Optionally, in one possible implementation of the first aspect, the above steps further include: acquiring facial information, which includes facial expression types and corresponding expression attributes, wherein the facial expression type describes the facial movements of the target object, and the expression attributes describe the amplitude of the facial movements; acquiring a facial expression sequence based on the facial information and a first association relationship, wherein the first association relationship represents the association between the facial information and the facial expression sequence; processing the target object's limb movements based on a first action state to obtain a first image, including: processing the target object's limb movements and facial movements based on the first action state and the facial expression sequence to obtain the first image; processing the target object's limb movements based on n sets of action states to obtain n images, including: processing the target object's limb movements and facial movements based on n sets of action states and the facial expression sequence to obtain n images. The facial expression type may include any one of neutral, happy, sad, surprised, angry, disgusted, fearful, delighted, tired, embarrassed, or contemptuous expressions.
[0034] In this possible implementation, facial expression sequences are obtained through facial information, so that the subsequently generated video includes not only action types but also facial expressions. This can be used to generate animations with high detail requirements, or in other words, to generate high-quality animations.
[0035] Optionally, in one possible implementation of the first aspect, the above steps further include: acquiring text information, which may include the target object's lines and the tone corresponding to the lines; generating speech segments based on the text information; generating lip-sync sequences based on the speech segments, the lip-sync sequences being used to describe the lip movements of the target object; processing the target object's limb movements based on a first action state to obtain a first image, including: processing the target object's limb movements and facial movements based on the first action state and the lip-sync sequence to obtain the first image; processing the target object's limb movements based on n sets of action states to obtain n images, including: processing the target object's limb movements and facial movements based on n sets of action states and the lip-sync sequence to obtain n images.
[0036] In this possible implementation, the lip-sync sequence can also be obtained through the fifth piece of information, so that the lip-sync of the target object in the subsequently generated video can change according to the speech segment (or the target object's lines), making the generated video more detailed.
[0037] Optionally, in one possible implementation of the first aspect, the above step of generating a target video based on the first image and n images includes: generating a target video based on the first image, the n images, and an audio segment.
[0038] In this possible implementation, voice segments can also be introduced, thereby generating animations that correspond to actions, expressions, and lip movements, which is suitable for high-quality animation production scenarios.
[0039] Optionally, in one possible implementation of the first aspect, the aforementioned first action type includes at least one of walking, running, jumping, sitting down, standing up, squatting down, lying down, hugging, punching, waving a sword, dancing, etc.
[0040] Optionally, in one possible implementation of the first aspect, the aforementioned first action attribute includes at least one of target position, displacement, travel path, action speed, frequency of action occurrence, action amplitude, and action orientation.
[0041] For example, if the first action type is kicking, then the first action attribute can be action speed or the direction of the action. As another example, if the first action type is sitting down, then the first action attribute is the target position.
[0042] In this possible implementation, the higher-level semantics of the first action type and the first action attribute are more efficient, intuitive, and understandable. Compared with the existing technology that requires users to input low-level control signals, it can reduce the user's workload and professional requirements, thereby improving the efficiency of subsequent target video generation.
[0043] The second aspect of this application provides a data processing device that can be applied to the automatic and rapid production of children's educational animations, short video animations, promotional animations, variety show animations, film and television pre-show animations, etc., or to automatic motion control in games, or to automatically generate character actions in interactive applications. The data processing device includes:
[0044] The acquisition unit is used to acquire first information, which includes a first action type and a first action attribute. The first action type is used to describe the first limb action, and the first action attribute is used to describe the process in which the first limb action occurs.
[0045] The acquisition unit is also used to acquire the first action state;
[0046] The prediction unit is used to obtain n sets of action states based on the first action state and the first information, where n is a positive integer;
[0047] The acquisition unit is also used to acquire the target object;
[0048] The processing unit is used to process the limb movements of the target object based on the first action state to obtain a first image;
[0049] The processing unit is also used to process the limb movements of the target object based on n sets of action states to obtain n images;
[0050] The generation unit is used to generate a target video based on a first image and n images, wherein the target video is related to a first action type and a first action attribute.
[0051] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is specifically used to acquire a first script input by the user, the first script including a first action type and a first action attribute; or, the acquisition unit is specifically used to display a first user interface, the first user interface including action type options and / or action attribute options; the acquisition unit is specifically used to respond to a first operation by the user on the first user interface, determine the first action type from the action type options, and / or select the first action attribute from the action attribute options.
[0052] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is specifically used to acquire a second action state, which represents a historical action state prior to the first action state; the acquisition unit is specifically used to input the second action state and first information into a trained first network to obtain a first action state; or, the acquisition unit is specifically used to display a second user interface to the user, the second user interface including an action state area; the acquisition unit is specifically used to respond to a second operation by the user on the second user interface and determine the first action state from the action state area; or, the acquisition unit is specifically used to acquire a pre-set first action state.
[0053] Optionally, in one possible implementation of the second aspect, the prediction unit described above is specifically used to input the first action state and the first information into the trained first network to obtain n sets of action states. The first network is used to obtain the fourth action state based on the third action state, the first action type and the first action attribute. The change from the third action state to the fourth action state is related to the first action type.
[0054] Optionally, in one possible implementation of the second aspect, the trained first network is obtained by training the first network with the first training data as input and the first loss function value being less than a first threshold. The first training data includes a third action state, a first action type, and a first action attribute. The third action state includes at least one of the third coordinate, third angle, and third velocity of the joint corresponding to the first limb action. The first loss function is used to indicate the difference between the fourth action state output by the first network and the first target action state. The fourth action state includes at least one of the fourth coordinate, fourth angle, and fourth velocity of the joint corresponding to the first limb action. The first target action state includes at least one of the first target coordinate, first target angle, and first target velocity. The first target action state and the third action state belong to the action states corresponding to two adjacent frames in the same action video.
[0055] Alternatively, in one possible implementation of the second aspect, the target object described above is a three-dimensional model, and the target object is used to perform the first limb action.
[0056] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is further used to acquire the value of n; the prediction unit is specifically used to obtain the output through the first network based on the first action state and the first information and to perform iteration to obtain n sets of action states.
[0057] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is further configured to acquire an end condition, which includes at least one of the following: the completion progress of the first limb action, the end time of the first action type, and the parameters in the first action attribute. The end condition is either preset or input by the user on a third user interface. Specifically, the acquisition unit is configured to determine the value of n based on the end condition.
[0058] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is further configured to acquire environmental information, which includes at least one of props and objects that interact with the target object; the prediction unit is specifically configured to obtain n sets of action states and n environmental contact information corresponding to the n sets of action states based on the first action state, the first information, and the environmental information, wherein one of the environmental contact information is used to indicate whether the joint in the action state corresponding to the environmental contact information is in contact with the environmental information; and the processing unit is specifically configured to process the limb movements of the target object based on the n sets of action states and the environmental contact information to obtain n images.
[0059] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is further configured to acquire second information, the second information including a second action type and a second action attribute, the second action type describing a second limb action and the second action attribute describing the process of the second limb action occurring; the acquisition unit is further configured to acquire a fifth action state; the prediction unit is further configured to obtain m sets of action states based on the fifth action state and the second information, where m is a positive integer; the processing unit is further configured to process the limb actions of the target object based on the fifth action state to obtain a second image; the processing unit is further configured to process the limb actions of the target object based on the m sets of action states to obtain m images; and the generation unit is specifically configured to generate a target video based on the first image, n images, the second image, and the m images.
[0060] Optionally, in one possible implementation of the second aspect, the first information mentioned above further includes a first time period of a first action type, and the second information further includes a second time period of a second action type; the generation unit is specifically used to generate a target video based on a first image, n images, a second image, m images, a first time period, and a second time period, wherein the first time period corresponds to the first image and n images, and the second time period corresponds to the second image and m images.
[0061] Optionally, in one possible implementation of the second aspect, the prediction unit described above is specifically used to input the fifth action state and the second information into the trained second network to obtain m sets of action states. The second network is used to obtain the seventh action state based on the sixth action state, the second action type, and the second action attribute. The change from the sixth action state to the seventh action state is related to the second action type.
[0062] Optionally, in one possible implementation of the second aspect, the trained second network is obtained by training the second network with the second training data as input and the second loss function value being less than a second threshold as the objective. The second training data includes a sixth action state, a second action type, and second action attributes. The sixth action state includes at least one of the sixth coordinate, sixth angle, and sixth velocity of the joint corresponding to the second limb action. The second loss function is used to indicate the difference between the seventh action state output by the second network and the second target action state. The seventh action state includes at least one of the seventh coordinate, seventh angle, and seventh velocity of the joint corresponding to the second limb action. The second target action state includes at least one of the second target coordinate, second target angle, and second target velocity. The second target action state and the sixth action state belong to the action states corresponding to two adjacent frames in the same action video.
[0063] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is further configured to acquire third information, which includes a third action type and a third action attribute. The third action type describes a third limb action, and the third action attribute describes the process of the third limb action. The third limb corresponding to the third limb action is a local limb in the first limb corresponding to the first limb action. The acquisition unit is further configured to acquire an eighth action state. The prediction unit is further configured to obtain p groups of action states based on the eighth action state and the third information, where p is a positive integer. The processing unit is specifically configured to use the limb action of the target object corresponding to the eighth action state to cover the limb action of the target object corresponding to the first action state based on the coverage relationship between the third limb and the first limb to obtain a first image. The processing unit is specifically configured to use the limb action of the target object corresponding to p groups of action states to cover the limb action of the target object corresponding to n groups of action states based on the coverage relationship between the third limb and the first limb to obtain n images.
[0064] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is further configured to acquire facial information, including facial expression types and expression attributes corresponding to the facial expression types. The facial expression types are used to describe the facial movements of the target object, and the expression attributes are used to describe the amplitude of the facial movements. The acquisition unit is further configured to acquire a facial expression sequence based on the facial information and a first association relationship, whereby the first association relationship represents the association between the facial information and the facial expression sequence. The processing unit is specifically configured to process the limb movements and facial movements of the target object based on the first action state and the facial expression sequence to obtain a first image. The processing unit is specifically configured to process the limb movements and facial movements of the target object based on n sets of action states and facial expression sequences to obtain n images.
[0065] Optionally, in one possible implementation of the second aspect, the aforementioned acquisition unit is further configured to acquire text information; the generation unit is further configured to generate speech segments based on the text information; the generation unit is further configured to generate lip-sync sequences based on the speech segments, the lip-sync sequences being used to describe the lip-sync of the target object; the processing unit is specifically configured to process the limb and facial movements of the target object based on the first action state and the lip-sync sequence to obtain a first image; the processing unit is specifically configured to process the limb and facial movements of the target object based on n sets of action states and lip-sync sequences to obtain n images.
[0066] Alternatively, in one possible implementation of the second aspect, the aforementioned generation unit is specifically used to generate a target video based on the first image, n images, and a speech segment.
[0067] Optionally, in one possible implementation of the second aspect, the first action type mentioned above includes at least one of walking, running, jumping, sitting down, standing up, squatting down, lying down, hugging, punching, swinging a sword, dancing, etc.
[0068] Optionally, in one possible implementation of the second aspect, the aforementioned first action attribute includes at least one of target position, displacement, travel path, action speed, frequency of action occurrence, action amplitude, and action orientation.
[0069] A third aspect of this application provides a data processing apparatus that performs the methods described in the first aspect or any possible implementation thereof.
[0070] A fourth aspect of this application provides a data processing apparatus, comprising: a processor coupled to a memory for storing programs or instructions, wherein when the programs or instructions are executed by the processor, the data processing apparatus implements the methods described in the first aspect or any possible implementation thereof.
[0071] The fifth aspect of this application provides a computer-readable medium having a computer program or instructions stored thereon, which, when run on a computer, cause the computer to perform the methods of the first aspect or any possible implementation thereof.
[0072] The sixth aspect of this application provides a computer program product that, when executed on a computer, causes the computer to perform the methods of the first aspect or any possible implementation thereof.
[0073] The technical effects of the second, third, fourth, fifth, and sixth aspects or any of their possible implementations can be found in the first aspect or the technical effects of different possible implementations of the first aspect, and will not be repeated here.
[0074] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: n sets of action states are obtained through the first information and the first action state; and the limb movements of the target object are processed based on the first action state and the n sets of action states to obtain a first image and n images, thereby generating a target video based on the first image and the n images. Since the higher-level semantics of the first action type and the first action attribute in the first script are more efficient, intuitive, and understandable, compared to the prior art which requires user input of lower-level control signals, it can reduce the user's workload and professional requirements, thereby improving the efficiency of subsequent target video generation. Attached Figure Description
[0075] Figure 1 A schematic diagram of an artificial intelligence main framework provided in an embodiment of the present invention;
[0076] Figure 2 This is a schematic diagram of the system architecture provided in the embodiments of this application;
[0077] Figure 3 A flowchart illustrating the prediction network training method provided in this application embodiment;
[0078] Figure 4 An example diagram of a joint architecture provided in this application embodiment;
[0079] Figure 5 A flowchart illustrating the data processing method provided in this application embodiment;
[0080] Figure 6 Another flowchart illustrating the data processing method provided in this application embodiment;
[0081] Figure 7 A structural example diagram of a first script provided in an embodiment of this application;
[0082] Figures 8 to 11 Schematic diagrams of several user interfaces provided for embodiments of this application;
[0083] Figure 12 A structural example diagram of a second script provided in an embodiment of this application;
[0084] Figure 13 An example diagram of environmental information provided in an embodiment of this application;
[0085] Figure 14 Another flowchart illustrating the data processing method provided in this application embodiment;
[0086] Figure 15 An example diagram of a target object provided in an embodiment of this application;
[0087] Figure 16 A schematic diagram of a first limb movement provided in an embodiment of this application;
[0088] Figure 17 A schematic diagram of a first image provided for an embodiment of this application;
[0089] Figure 18 A schematic diagram illustrating another first limb movement provided in an embodiment of this application;
[0090] Figure 19 This is a schematic diagram of one of n images provided in an embodiment of this application;
[0091] Figure 20 An example diagram illustrating the structure of another first script provided in this application embodiment;
[0092] Figure 21 An example diagram illustrating the structure of another first script provided in this application embodiment;
[0093] Figure 22 A schematic diagram of another user interface provided for an embodiment of this application;
[0094] Figure 23 A schematic diagram of a second limb movement provided in an embodiment of this application;
[0095] Figure 24 A schematic diagram of a limb movement provided for an embodiment of this application;
[0096] Figure 25 This is a schematic diagram of one of m images provided in an embodiment of this application;
[0097] Figure 26 A structural example diagram of a second script provided in an embodiment of this application;
[0098] Figures 27 to 30 Schematic diagrams of several other user interfaces provided for embodiments of this application;
[0099] Figure 31 A schematic diagram of a second partial limb movement provided in an embodiment of this application;
[0100] Figure 32 A schematic diagram illustrating an embodiment of this application for updating a first limb action based on a second local limb action;
[0101] Figure 33 A schematic diagram illustrating an embodiment of this application for updating limb movements based on a second local limb movement;
[0102] Figure 34 This is a schematic diagram of one of n images updated using a second local limb movement, provided as an embodiment of this application.
[0103] Figure 35 A structural example diagram of a third script provided in an embodiment of this application;
[0104] Figures 36 to 39 Schematic diagrams of several other user interfaces provided for embodiments of this application;
[0105] Figure 40 A structural example diagram of a fourth script provided in an embodiment of this application;
[0106] Figures 41 to 43 Schematic diagrams of several other user interfaces provided for embodiments of this application;
[0107] Figure 44 A schematic diagram of the structure of the data processing device provided in the embodiments of this application;
[0108] Figure 45 Another structural schematic diagram of the data processing device provided in the embodiments of this application;
[0109] Figure 46 Another structural schematic diagram of the data processing device provided in the embodiments of this application. Detailed Implementation
[0110] This application provides a data processing method and related equipment. Compared to existing technologies that require user input of low-level control signals, this method reduces user workload and professional requirements, thereby improving the efficiency of subsequent target video generation.
[0111] The technical solutions of the embodiments of the present invention will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0112] To facilitate understanding, the relevant terms and concepts mainly involved in the embodiments of this application will be introduced below.
[0113] 1. Neural Networks
[0114] Neural networks can be composed of neural units, which can refer to units such as X. s The arithmetic unit that takes an intercept of 1 as input can output the following:
[0115]
[0116] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For X s The weights are denoted by b, where b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input to the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.
[0117] 2. Deep Neural Networks
[0118] A deep neural network (DNN), also known as a multilayer neural network, can be understood as a neural network with many hidden layers, though there's no specific metric for "many." DNNs can be categorized into three types based on their layer positions: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. However, deep neural networks may not necessarily include hidden layers; this is not a limitation here.
[0119] The function of each layer in a deep neural network can be expressed mathematically. To describe it: From a physical perspective, the work of each layer in a deep neural network can be understood as transforming the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / decrease; 2. Magnification / scaling; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are performed by [the layer name is missing in the original text], while operation 4 is performed by [the layer name is missing in the original text]. The operation 5 is then implemented by α(). The term "space" is used here because the object being classified is not a single thing, but a class of things; space refers to the set of all individuals within this class of things. Here, is the weight vector, where each value represents the weight of a neuron in that layer of the neural network. This vector W determines the spatial transformation from the input space to the output space, as described above; that is, the weights W of each layer control how the space is transformed. The purpose of training a deep neural network is to ultimately obtain the weight matrix of all layers of the trained neural network (a weight matrix formed by the vectors W from many layers). Therefore, the training process of a neural network is essentially learning how to control spatial transformation, more specifically, learning the weight matrix.
[0120] 3. Convolutional Neural Networks
[0121] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A CNN contains a feature extractor consisting of convolutional layers and subsampling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as performing convolution between the same trainable filter and an input image or a convolutional feature map. A convolutional layer is a layer of neurons in a CNN that performs convolutional processing on the input signal. In a convolutional layer of a CNN, a neuron may only be connected to some of its neighboring neurons. A convolutional layer typically contains several feature maps, each composed of rectangularly arranged neural units. Neural units within the same feature map share weights, which are the convolutional kernel. Shared weights can be understood as the way image information is extracted regardless of location. The underlying principle is that the statistical information of one part of an image is the same as that of other parts. This means that image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all locations in an image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.
[0122] Convolutional kernels can be initialized as matrices of random size. During the training of the convolutional neural network, the kernels can learn to acquire appropriate weights. Furthermore, the direct benefit of shared weights is reducing the connections between layers of the convolutional neural network, while also reducing the risk of overfitting. The separation network, recognition network, detection network, depth estimation network, and other networks in the embodiments of this application can all be CNNs.
[0123] 4. Backpropagation Algorithm
[0124] Convolutional neural networks can employ backpropagation (BP) to correct the parameters in the initial super-resolution model during training, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates an error loss; this error loss information is then propagated back to update the parameters in the initial super-resolution model, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the super-resolution model, such as the weight matrix.
[0125] 5. Recurrent Neural Networks
[0126] In traditional neural network models, layers are fully connected, but nodes within each layer are unconnected. However, this type of ordinary neural network cannot solve many problems. For example, predicting the next word in a sentence, because words in a sentence are not independent; the preceding words are usually needed. Recurrent neural networks (RNNs) refer to a sequence where the current output is related to the previous outputs. Specifically, the network memorizes previous information, storing it in its internal state, and applies it to the calculation of the current output.
[0127] 6. Feedforward Neural Network
[0128] Feedforward neural networks (FNNs) are the earliest invented simple artificial neural networks. In a feedforward neural network, each neuron belongs to a different layer. Neurons in each layer receive signals from neurons in the previous layer and output signals to the next layer. Layer 0 is called the input layer, the last layer is called the output layer, and the other intermediate layers are called hidden layers. There is no feedback in the entire network; signals propagate unidirectionally from the input layer to the output layer.
[0129] 7. Multilayer perceptron (MLP)
[0130] A multilayer perceptron, also known as a multilayer perceptron, is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset.
[0131] 8. Transformer: A neural network structure consisting of a self-attention mechanism and a feedforward network.
[0132] 9. Loss Function
[0133] In training a deep neural network, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value and update the weight vector of each layer based on the difference. (Of course, there's usually an initialization process before the first update, where parameters are pre-configured for each layer.) For example, if the network's prediction is too high, the weight vector is adjusted to predict a lower value, and this adjustment continues until the network can predict the target value accurately. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss.
[0134] 10. Animation
[0135] Virtually created video content includes animated videos displayed on a 2D plane, as well as 3D animated content displayed on 3D display devices such as augmented reality (AR), virtual reality (VR), and holographic displays; its style is not limited to cartoon style, but also includes realistic style, such as digital human animation, special effects films, etc.
[0136] 11. Roles
[0137] Characters can include bipedal characters or quadrupedal characters. Bipedal characters refer to humanoid characters, anthropomorphic animals, robots, monsters, etc., that stand on two legs; quadrupedal characters refer to characters that stand on four legs, which can be quadrupedal animals, quadrupedal monsters, etc.
[0138] 12. Standard Role
[0139] Standard character: For bipedal characters, the standard character is a character model that conforms to the standard human form; for quadrupedal characters, the standard character is a character model that conforms to the standard quadrupedal animal form.
[0140] 13. Character binding: Embed the skeleton into the model and calculate the weight of each vertex of the model relative to each bone.
[0141] 14. Action Redirection: Transfer the body and facial movements of one character to another character with a different appearance while ensuring that the movements are consistent with the original movements.
[0142] 15. Blendshape: A technique for deforming a mesh to achieve a combination of multiple predefined meshes.
[0143] 16. General Character Animation: This includes character body movements, facial expressions, and lip movements. It is able to interact with the environment and has a storyline, such as short videos, cartoons, variety shows, and films.
[0144] 17. Redirection
[0145] Redirection technology is a process of copying animation data from one skeleton to another, essentially a "copying" process. Animation redirection technology has been widely used in many areas. For example, the motion capture technology commonly used in AAA console games is based on this principle—using image recognition and other technologies to generate animation information from real-life human movements and applying it to virtual characters, saving it as animation data.
[0146] The system architecture provided in the embodiments of this application is described below.
[0147] See appendix Figure 1This invention provides a system architecture 100. As shown in the system architecture 100, a data acquisition device 160 is used to collect training data. In this embodiment, the training data includes: action state, action type, and action attributes corresponding to the action type. The action type describes a sequence of limb movements, the action attributes describe parameters related to the sequence of limb movements, and the action state determines the limb movements in the sequence. Further, the training data may also include environmental information (e.g., props, terrain, other characters, etc.) and / or environmental contact information (e.g., contact information between limb movements and props), and completion progress (e.g., values between 0 and 1). The training data is stored in a database 130, and a training device 120 trains a target model / rule 101 based on the training data maintained in the database 130. The following describes in more detail how the training device 120 obtains the target model / rule 101 based on the training data. This target model / rule 101 can be used to implement the data processing method provided in this embodiment. Specifically, the target model / rule 101 in this embodiment can be a prediction network. It should be noted that in practical applications, the training data maintained in the database 130 may not all come from the data acquisition device 160; it may also be received from other devices. Furthermore, it should be noted that the training device 120 may not necessarily train the target model / rule 101 entirely based on the training data maintained in the database 130; it may also obtain training data from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application.
[0148] The target model / rule 101 trained using training device 120 can be applied to different systems or devices, such as... Figure 1 The execution device 110 shown can be a terminal, such as a mobile phone terminal, tablet computer, laptop computer, AR / VR, vehicle terminal, etc., or it can be a server or cloud platform. (See attached...) Figure 1 In this embodiment, the execution device 110 is equipped with an I / O interface 112 for data interaction with external devices. Users can input data into the I / O interface 112 through the client device 140. This input data may include: action type, action attributes, and action state. Optionally, the input data may also include environmental information surrounding the character, contact information between the character's joints and the environment, and completion progress. Of course, the input data can also be a video of the character's actions. Furthermore, the input data can be input by the user, uploaded by the user through a shooting device, or from a database; specific details are not limited here.
[0149] The preprocessing module 113 is used to preprocess the input data (if the input data is an action video) received by the I / O interface 112. In this embodiment, the preprocessing module 113 can be used to group the action segments in the action video based on the action attributes, so as to train a machine learning model for each group of data.
[0150] During the preprocessing of input data by the execution device 110, or during the calculation module 111 of the execution device 110 performing calculations and other related processes, the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0151] Finally, I / O interface 112 returns the processing result, such as the predicted action status obtained above, to client device 140, thereby providing it to the user.
[0152] It is worth noting that the training device 120 can generate corresponding target models / rules 101 based on different training data for different objectives or tasks. The corresponding target models / rules 101 can be used to achieve the above objectives or complete the above tasks, thereby providing the user with the required results.
[0153] In the appendix Figure 1 In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 112. Alternatively, the client device 140 can automatically send input data to I / O interface 112. If user authorization is required for the client device 140 to automatically send input data, the user can set the corresponding permissions in the client device 140. The user can view the output results of the execution device 110 on the client device 140, which can be presented in various forms such as display, sound, or animation. The client device 140 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130. Alternatively, data can be collected directly from the I / O interface 112 without going through the client device 140, using the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130.
[0154] It is worth noting that, attached Figure 1 This is merely a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationships between the devices, components, modules, etc. shown in the diagram do not constitute any limitation. For example, in the attached diagram... Figure 1In this context, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 may also be placed within the execution device 110.
[0155] like Figure 1 As shown, the target model / rule 101 is trained by the training device 120. In this embodiment of the application, the target model / rule 101 can be a prediction network. Specifically, in the network provided in this embodiment of the application, the prediction network can use multilayer perceptron (MLP), long short-term memory network (LSTM), graph convolutional network (GCN), graph neural network (GNN), transformer, etc., and no specific limitation is made here.
[0156] The following describes a chip hardware structure provided by an embodiment of this application.
[0157] Figure 2 A chip hardware structure provided in this embodiment of the invention includes a neural network processor 20. This chip can be configured as follows: Figure 1 The execution device 110 shown is used to perform the calculations of the calculation module 111. This chip can also be located in, for example... Figure 1 The training device 120 shown is used to complete the training work of the training device 120 and output the target model / rule 101.
[0158] The neural network processor 20 can be any processor suitable for large-scale XOR operations, such as a neural network processing unit (NPU), tensor processing unit (TPU), or graphics processing unit (GPU). Taking an NPU as an example: the neural network processor 20 is mounted as a coprocessor on the main central processing unit (CPU) (host CPU), and tasks are assigned by the main CPU. The core of the NPU is the arithmetic circuit 203, and the controller 204 controls the arithmetic circuit 203 to retrieve data from the memory (weight memory or input memory) and perform operations.
[0159] In some implementations, the arithmetic circuit 203 internally includes multiple process engines (PEs). In some implementations, the arithmetic circuit 203 is a two-dimensional pulsating array. The arithmetic circuit 203 can also be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 203 is a general-purpose matrix processor.
[0160] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 202 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 201 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 208.
[0161] The vector computation unit 207 can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector computation unit 207 can be used for network computation in non-convolutional / non-FC layers of neural networks, such as pooling, batch normalization, local response normalization, etc.
[0162] In some implementations, the vector computation unit 207 can store the processed output vector into a unified buffer 206. For example, the vector computation unit 207 can apply a nonlinear function to the output of the arithmetic circuit 203, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit 207 generates normalized values, merged values, or both. In some implementations, the processed output vector can be used as activation input to the arithmetic circuit 203, for example, for use in subsequent layers of a neural network.
[0163] The unified memory 206 is used to store input data and output data.
[0164] The weight data is directly transferred from the external memory to the input memory 201 and / or the unified memory 206 through the direct memory access controller 205 (DMAC), the weight data in the external memory is stored in the weight memory 202, and the data in the unified memory 206 is stored in the external memory.
[0165] The bus interface unit (BIU) 210 is used to enable interaction between the main CPU, DMAC and instruction fetch memory 209 via the bus.
[0166] The instruction fetch buffer 209, which is connected to the controller 204, is used to store the instructions used by the controller 204.
[0167] The controller 204 is used to call the instructions cached in the instruction memory 209 to control the operation of the computing accelerator.
[0168] Generally, the unified memory 206, input memory 201, weight memory 202, and instruction fetch memory 209 are all on-chip memories, while the external memory is memory outside the NPU. This external memory can be double data rate synchronous dynamic random access memory (DDR SDRAM), high bandwidth memory (HBM), or other readable and writable memory.
[0169] The training method and data processing method of the prediction network according to the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0170] First, combined Figure 3 The training method of the prediction network in the embodiments of this application will be described in detail. Figure 3 The method shown can be executed by a training device for a prediction network. This training device can be a cloud server or a terminal device, such as a computer or server with sufficient computing power to execute the prediction network training method. It can also be a system composed of a cloud server and terminal devices. For example, the training method can be performed by… Figure 1 Training equipment 120 Figure 2 The neural network processor 20 in the system executes the commands.
[0171] Optionally, the training method can be processed by the CPU, or by both the CPU and GPU, or it can use other processors suitable for neural network computing instead of a GPU. This application does not impose any restrictions.
[0172] Please see Figure 3 The training method for the prediction network provided in this application may include steps 301 to 304. Steps 301 to 304 are described in detail below.
[0173] Step 301: Obtain the action dataset.
[0174] Obtain motion datasets (also known as motion capture datasets or multiple motion videos) of various commonly used characters (or target objects). These datasets include multiple motion segments, each of which may contain one or more frames. Each motion segment corresponds to a specific motion type, and this motion type is identified. Alternatively, after obtaining the motion dataset, it can be split into multiple motion segments according to the rule that one motion segment corresponds to one motion type. Furthermore, motion segments in the dataset that contain multiple motion types at the same moment can be discarded. This ensures that one motion segment corresponds to one motion type, which helps improve the accuracy of the prediction network trained subsequently.
[0175] In this application embodiment, the character can refer to a human body, animal, robot, anthropomorphic animal, monster, etc., and can be a bipedal character, a quadrupedal character, or a multi-legged character, etc., without being limited here. This method embodiment is only described using the human body as an example.
[0176] In this embodiment of the application, the action dataset can be obtained by user input, by user collection and uploading through a shooting device, or from a database, etc., and the specific method is not limited here.
[0177] The above-mentioned action types are used to describe the limb movements (or sequences of limb movements) of humans or animals.
[0178] In this embodiment of the application, the limb movements are described only as examples of movements of the whole body. It can be understood that limb movements can also refer to the movements of some limbs (or local limbs) in the whole body, etc., and are not limited here.
[0179] Optionally, the action type of an action segment can describe full-body movements, such as walking, running, jumping, sitting down, standing up, squatting, lying down, hugging, boxing, swinging a sword, dancing, etc. The action type of an action segment can also describe partial body movements, such as raising one's head, waving one's hand, kicking one's leg, wagging one's tail, etc., without any specific limitations here.
[0180] Optionally, to obtain animations of character interaction with the environment, one or more prop models and other character models can be selected from a database (e.g., a prop library) and inserted into the aforementioned action clips. The dimensions, aspect ratios, etc., of the props or other characters can be randomly modified, and the actions in the action clips can be adjusted based on the inserted props or other characters (or other objects), thus ensuring the action clips match the props or other characters. This allows for the acquisition of environmental geometry information and contact tags between the character's joints and the environment. Environmental geometry information represents props, other characters, terrain, etc., within a certain range from the character. For example, environmental geometry information represents props or other characters inserted within a 5-meter radius and 2-meter height of the character's current position. Contact tags indicate whether the character's joints are in contact with props or other characters (or the contact point).
[0181] In this application's embodiments, joints can be used to determine a character's limb movements. A joint can be understood as a movable bone connection, or a junction between bones. There are various types of joints; this application's embodiments only use one example. Figure 4 Taking the joint type shown as an example, it's understandable that the joint type can also refer to the nine major joints (foot joint, ankle joint, knee joint, hip joint, waist joint, shoulder joint, elbow joint, wrist joint, etc.). Of course, the joint type can also be set according to actual needs (for example, it can also include joints corresponding to the tail, horns, etc.), and this is not specifically limited here. Figure 4 The joint types shown include head, neck, thorax, chest, spine, root, collar, shoulder, elbow, wrist, hip, knee, ankle, and toe. The root joint is the apex of the pelvis. A character can have one or more joints. For example, in a human character, there are two collar, shoulder, elbow, wrist, knee, ankle, and toe joints: collar 1, collar 2, shoulder 1, shoulder 2, elbow 1, elbow 2, wrist 1, wrist 2, knee 1, knee 2, ankle 1, ankle 2, and toe 1 and toe 2.
[0182] Step 302: Group multiple action segments based on action attributes.
[0183] After acquiring multiple action segments, the action attribute corresponding to the action type of each action segment is determined. This action attribute describes the parameters related to the action of the first limb. The multiple action segments are then grouped based on the action attribute. That is, multiple action segments with the same action attribute are grouped together. In this embodiment, one action attribute can correspond to one or more action types. For example, the action types and the action attributes related to the action types can be shown in Table 1:
[0184] Table 1
[0185] Walk Goals and Paths jump Direction, amplitude run Goals and Paths look up amplitude Shaking head Amplitude, frequency wave amplitude … …
[0186] The action types and attributes in Table 1 are just examples. In actual applications, they can be set according to actual needs. For example, the action attribute for walking can also include stride length, and the action type for running can also include speed, etc. Specific limitations are not specified here. As can be seen from Table 1, the action attributes corresponding to the action types "walking" and "running" are both "target" and "path." Therefore, action segments with the action type "walking" and action segments with the action type "running" can be grouped together. Here, "target" can be understood as the target position or displacement, used to indicate the position where the limb ends the action. "Path" can be understood as the movement path. "Direction" can be understood as the direction of the action, "amplitude" as the amplitude of the action or the amplitude of the action corresponding to a certain part of the limb, and "frequency" as the frequency of the action or the frequency of the action corresponding to a certain part of the limb.
[0187] To facilitate understanding of the above grouping of multiple action fragments based on action attributes, the following description uses two groups of action fragments as examples. Please refer to Table 2:
[0188] Table 2
[0189] First set of action clips Walk, run, jump Goals and Paths Second set of action clips Throw, kick, throw Target, direction, speed … … …
[0190] Among them, the action attributes corresponding to the action types "walk", "run", and "jump" are all "target" and "path". Therefore, the action fragments with the action type "walk", "run", and "jump" can be grouped into one group (i.e., the first group of action fragments). The action attributes corresponding to the action types "throw", "kick", and "throw" are all "target", "direction", and "speed". Therefore, the action fragments with the action type "throw", "kick", and "throw" can be grouped into one group (i.e., the second group of action fragments).
[0191] Step 303: Obtain the first frame information of the first frame data and the second frame information of the second frame data in each action segment.
[0192] After grouping multiple action segments according to their action attributes, the first frame information of the first frame data and the second frame information of the second frame data in each group of action segments can be obtained. The first frame data occurs before the second frame data occurs; furthermore, the second frame data can be the frame following the first frame. The first frame information can include the action state of the first frame data, the action type corresponding to the first frame data, and the action attributes. The action state of the first frame data can include at least one of the coordinates, angle, and velocity of the joint corresponding to the first limb action in the first frame data. The second frame information is similar to the first frame information; it can include the state of the second frame data (or the first target action state), the action type corresponding to the second frame data, and the action attributes. The first target action state can include at least one of the first target coordinates, first target angle, and first target velocity of the joint corresponding to the first limb action in the second frame data.
[0193] In other words, the above can also be understood as: for each frame of data in each group of action segments, or taking the action state of the current frame data and the next frame data, the action state includes at least one of the following: joint coordinates, rotation angle, speed, etc.
[0194] Optionally, if the first frame data corresponds to the first frame in the action segment, the action state of the first frame data can also be understood as the initial action state. This initial action state includes the initial orientation of the limb movement, the initial rotation angle of the joint corresponding to the limb movement, and the coordinates of the root node in the limb movement. The initial orientation can be understood as the orientation of the character's first limb (or the orientation of the character's body), or the overall rotation of the character's limbs. This orientation can include east, south, west, and north, and can also include more detailed divisions, such as southeast, northeast, and northwest, etc., which are not limited here. For example, 0 represents east, 90 represents north, 180 represents west, and -90 represents south. The initial rotation angle is the rotation angle of each joint corresponding to the first limb movement. This rotation angle can be represented by attitude angle, quaternion, Euler angle, axis angle, etc. The attitude angle can be understood as representing the attitude of the joint, equivalent to the angle around the x, y, and z axes. The initial position of the root node can be represented by (x, z), and the initial position of the root node can be understood as the projection of the root node onto the ground.
[0195] For example, the initial orientation is 0, the root node coordinates are (0,0), and the initial orientation is 0 (i.e., the limb is facing east).
[0196] Optionally, each group of action segments may include one or more action segments, which is not limited here. As shown in Table 2, the first group of action segments and the second group of action segments each include action segments of 3 action types.
[0197] Optionally, the first frame information may further include the completion progress of the first frame data, and the second frame information may further include the completion progress of the second frame data. The completion progress of the first frame data represents the ratio of the time consumed to reach the first frame data to the total duration of the action segment containing the first frame data. The completion progress of the second frame data represents the ratio of the time consumed to reach the second frame data to the total duration of the action segment containing the second frame data, with the ratio ranging from 0 to 1 (or 0% to 100%). Furthermore, the first and second frame data can be in the same action segment or in different action segments; this is not specifically limited here.
[0198] Step 304: For each group of action segments, train the prediction network corresponding to the first action segment of each group.
[0199] For each set of action segments with consistent action attributes, a prediction network is trained. This ensures that the variables included in the action attributes of the action segments corresponding to a prediction network are consistent.
[0200] The trained prediction network is trained using training data as input and with the goal of reducing the loss function value to a certain threshold. This training data includes information from the first frame of each action segment (e.g., the action state, action type, and action attribute of the first frame). The loss function indicates the difference between the action state predicted by the prediction network and the first target action state in the second frame.
[0201] In this scenario, the prediction network is trained with the goal of reducing the loss function value to a certain threshold. This means continuously narrowing the difference between the action state output by the prediction network in the next frame and the actual action state in the next frame (i.e., the first target action state in the second frame information). This training process can be understood as a prediction task. The loss function can be understood as the loss function corresponding to the prediction task. The first target action state and the action state of the first frame data belong to the same collected action data, or can be understood as the action states of two adjacent frames within the same action data video.
[0202] In other words, for each action segment, a machine learning model is trained, taking the action state, action type, and action attributes of the current frame as input, and outputting the action state of the next frame. The change from the action state of the current frame to the action state of the next frame is related to the corresponding action segment.
[0203] For example, for the first set of action segments with action attributes of "target" and "path" (e.g., walking, running, jumping), the trained prediction network can be called the first network. For whole-body limb action segments with action attribute of "target" (e.g., sitting down), the trained prediction network can be called the second network. For partial limb action segments with action attribute of "target" (e.g., looking / staring), the trained prediction network can be called the third network. For action segments with action attribute of "amplitude" (e.g., waving), the trained prediction network can be called the fourth network.
[0204] Optionally, the input to the prediction network can include the action state, action type, and action attributes of the current frame, and the output is the action state of the next frame. Alternatively, the input to the prediction network can include the action state, action type, action attributes, and environmental geometry at a first time step, and the output includes the action state at a second time step and environmental contact labels. Another possible approach is to include the action state, action type, action attributes, and environmental geometry at a first time step, and the output includes the action state at a second time step, environmental contact labels, and completion progress. Here, the second time step is the time step following the first time step. The specific input and output settings can be configured according to actual needs and are not limited here.
[0205] The loss function in this embodiment is adjusted according to the actual input and output, and is not limited here.
[0206] In this embodiment of the application, training data can be obtained by directly recording character actions, or by following steps 301 to 303 above, or by user input of image or video information. In practical applications, there are other ways to obtain training data, and the specific method of obtaining training data is not limited here.
[0207] The prediction network in this application embodiment can use multilayer perceptron (MLP), long short-term memory (LSTM), graph convolutional networks (GCN), graph neural networks (GNN), transformer, etc., and is not specifically limited here.
[0208] For example, the structure of a prediction network can be as follows Figure 5As shown, the prediction network includes a state encoding network, an environment encoding network, an attribute encoding network, a gating network, an action generation network, and a state decoding network. The state encoding network encodes the action state at the first moment to obtain a state vector. The environment encoding network encodes the environment geometry (or environmental geometric information) to obtain an environment vector. The attribute encoding network encodes action attributes to obtain an attribute vector. The gating network generates mixing coefficients based on the action type. The action generation network includes multiple sets of network parameters (the number is greater than or equal to the number of action types the model is responsible for). The state vector, environment vector, and attribute vector are concatenated to obtain a single vector (called the concatenated vector). This concatenated vector and the mixing coefficients are then input into the action generation network. The network parameters obtained by weighting the multiple sets of network parameters of the action generation network according to the mixing coefficients are used as the current network parameters of the action generation network for inference (which can also be understood as the process of updating the network parameters in the action generation network). The action generation network generates a vector based on the concatenated vector and the mixing coefficients. The output vector of the action generation network is input into the state decoding network, which outputs the motion state, environmental contact, and completion progress at the second moment.
[0209] certainly, Figure 5 The prediction network shown may include more or fewer structures; for example, the prediction network may not include an environment coding network, meaning the input may not include environment geometry. The specific structure of the prediction network is not limited here.
[0210] The different networks mentioned above can have the same or different structures. For example, a network structure could be a multilayer perceptron, which includes three fully connected layers and a nonlinear activation layer between adjacent layers.
[0211] Optionally, the backpropagation algorithm is used to propagate the error from the last layer to the first layer, and the gradient descent algorithm is used to optimize and update the model parameters. The above process is iterated until the model converges, or the error stops decreasing, or the number of iterations reaches a threshold, etc.
[0212] It should be noted that the training process may also use other training methods instead of the aforementioned methods, and no restrictions are imposed here.
[0213] The data processing method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings. The data processing method in the embodiments of this application can be executed by a terminal device or a cloud server, or it can be executed by both a terminal device and a cloud server, which will be described separately below.
[0214] Please see Figure 6This application provides an embodiment of a data processing method, which can be executed by a data processing device (terminal device / cloud server) or by a component of the data processing device (such as a processor, chip, or chip system). This embodiment includes steps 601 to 607.
[0215] The data processing method provided in this application can be applied to the automatic and rapid production of children's educational animations, short video animations, promotional animations, variety show animations, film and television preview animations, etc., or to automatic motion control in games, or to scenes for automatically generating character actions in interactive applications.
[0216] It is understandable that the above scenarios are just examples. In actual applications, there may be other application scenarios, which are not limited here.
[0217] Step 601: Obtain the first information.
[0218] The first information in this embodiment includes a first action type and a first action attribute. The first action type describes a first limb action, and the first action attribute describes the process in which the first limb action occurs.
[0219] Optionally, the first limb action can be a full-body limb action, and the first action type includes walking, running, jumping, sitting down, standing up, squatting down, lying down, hugging, punching, swinging a sword, dancing, etc. The attributes of the first action include target position, displacement, path (i.e., movement path), action speed, frequency of action occurrence, amplitude of action, and direction of action.
[0220] For example, the first action type is "walking", and the first action attributes are "target" and "path". Among them, the limb movement corresponding to "walking" is a full-body limb movement.
[0221] Optionally, the first information may also include the first start time and / or the first end time of the first action type, and including the first start time and the first end time can be understood as including the first time period.
[0222] In this embodiment, the descriptions of action type, action attributes, limb movements, joints, root nodes, etc., are the same as those described above. Figure 3 The corresponding descriptions in the illustrated embodiments are similar and will not be repeated here.
[0223] The first information in the embodiments of this application may be referred to as the first script block.
[0224] There are multiple ways to obtain the first information in this application embodiment, and the first information of one or more roles can be obtained. For ease of description, this application embodiment only describes the acquisition of the first information from the perspective of one role (e.g., Tom).
[0225] The first method is to obtain the first piece of information (also known as the first script block).
[0226] Optionally, the first script block includes a first action type, a first action attribute, and a first start time. Of course, the first script block may also include a first end time and / or a first end point.
[0227] For example, the format of the first script block can be as follows: Figure 7 As shown, the first start time of the first script block is 00:00:02.2; the first end time is 00:00:05.5; the first action type is "walk"; the first action attributes include target and trajectory. The target is Jerry, and the trajectory is "auto". The target being Jerry can be understood as the target location being Jerry's current location.
[0228] The second method involves obtaining the first information through the user's first operation on the first user interface.
[0229] Optionally, the device (terminal device or cloud server) displays a first user interface to the user. This first user interface may include action type options and action attribute options. The user can determine the first action type from the action type options and the first action attribute from the action attribute options through a first operation on the first user interface (e.g., clicking, swiping, dragging, etc.). Of course, the method of determining the first action attribute can also be based on automatically displaying the first action attribute that matches the first action type selected by the user. The specific method of determining the first action attribute is not limited here.
[0230] Optionally, the first user interface is used by the user to edit a first script block. The first user interface may include a script block editing interface, and further, it may also include a script block timeline interface and an animation preview interface. The script block editing interface is used by the user to select the action type and action attributes of the script block, and may further include a start time and an end time.
[0231] For example, taking a first script block that includes a time period (i.e., a start time and an end time) as an example, the first user interface can be as follows: Figure 8As shown, the first user interface includes an animation preview interface (animation not shown), a script block editing interface, and a script timeline interface. The animation preview interface may include a play icon (not shown) and a progress time period (not shown). The script block editing interface includes a script block name area 102, a start time area 103, an end time area 104, a type option 105, and an attribute area 106. The script timeline interface includes a first character area 107, a second character area 108, and a first script block area 101. Area 104 can be automatically displayed based on the type or determined by user operation; this is not limited here. Users can select the script block to edit (e.g., the first script block) by entering the script block name in area 102, or by clicking area 101. For example, clicking area 101 to select the script block to edit... Figure 9 As shown, the user can click on area 101, and the device responds to the user's click, identifying the script block to be edited as the first script block. Since the first script block corresponds to the action type and action attribute, the device can change the displayed "Type" to "Action Type" and "Attribute" to "Action Attribute." Furthermore, as... Figure 10 As shown, users can edit the start and end times through areas 103 and 104. They can also click on an action type. The device responds to the user's click by displaying a drop-down menu containing multiple action type options. The device then determines the first action type based on the user's selection. Taking the user clicking the "Walk" icon 201 as an example, the action attribute area automatically displays the first action attribute matching "Walk" (i.e., "Target" and "Path"). At this point, the editing operation of the first script block (i.e., the first information) is complete, and the device can display... Figure 11 The interface shown has the following settings: the first start time of the first script block is 00:00:02.2; the first end time is 00:00:05.5; the first action type is "walk"; and the first action attributes include target and path. The target can be a default position, character, or item, or it can be determined by the user clicking on other characters, items, or blank spaces in the animation preview interface. Similarly, the path can be a default value or a line drawn by the user in the animation preview interface. Furthermore, the user can slide and drag the cursor 301 in the script timeline interface to adjust the first start time and / or the first end time of the first script block; specific adjustments are not limited here.
[0232] Understandable Figures 8 to 11 These are just a few examples of the first user interface for editing the first information. In practical applications, there may be other forms of user interface, such as those excluding the end time or animation preview interface. Specific details are not limited here.
[0233] The third method is to obtain pre-set first information.
[0234] Optionally, the user can pre-set the first information, which includes the first action type and the first action attributes. Additionally, the first information may also include the first start time and / or the first end time.
[0235] For example, the pre-set first information (or first script block) includes a first start time of 00:00:02.2; a first end time of 00:00:05.5; a first action type of walk; and first action attributes including target and trajectory, with the target being Jerry and the trajectory being auto.
[0236] It is understandable that the above methods for obtaining the first information are just examples. In practical applications, there may be other ways to obtain the first information, which are not limited here.
[0237] Step 602: Obtain the first action state.
[0238] In this application embodiment, the first action state has several possibilities, which are described below:
[0239] The first type, the first action state is the initial motion state.
[0240] Optionally, the first action state can be the initial action state of the first limb action, which includes the initial orientation of the first limb action, the initial rotation angle of the joint corresponding to the first limb action, and the coordinates of the root node in the first limb action. The descriptions of orientation, rotation angle, root node, etc., can be found in the foregoing. Figure 3 The description in step 303 of the illustrated embodiment is not limited here.
[0241] Optionally, if the first action state is the initial motion state, the first action state can be obtained through a second operation of the user on the second user interface. Specifically, the device can display the second user interface to the user, which includes an action state area. The user can click, fill in, or perform other operations on the action state area to determine the first action state.
[0242] Alternatively, the first action state can be obtained by user input of a first script, which is similar to the first method of obtaining the first information described above. For example, this method can be referred to... Figure 12 ,in Figure 12 Including the first action state and the first information mentioned above, the initial position in the first action state is (0,0) and the initial orientation is 0.
[0243] The second type is where the first action state is not the initial action state.
[0244] In this case, the first action state may include at least one of the coordinates, rotation angle, velocity, etc. of the joint corresponding to the first limb action.
[0245] In this case, the first action state can be obtained either by pre-setting it by the user or by inputting the second action state and the first information into the first network. The second action state includes at least one of the following: the second coordinates, rotation angle, velocity, etc., of the joint corresponding to the first limb movement. This can be understood as the first and second action states being the action states corresponding to two adjacent frames in the subsequent target video, or the second action state preceding the first action state (i.e., the second action state is used to represent the historical action state preceding the first action state). In other words, the subsequent action state (i.e., the first action state) is predicted using the previous action state (i.e., the second action state), the action type, and the action attributes.
[0246] It is understandable that the above methods for obtaining the first action state are just examples. In practical applications, there may be other ways to obtain the first action state, which are not limited here.
[0247] Optionally, to acquire animations that interact with the environment, the device can also acquire the scene's environmental geometry information (or environmental information), which includes props, terrain, other characters, etc. Props include trees, tools, or stools, vehicles, etc., that may interact with the characters; specific examples are not limited here. The method of acquiring the scene can be by the user importing a scene model file (e.g., the user creates and imports a scene, or the user selects a scene from a database, a default environment scene, etc.), or by the user building the scene on the device's user interface; specific examples are not limited here.
[0248] For example, the acquired environmental information can be as follows: Figure 13 As shown, the environmental information includes trees, boxes, etc.
[0249] Step 603: Obtain n sets of action states based on the first action state and the first information.
[0250] After obtaining the first action state and the first information, n sets of action states can be obtained based on the first action state and the first information. Each set of action states in the n sets of action states can be used to obtain a limb action. In other words, the first limb action sequence (or the first action segment) can be obtained based on the n sets of action states, where n is a positive integer.
[0251] Optionally, the first action state and the first information can be input into the trained first network to obtain n sets of action states. For details, please refer to... Figure 14 The prediction network shown is related to... Figure 5 The prediction network shown is similar, and the same structure will not be described again here. Figure 14 The prediction network shown, during the inference phase (i.e., in the process of predicting n sets of action states), can input the first action state and the first information into the first network to obtain a set of action states, and then input this set of action states and the first information into the first network to obtain the next set of action states. This process is repeated until n sets of action states are obtained. In other words, the prediction can iterate multiple times until a termination condition is met, at which point the iteration ends, thus obtaining the predicted n sets of action states. The trained first network can be obtained through the aforementioned... Figure 3 The training method shown can be used to train the first limb, but other methods are also possible; specific methods are not limited here. Each of the n sets of action states can include at least one of the following: coordinates, rotation angle, velocity, etc., of the joint corresponding to the first limb action. Additionally, the input to the first network can include environmental geometry information, and the output includes environmental contact information. This environmental contact information indicates whether the joint corresponding to the first limb action is in contact with props or other characters in the environmental geometry information. The first action segment can then be adjusted based on the environmental contact information.
[0252] Optionally, the value of n can also be obtained. Specifically, the termination condition can be obtained, and then the value of n can be determined based on the termination condition. The termination condition of the first action type can include at least one of the following: the completion progress of the first limb action, the termination time of the first action type, and the parameters in the first action attribute. For example, if the first action type is "walk" and the first action attribute includes "target" and "path", then the termination condition can be the first termination time of the first action type, or it can be that the prediction ends when the target is reached.
[0253] In addition, the termination condition can be preset, entered by the user on a third user interface, or predicted by the first network.
[0254] Optionally, if the first information includes a first time period or a first end time, the first end time, the first action state, and the first information can be input into the first network to obtain the action state at the next moment and the completion progress of the action state at the next moment. The completion progress represents the ratio of the time from the initial to the next action state to the time of the first action segment corresponding to the first action type. Generally, the completion progress is represented by 0-1, where 0 represents the start and 1 represents the end of the action segment corresponding to the first action type.
[0255] Optionally, if environmental geometry information has been obtained, the environmental geometry information, the first action state, and the first information can be input into the first network to obtain n sets of action states and n environmental contact information corresponding to the n sets of action states. Each environmental contact information in the n environmental contact information indicates whether the joint of the action state corresponding to the environmental contact information is in contact with props, characters, etc. in the environmental geometry information.
[0256] For example, the trained first network is acquired by training the first network with first training data as input and aiming for a first loss function value to be less than a first threshold. The first training data includes a third action state, a first action type, and attributes related to the first action. The third action state includes at least one of the third coordinate, third angle, and third velocity of the joint corresponding to the first limb action. The first loss function is used to indicate the difference between the fourth action state output by the first network and the first target action state. The fourth action state includes at least one of the fourth coordinate, fourth angle, and fourth velocity of the joint corresponding to the first limb action. The first target action state includes at least one of the first target coordinate, first target angle, and first target velocity. The first target action state and the third action state belong to the same collected action data, or to the action states corresponding to two adjacent frames in the same action video. It is understood that during the training of the first network, at least one of environmental geometric information, a first time period, and a first end time can be added to the input, that is, the input of the training process is consistent with the input of the inference process.
[0257] Step 604: Obtain the target object.
[0258] There are multiple ways to obtain the target object in this application embodiment. It can be to select a 3D model from the database as the target object, or to input the 3D model of the user's role into the device, or to receive the target object sent by other devices. The specific method is not limited here.
[0259] In this embodiment, the target object is related to the first action type and the first action attribute. This can also be understood as redirecting the first action fragment corresponding to the first action type to the target object, thus obtaining the first action fragment of the target object.
[0260] For example, the target object is such as Figure 15 As shown.
[0261] Step 605: Process the limb movements of the target object based on the first action state to obtain the first image.
[0262] After the device acquires the first action state, it can process the target object's limb movements (i.e., the first limb movements) based on the first action state to obtain the first image.
[0263] Optionally, after the device obtains the first action state, it can first determine a first limb action based on the joint information corresponding to the first action state (such as the initial orientation of the first limb action, the initial rotation angle of the joint corresponding to the first limb action, the coordinates of the root node in the first limb action, or at least one of the coordinates, rotation angle, speed, etc. of the joint corresponding to the first limb action), and redirect the first limb action to the target object to obtain the first image.
[0264] For example, if the first action state is the initial action state of the first action type "walking", then a first limb action determined based on the joint information in the initial state can be as follows: Figure 16 As shown, the first image obtained by redirecting the first limb movement to the first limb movement of the target object can be as follows: Figure 17 As shown. It is understandable that... Figure 17 The image shown includes environmental geometry information (i.e. Figure 17 In practice, environmental geometry information may not be required for elements such as trees, green spaces, and boxes.
[0265] Step 606: Process the limb movements of the target object based on n sets of action states to obtain n images.
[0266] After acquiring n sets of action states, the device can process the target object's limb movements (i.e., the first limb movements) based on the n sets of action states to obtain n images.
[0267] Optionally, after acquiring n sets of action states, the device can first determine n first limb actions based on the joint information corresponding to the n sets of action states (e.g., the initial orientation of the first limb action, the initial rotation angle of the joint corresponding to the first limb action, the coordinates of the root node in the first limb action, or at least one of the coordinates, rotation angle, speed, etc. of the joint corresponding to the first limb action), and redirect the n first limb actions to the target object, thereby obtaining n images.
[0268] For example, assuming n is 6, the 6 first limb movements associated with n sets of action states can be referenced. Figure 18 The limb actions corresponding to numbers 2-7 are shown in the table, where limb action number 1 is equivalent to one of the first limb actions associated with the first action state.
[0269] Optionally, after acquiring n sets of environmental contact information corresponding to n sets of action states, the device can process the first limb movement of the target object based on the n sets of action states and the n environmental contact information to obtain n images. This can also be understood as determining n first limb movements based on the n sets of action states and the n environmental contact information, then redirecting these n first limb movements to the target object to obtain n images.
[0270] For example, Figure 19 This can be understood as processing the first limb movement of the target object based on n sets of action states to obtain one of the n images (e.g., corresponding to...). Figure 18 (Image corresponding to the first limb movement of number 5 in the image).
[0271] Step 607: Generate the target video based on the first image and n images.
[0272] After acquiring the first image and n images, the device can generate a target video based on the timing of the predicted n sets of action states, the first image, and the n images. This target video is related to the first action type and the first action attribute.
[0273] For example, if the first action type is walking, then the target video is a video about the target object walking.
[0274] The frame rate can be optionally set according to actual needs, and the target video can be generated based on the frame rate, the generation sequence of the action state, the first image, and n images. The images corresponding to the first script block include: the first image and n images.
[0275] For example, Figure 17 and Figure 19 This can be understood as two specific frames from the target video.
[0276] It is understood that the target video in the embodiments of this application can be understood as an animation about the target object.
[0277] In this embodiment, on the one hand, the higher-level semantics of types and attributes in the script are more efficient, intuitive, and understandable. Compared to the prior art that requires users to input low-level control signals, this reduces the user's workload and improves the efficiency of generating the target video. On the other hand, compared to the prior art that requires users to input control signals for each frame, this application can predict n sets of action states based on the first action state and the first information, which is equivalent to obtaining the action sequence over a period of time based on the script. This reduces the user's operation and technical requirements, improves the user experience, and increases the efficiency of animation generation. Furthermore, the first information can be flexibly adjusted through interactive interfaces, making the generated animation more versatile.
[0278] In one possible implementation, it may further include acquiring second information; acquiring a fifth action state; predicting m sets of action states based on the fifth action state and the second information; processing the target object's limb movements based on the fifth action state to obtain a second image; and processing the target object's limb movements based on the m sets of action states to obtain m images. In this case, step 601 may specifically generate a target video based on the first image, n images, the second image, and m images. These are described below:
[0279] Optionally, the device can also acquire second information, similar to the first information, including a second action type and a second action attribute. The second action type describes the action of the second limb, and the second action attribute describes the process by which the second limb action occurs. The first limb and the second limb can refer to the entire body, and the joints driven by the first limb and the second limb can be the same or different.
[0280] Optionally, the second limb movement can be a full-body movement, and the types of the second movement include walking, running, jumping, sitting down, standing up, squatting down, lying down, hugging, punching, swinging a sword, dancing, etc. The attributes of the second movement include target position, displacement, path (i.e., movement path), movement speed, frequency of movement, amplitude of movement, and direction of movement.
[0281] For example, the second action type is "sit down" and the second action attribute is "target". Among them, the limb action corresponding to "sit down" is a full-body limb action.
[0282] It is understandable that more or fewer action types and action attributes corresponding to full-body limb movements can be obtained. The specifics are not limited here. The following description only uses the example of obtaining the action types (i.e., the first action type and the second action type) and action attributes corresponding to two full-body limb movements.
[0283] Optionally, the first information may also include a first start time and / or a first end time for the first action type, where the first start time and first end time can be understood as including a first time period. If the first information does not include a first end time, it can be understood that after performing the first limb action corresponding to the first action type, the limb action corresponding to the subsequent action type is performed immediately. If the first information includes a first end time, then subsequent operations are performed based on the start time of the information corresponding to the subsequent action type.
[0284] Optionally, the second information is similar to the first information described above. The second information also includes a second start time and / or a second end time for the second action type. Including the second start time and the second end time can be understood as including a second time period. If the second information does not include the second end time, it can be understood that after performing the second limb action corresponding to the second action type, the limb action corresponding to the subsequent action type is performed immediately. If the second information includes the second end time, then subsequent operations are performed according to the start time of the information corresponding to the subsequent action type. The first time period and the second time period may not overlap or may be consecutive. Optionally, the first start time of the first time period is earlier than the second start time of the second time period. Consecutive contiguous ...
[0285] Optionally, the first and second information mentioned above correspond to whole-body limb movements.
[0286] In this embodiment, the first information can be referred to as the first script block, and the second information can be referred to as the second script block. That is, the first script corresponding to a full-body limb movement includes both the first script block and the second script block. Of course, a full-body limb movement can include more script blocks, which is not limited here. In other words, more or fewer action types and action attributes corresponding to full-body limb movements can be obtained. The following description only uses the example of obtaining the action types and action attributes corresponding to two full-body limb movements.
[0287] There are multiple ways to obtain the first script in this application embodiment. The following description uses obtaining the first information and the second information as examples. In this application embodiment, the first script of one or more roles can be obtained. For ease of description, this application embodiment only describes obtaining the first script from the perspective of one role (e.g., Tom).
[0288] The first method is to obtain the second script block.
[0289] Optionally, the second script block includes a second action type and a second action attribute. It is understood that the second script block may also include a second time period (i.e., a second start time and a second end time).
[0290] Of course, the aforementioned methods of obtaining the first script block and obtaining the second script block can be used to obtain the first script block.
[0291] For example, taking a first script that includes two script blocks as an example, the format of the first script can be as follows: Figure 20As shown, the first script block has a first start time of 00:00:02.2 and a first end time of 00:00:05.5; the first action type is "walk," and the first action attributes include a target and a trajectory, with the target being Jerry and the trajectory being "auto." The second script block has a second start time of 00:00:05.5 and a second end time of 00:00:08; the second action type is "seat," and the second action attributes include a target, with the target being a box.
[0292] For example, the format of the first script can also be as follows: Figure 21 As shown, it includes a first action state, a first script block, and a second script block, wherein the first script block and the second script block are related to... Figure 10 The first script block in the code is similar to the second script block, and the first action state is... Figure 12 The first action state is similar, so it will not be described again here.
[0293] The second method involves obtaining the second script block through the user's first operation on the first user interface.
[0294] Optionally, the device (terminal device or cloud server) displays a first user interface to the user. This first user interface may include action type options and action attribute options. The user can determine the first action type from the action type options and the first action attribute from the action attribute options through a first operation on the first user interface (e.g., clicking, swiping, dragging, etc.). Of course, the method of determining the first action attribute can also be based on automatically displaying the first action attribute that matches the first action type selected by the user. The specific method of determining the first action attribute is not limited here.
[0295] Optionally, the first user interface is used by the user to edit the first script. The first user interface may include a script block editing interface, and further, it may also include a script block timeline interface and an animation preview interface. The script block editing interface is used by the user to select the action type and action attributes of the script block, and may further include a start time and an end time.
[0296] For example, the operation of the second script block (i.e., the second information) is similar to that of the first script block, and will not be repeated here. Figure 22 As shown, the second start time of the second script block is 00:00:05.5; the second end time is 00:00:08; the second action type is sit down; and the second action attributes include the target (e.g., a box or the location clicked by the user in the animation preview interface).
[0297] Understandable Figure 22This is just one example of the first user interface for editing the first script. In practical applications, there may be other forms of user interface, such as those that do not include the end time or animation preview interface. No specific restrictions are made here.
[0298] The third method is to obtain pre-set second information.
[0299] Optionally, the user can pre-set second information, which includes a second action type and second action attributes. Additionally, the second information may include a second start time and / or a second end time.
[0300] For example, the second start time of the pre-set second information (or second script block) is 00:00:05.5; the second end time is 00:00:08; the second action type is sit down; and the second action attribute includes target, which is box.
[0301] It is understandable that the above methods of obtaining second information are just examples. In practical applications, there may be other ways to obtain second information, which are not limited here.
[0302] In one possible implementation, a fifth action state corresponding to the second information can also be obtained. If the first end time of the first information is the first start time of the second information, then the fifth action state can be the nth action state in the above n sets of action states, that is, the first action type and the second action type are continuous. Of course, the fifth action state can also be similar to the first action state, and can also include the initial orientation of the second limb action, the initial rotation angle of the joint corresponding to the second limb action, and the coordinates of the root node in the second limb action. Or the fifth action state includes at least one of the coordinates, rotation angle, velocity, etc. of the joint corresponding to the second limb action. It can be understood that, for whole-body limb actions, the joints corresponding to whole-body limb actions can be understood as joints throughout the body.
[0303] Optionally, after obtaining the second information and the fifth action state, m sets of action states can be predicted based on the second information and the fifth action state, where m is a positive integer. Each set of action states in the m sets of action states can be used to obtain a limb action. In other words, a second limb action sequence (or second action segment) can be obtained based on the m sets of action states.
[0304] Optionally, the fifth action state and the second information can be input into the trained second network to obtain m sets of action states. This is similar to the input and output of the first network described above, and will not be repeated here. Each of the m sets of action states can include at least one of the following: the coordinates, rotation angle, velocity, etc., of the joint corresponding to the second limb action. Of course, if the first action attribute in the first information is consistent with the second action attribute in the second information, then the first network and the second network can be the same.
[0305] For example, the trained second network is acquired by training the second network with the second training data as input and aiming for the value of the second loss function to be less than a second threshold. The second training data includes a sixth action state, a third action type, and a third action attribute. The sixth action state includes at least one of the seventh coordinate, seventh angle, and seventh velocity of the joint corresponding to the second limb action. The second loss function is used to indicate the difference between the seventh action state output by the second network and the second target action state. The seventh action state includes at least one of the seventh coordinate, seventh angle, and seventh velocity of the joint corresponding to the second limb action. The second target action state includes at least one of the second target coordinate, second target angle, and second target velocity. The second target action state and the sixth action state belong to the same collected action data, or can be understood as belonging to the action states corresponding to two adjacent frames in the same action video.
[0306] Furthermore, in predicting m sets of action states based on the second information and the fifth action state, the termination condition of the second action type and the environmental geometric information of the second action type can also be introduced. The termination condition of the second action type can be used to determine the number of m. The termination condition of the second action type can include at least one of the following: the completion progress of the second limb action, the end time of the second action type, and parameters in the second action attribute. For example, if the second action type is "sit down" and the second action attribute includes "target", then the termination condition can be the second end time of the second action type, or it can be that the prediction ends when the user sits down at the location of the "target". In addition, this termination condition is either preset, input by the user on the third user interface, or predicted by the second network. The remaining descriptions can refer to the relevant descriptions of the termination conditions of the first action state mentioned above, and will not be repeated here.
[0307] Optionally, after acquiring the fifth action state, the device can process the limb movement of the target object (i.e., the second limb movement) based on the fifth action state to obtain the second image. Specifically, a second limb movement can be determined first based on the joint information of the fifth action state (e.g., the initial orientation of the second limb movement, the initial rotation angle of the joint corresponding to the second limb movement, the coordinates of the root node in the second limb movement, or at least one of the coordinates, rotation angle, velocity, etc. of the joint corresponding to the second limb movement), and the second limb movement can be redirected to the limb movement of the target object to obtain the second image.
[0308] Optionally, after acquiring m sets of action states, the device can process the second limb movements of the target object based on the m sets of action states to obtain m images.
[0309] Similarly, we can obtain m sets of environmental contact information corresponding to m sets of action states. Based on these m sets of action states and the m sets of environmental contact information, we can process the second limb movements of the target object to obtain m images. This can also be understood as determining m second limb movements based on the m sets of action states and the m sets of environmental contact information, then redirecting these m second limb movements to the target object to obtain m images. The method for obtaining these m sets of environmental contact information can refer to the method for obtaining n sets of environmental contact information mentioned earlier, and will not be elaborated here.
[0310] For example, assuming m is 4, the four second limb movements associated with m sets of action states can be referenced. Figure 23 The limb movements corresponding to 2-5 in the image are as follows: the limb movement of number 1 is equivalent to a second limb movement associated with the second action state, or it can be understood as a second limb movement corresponding to the second image.
[0311] Additionally, if n first limb movements and m second limb movements are obtained, then the n first limb movements and m second limb movements can be concatenated to obtain n+m limb movements.
[0312] For example, the limb movements of 2+n+m can be as follows: Figure 24 As shown, the first image corresponds to the first limb movement of number 1, and n images correspond to the first limb movements of numbers 2-7. The second image corresponds to the second limb movement of number 8, and m images correspond to the second limb movements of numbers 9-12.
[0313] For example, Figure 25 This can be understood as processing the second limb movements of the target object based on m sets of action states to obtain one of the m images.
[0314] Of course, if each script includes its own time period, the target video can be generated based on the time period and the images corresponding to each script (the first information mentioned above also includes the first time period of the first action type, and the second information also includes the second time period of the second action type; generating the target video based on the first image, n images, the second image, and m images includes: generating the target video based on the first image, n images, the second image, m images, the first time period, and the second time period, where the first time period corresponds to the first image and n images, and the second time period corresponds to the second image and m images. In this possible implementation, the first image, n images, the second image, and m images are processed based on the first time period of the first action type and the second time period of the second action type to generate the target video. This ensures the continuity of the target video.). Additionally, the frame rate can be set according to actual needs, and the target video can be generated based on the frame rate, time period, and images. The images corresponding to the first script include: the first image, n images, the second image, and m images.
[0315] In this case, step 601 can specifically generate a target video based on the first image, n images, the second image, and m images.
[0316] For example, Figure 25 It can be understood as an image of a specific frame in the target video.
[0317] For example, if the first action type is walking and the second action type is sitting, then the target video is a video about the target object walking and sitting.
[0318] Furthermore, the first script can include one or more script blocks. If there is one, the first script includes first information; if there are two, the first script includes both first and second information. Of course, it can also include more information (including the action type and action attributes corresponding to whole-body limb movements).
[0319] In another possible implementation, it may also include acquiring third information; acquiring an eighth action state; and predicting p groups of action states based on the eighth action state and the third information. In this case, step 605 may specifically be based on the coverage relationship between the third limb and the first limb, using the limb action of the target object corresponding to the eighth action state to cover the limb action of the target object corresponding to the first action state to obtain a first image; step 606 may specifically be based on the coverage relationship between the third limb and the first limb, using the limb action of the target object corresponding to p groups of action states to cover the limb action of the target object corresponding to n groups of action states to obtain n images. These are described below:
[0320] Optionally, third information corresponding to the first partial limb can also be obtained. This third information includes a third action type and a third action attribute. The third action type describes the action of the third limb, and the third action attribute describes the process by which the action of the third limb occurs. The third limb can be either the first limb or the first partial limb within the second limb.
[0321] Optionally, the third limb movement is a partial limb movement, and the types of third movements include raising the head, kicking the leg, waving the hand, wagging the tail, etc. The attributes of the third movement include target position, displacement, path (i.e., movement path), movement speed, frequency of movement, amplitude of movement, and direction of movement.
[0322] For example, the third action type is "staring", and the third action attribute is "target". Among them, the limb action corresponding to "staring" is a partial limb action, such as mainly the head.
[0323] Of course, it is also possible to obtain fourth information corresponding to the second partial limb. This fourth information includes a fourth action type and a fourth action attribute. The fourth action type describes the action of the fourth limb, and the fourth action attribute describes the process by which the action of the fourth limb occurs. The fourth limb can be the second partial limb within either the first or second limb.
[0324] Optionally, the fourth limb movement is a partial limb movement, and the types of the fourth movement include raising the head, kicking the leg, waving the hand, wagging the tail, etc. The attributes of the fourth movement include target position, displacement, path (i.e., movement path), movement speed, frequency of movement, amplitude of movement, and direction of movement.
[0325] For example, the fourth action type is "waving", and the fourth action attribute is "amplitude". Among them, the limb action corresponding to "waving" is a partial limb action, such as mainly the arm.
[0326] Optionally, the third and fourth information correspond to local limb movements.
[0327] Optionally, the third information also includes the third start time and / or the third end time of the third action type, which can be understood as including the third time period. Similarly, the fourth information is similar to the third information, and also includes the fourth start time and / or the fourth end time of the fourth action type, which can be understood as including the fourth time period. The third and fourth time periods may overlap or not, and this is not specifically limited here.
[0328] In this embodiment, the third information can be referred to as the third script block, and the fourth information can be referred to as the fourth script block. That is, the second script corresponding to the local limb movement includes the third script block and the fourth script block. Of course, the local limb movement can include more script blocks, which is not limited here. In other words, more or fewer action types and action attributes corresponding to local limb movements can be obtained. The following description only uses the example of obtaining the action types and action attributes corresponding to two local limb movements.
[0329] There are multiple ways to obtain the second script in this application embodiment. The following description uses obtaining the third and fourth information as examples. In this application embodiment, the second script of one or more roles can be obtained. For ease of description, this application embodiment only describes obtaining the second script from the perspective of one role (e.g., Tom).
[0330] The first method is to obtain the second script.
[0331] Optionally, the method for obtaining the third and fourth information corresponding to the local limb movement can be similar to the method for obtaining the first script. That is, obtaining the second script, which includes the third information (also called the third script block) and the fourth information (also called the fourth script block). The third script block includes the third action type, the third action attribute, and the third start time. Of course, the third script block can also include the third end time. The fourth script block includes the fourth action type and the fourth action attribute. It is understood that the third script block can also include the third end time, and further, the fourth script block can also include a fourth time period (i.e., the fourth start time and the fourth end time).
[0332] For example, taking a second script that includes two script blocks as an example, the format of the second script can be as follows: Figure 26 As shown, the third script block has a third start time of 00:00:0.1 and a third end time of 00:00:05.5; the third action type is "look at," and the third action attribute includes a target, which is Jerry's face. The fourth script block has a fourth start time of 00:00:03 and a fourth end time of 00:00:05; the fourth action type is "wavehand," and the fourth action attribute includes a periodic amplitude of 0.5.
[0333] The second method involves obtaining the second script through the user's first operation on the first user interface.
[0334] Optionally, the method for obtaining the second script (i.e., the third and fourth information) corresponding to the local limb movement can be similar to the method for obtaining the first script (i.e., the first and second information). The device (terminal device or cloud server) displays a first user interface to the user. This first user interface may include action type options and action attribute options. The user can determine the third action type from the action type options and the third action attribute from the action attribute options by operating the first user interface. Of course, the method for determining the third action attribute can also be based on the user's selected third action type, automatically displaying the third action attribute that matches the third action type. The specific method for determining the third action attribute is not limited here.
[0335] Optionally, the first user interface is used by the user to edit the second script. The first user interface may include a script block editing interface, and further, it may include a script block timeline interface and an animation preview interface. The script block editing interface is used by the user to select the action type and action attributes of the script block, and may further include a start time and an end time.
[0336] For example, taking a second script that includes two script blocks and where both script blocks include time periods (i.e., start and end times) as an example, the first user interface can be as follows: Figure 8 As shown, the first user interface includes an animation preview interface (animation not shown), a script block editing interface, and a script timeline interface. The animation preview interface may include a play icon (not shown) and a progress time period (not shown). The script block editing interface includes a script block name area 102, a start time area 103, an end time area 104, a type option 105, and an attribute area 106. The script timeline interface includes a first character area 107, a second character area 108, and a first script block area 101. Area 104 can be automatically displayed based on the type or determined by user operation; this is not limited here. Users can select the script block to edit (e.g., the first script block) by entering the script block name in area 102, or by clicking area 101. For example, clicking area 101 to select the script block to edit... Figure 27 As shown, the user can click on area 109. The device responds to the user's click, determining whether the script block to be edited belongs to a secondary action script (i.e., the second script) or a fourth script block. Since the second script corresponds to action type and action attribute, the device can change the displayed "Type" to "Action Type" and "Attribute" to "Action Attribute". Furthermore, the user can edit the start time through areas 103 and 104, such as... Figure 28As shown, the user can also click on the action type. The device responds to the user's click by displaying a drop-down menu containing multiple action type options. The fourth action type is then determined based on the user's selection. Here, we take the user clicking the "wave" icon 401 as an example. At this point, the editing operation of the fourth script block is complete, and the device can display as shown... Figure 29 The interface shown has the fourth script block's start time set to 00:00:03 and end time set to 00:00:05. The fourth action type is a wave, and its attributes include amplitude, which can be a default value or user-defined. Here, an amplitude of 0.5 is used as an example. Furthermore, users can slide across the script timeline interface, similar to... Figure 11 Drag the cursor 301 in the middle to adjust the fourth start time and / or the fourth end time of the fourth script block; the specifics are not limited here.
[0337] Similarly, the operations of the third script block are similar to those of the fourth script block, and will not be repeated here. Figure 30 As shown, the third start time of the third script block is 00:00:05.5; the third end time is 00:00:08; the third action type is "head up" (which can also be understood as "look"), and the third action attribute includes amplitude. In other words, the third limb corresponding to the third limb action (i.e., the first local limb) can be understood as the head (the corresponding joints include the head and neck joints), and the fourth limb corresponding to the fourth limb action (i.e., the second local limb) can be understood as the arm (the corresponding joints can include the shoulder joint, elbow joint, and wrist joint).
[0338] Understandable Figures 27 to 30 These are just a few examples of the first user interface for editing the second script. In practical applications, there may be other forms of user interface, such as those excluding the end time or animation preview interface. No specific restrictions are made here.
[0339] The third method is to obtain a pre-set second script.
[0340] Optionally, the user can pre-set third and fourth information. The third information includes a third action type and a third action attribute. Additionally, the third information may include a third start time and / or a third end time. The fourth information includes a fourth action type and a fourth action attribute. Additionally, the fourth information may include a fourth start time and / or a fourth end time.
[0341] For example, the pre-set third information (or third script block) includes a third start time of 00:00:05.5; a third end time of 00:00:08; and a third action type of "looking up" (which can also be understood as "looking"). The fourth information (or fourth script block) includes a fourth start time of 00:00:03; a fourth end time of 00:00:05; a fourth action type of "waving," and the fourth action attribute includes amplitude.
[0342] It is understood that the above-mentioned methods for obtaining third and / or fourth information are just examples. In practical applications, there may be other ways to obtain third and / or fourth information, which are not limited here.
[0343] Optionally, an eighth action state corresponding to the third information and a ninth action state corresponding to the fourth information can also be obtained. The eighth action state may include at least one of the coordinates, rotation angle, velocity, etc. of the joint corresponding to the third limb action (or the first local limb action). The ninth action state may include at least one of the coordinates, rotation angle, velocity, etc. of the joint corresponding to the fourth limb action (or the second local limb action).
[0344] Optionally, after obtaining the third information and the eighth action state, p groups of action states can be predicted based on the third information and the eighth action state, where p is a positive integer. Each group of action states in the p groups can be used to obtain a local limb action. In other words, the third limb action sequence (or the first local limb action sequence) can be obtained based on the p groups of action states.
[0345] Optionally, after obtaining the fourth information and the ninth action state, q sets of action states can be predicted based on the fourth information and the ninth action state, where q is a positive integer. Each set of action states in the q sets of action states can be used to obtain a local limb action. In other words, the fourth limb action sequence (or the second local limb action sequence) can be obtained based on the q sets of action states.
[0346] Optionally, the eighth action state and the third information can be input into the trained third network to obtain p sets of action states. This is similar to the input and output of the first network mentioned above, and will not be repeated here. Each of the p sets of action states can include at least one of the coordinates, rotation angle, velocity, etc. of the joint corresponding to the third limb action (first local limb).
[0347] Optionally, the ninth action state and the fourth information can be input into the trained fourth network to obtain q sets of action states. This is similar to the input and output of the first network mentioned above, and will not be repeated here. Each action state in the p sets of action states can include at least one of the coordinates, rotation angle, velocity, etc., of the joint corresponding to the third limb action (first local limb). The training methods for the third and fourth networks can be as described above. Figure 3 The training methods shown are similar and will not be repeated here.
[0348] For example, if the third limb movement is the head, then the joints corresponding to the head may include Figure 4 The head and neck joints. If the fourth limb movement is the arm, then the corresponding joints of the arm can include... Figure 4 The shoulder, elbow, and wrist joints.
[0349] Furthermore, in predicting the action states of group p based on the third information and the eighth action state, the termination condition of the third action type and the environmental geometric information of the third action type can also be introduced. The termination condition of the third action type can be used to determine the number of p. The termination condition of the third action type can include at least one of the following: the completion progress of the third limb action, the end time of the third action type, and the parameters in the third action attribute. Similarly, in predicting the action states of group q based on the fourth information and the ninth action state, the termination condition of the fourth action type and the environmental geometric information of the fourth action type can also be introduced. The termination condition of the fourth action type can be used to determine the number of q. The termination condition of the fourth action type can include at least one of the following: the completion progress of the fourth limb action, the end time of the fourth action type, and the parameters in the fourth action attribute.
[0350] Optionally, after the device obtains the eighth action state, it can first determine the first local limb action based on the joint information corresponding to the eighth action state, and then, based on the coverage relationship between the first local limb and the first limb, use the first local limb action to cover the first limb action, and then redirect the covered first limb action to the target object to obtain the first image.
[0351] It is understandable that if the first information also includes a first time period corresponding to the first action type, and the third information also includes a third time period corresponding to the third action type, then based on the coverage relationship between the first local limb and the first limb, and the coverage relationship between the third time period and the first time period, after covering the first limb action with the first local limb action, the covered first limb action is redirected to the target object, thereby obtaining the first image.
[0352] Optionally, after the device obtains the ninth action state, it can first determine the second local limb action based on the joint information corresponding to the ninth action state, and then, based on the coverage relationship between the second local limb and the first limb, use the second local limb action to cover the first limb action, and then redirect the covered first limb action to the target object to obtain the first image.
[0353] It is understandable that if the first information also includes a first time period corresponding to the first action type, and the fourth information also includes a fourth time period corresponding to the fourth action type, then based on the coverage relationship between the second local limb and the first limb and the coverage relationship between the fourth time period and the first time period, after the second local limb action is used to cover the first limb action, the covered first limb action is redirected to the target object, thereby obtaining the first image.
[0354] Optionally, if the first local limb actions corresponding to p groups of action states have been obtained, and if the first local limbs have an overlapping relationship with the first limbs, then based on the overlapping relationship between the first local limbs and the first limbs, the first local limb actions of the target object corresponding to p groups of action states are used to cover the first limb actions of the target object corresponding to the first action state to obtain n images. If the first local limbs have an overlapping relationship with the second limbs, then based on the overlapping relationship between the first local limbs and the second limbs, the first local limb actions of the target object corresponding to p groups of action states are used to cover the second limb actions of the target object corresponding to the fifth action state to obtain m images. Generally, p is less than n.
[0355] Optionally, if the second local limb actions corresponding to q sets of action states have been obtained, and if the first local limbs have an overlapping relationship with the first limb, then based on the overlapping relationship between the second local limbs and the first limbs, the second local limb actions of the target object corresponding to the q sets of action states can be used to cover the first limb actions of the target object corresponding to the first action state to obtain n images. If the first local limbs have an overlapping relationship with the second limbs, then based on the overlapping relationship between the second local limbs and the second limbs, the second local limb actions of the target object corresponding to the q sets of action states can be used to cover the second limb actions of the target object corresponding to the fifth action state to obtain m images. Generally, q is less than n.
[0356] For example, assuming q is 4, the four second local limb movements associated with the q group of action states can be referenced. Figure 31 The limb actions corresponding to 2-5 in the diagram are as follows: The second partial limb action of action 1 is equivalent to one second partial limb action associated with the ninth action state. Therefore, using the second partial limb action to overwrite the first limb action to obtain the updated first limb action can be done as follows: Figure 32 As shown. From Figure 32As can be seen, the arm movement in the first limb movement has been replaced by the movement of the second limb.
[0357] In addition, if n first limb actions, m second limb actions, and q second local limb actions are obtained, the n first limb actions and m second limb actions can be concatenated to obtain n+m limb actions, and the n+m limb actions can be updated according to the coverage relationship between the second local limbs and the first / second limbs.
[0358] For example, the updated limb movements of 2+n+m can be as follows: Figure 33 As shown, the first image corresponds to the first limb movement number 1, and n images correspond to the first limb movements numbered 2-7, where the first limb movements numbered 3-7 are limb movements covered by the second local limb movements. The second image corresponds to the second limb movement number 8, and m images correspond to the second limb movements numbered 9-12.
[0359] For example, Figure 34 This can be understood as using the second local limb movements of the target object corresponding to q sets of action states to cover the first limb movements of the target object corresponding to n sets of action states, and then processing them to obtain one of the n images.
[0360] For example, by splicing whole-body limb movements and partial limb movements according to their respective time periods and / or coverage relationships, the above-mentioned results can be obtained. Figure 33 The action sequence shown.
[0361] In the above method, the n images acquired by the device contain one or more full-body limb movements and one or more partial limb movements, facilitating the subsequent generation of complex character animations (i.e., videos about the target object). For ease of understanding, the process of obtaining limb movements for retargeting based on the first and second scripts is simplified: the first script can be understood as generating full-body movements, and the second script as generating partial limb movements. The partial limb movements generated by the second script sequentially overwrite the corresponding body parts' movements within the corresponding time periods of the full-body movements generated by the first script. Of course, if there is no second script, the full-body movements corresponding to the first script are determined as the movements of the target object for retargeting.
[0362] For example, if the first action type is walking and the fourth action type is waving, then the target video is a video of the target object walking and waving.
[0363] The second script can include one or more script blocks. If it has one block, the second script includes third information; if it has two blocks, the second script includes both third and fourth information. It can also include more information (including the action type and action attributes corresponding to the local limb movements).
[0364] Of course, if each script includes its own time period, the target video can be generated based on the time period and images corresponding to each script. Additionally, the frame rate can be set according to actual needs, and the target video can be generated based on the frame rate, time period, and images. Specifically, the images corresponding to the first script include: a first image, n images, a second image, and m images. The images corresponding to the first and second scripts include: the first image updated using local limb movements, the n images updated using local limb movements, the second image updated using local limb movements, and the m images updated using local limb movements.
[0365] Optionally, if the n images are n images updated based on the first local limb movement, then the target video is associated with the first action type, the first action attribute, the third action type, and the third action attribute. If the n images are n images updated based on the second local limb movement, then the target video is associated with the first action type, the first action attribute, the fourth action type, and the fourth action attribute. If the n images are n images updated based on both the first and second local limb movements, then the target video is associated with the first action type, the first action attribute, the third action type, the third action attribute, the fourth action type, and the fourth action attribute.
[0366] In another possible implementation, optionally, it may also include acquiring facial information; acquiring a facial expression sequence based on the first facial information and a first association relationship; in this case, step 605 may specifically be processing the limb movements and facial movements of the target object based on the first action state and the facial expression sequence to obtain a first image. Step 606 may specifically be processing the limb movements and facial movements of the target object based on n sets of action states and facial expression sequences to obtain n images. These are described below:
[0367] It is understandable that the limb and facial movements of the target object can be processed based on m sets of action states and facial expression sequences to obtain m images (i.e., update m images with the facial expression sequence). The following description only takes updating n images with facial expression sequences as an example.
[0368] Optionally, the device can also acquire facial information (also known as a third script). The number of facial information entries can be one or more; this embodiment only describes an example with two facial information entries. Facial information can include facial expression types and corresponding expression attributes. The facial expression type describes the facial movements of the target object, and the expression attribute describes the amplitude of the facial movements. In this embodiment, the facial expression types (e.g., a first facial expression type and a second facial expression type) can include neutral, happy, sad, surprised, angry, disgusted, fearful, delighted, tired, embarrassed, contemptuous, etc. Expression attributes (e.g., a first expression attribute and a second expression attribute) represent the level, amplitude, or range of the facial expression, and the expression attribute can be represented by 0-1 to reflect the degree of expression of the aforementioned facial expressions. This facial information is used to acquire the facial expression sequence of the target object.
[0369] The methods for obtaining facial information are similar to those for obtaining the first information, and can involve various situations, which are described below:
[0370] The first method is to obtain a third-party script.
[0371] Optionally, the third script includes first facial information and second facial information. The first facial information may include a first facial expression type and a corresponding first expression attribute. The first facial expression type describes the first facial movement of the target object, and the expression attribute describes the amplitude of the first facial movement. The second facial information may include a second facial expression type and a corresponding second expression attribute. The second facial expression type describes the second facial movement of the target object, and the expression attribute describes the amplitude of the second facial movement. It is understood that the first facial information may further include a fifth time period (a fifth start time and a fifth end time), and the second facial information may further include a sixth time period (a sixth start time and a sixth end time).
[0372] For example, taking a third script that includes two facial information entries as an example, the format of the third script can be as follows: Figure 35 As shown, the fifth start time of the first facial information is 00:00:0.5; the fifth end time is 00:00:02; the first facial expression type is surprised, and the first expression attribute includes a level of 0.8. The sixth start time of the second facial information is 00:00:05.5; the sixth end time is 00:00:08; the second facial expression type is a smile, and the second expression attribute includes a level of 0.6.
[0373] The second method involves obtaining a third-party script through user interactions with the user interface.
[0374] Optionally, the device (terminal device or cloud server) displays a user interface to the user. This user interface may include expression type options and expression attribute options. The user can determine a first facial expression type from the expression type options and a first expression attribute from the expression attribute options by operating the user interface. Of course, the method of determining the first expression attribute may also be based on the first facial expression type selected by the user, automatically displaying the first expression attribute that matches the first facial expression type. The specific method of determining the first expression attribute is not limited here.
[0375] Optionally, this user interface is used for users to edit third-party scripts. This user interface may include a script block editing interface, and further may include a script block timeline interface and an animation preview interface. The script block editing interface is used by users to select the action type and action attributes of the script block, and may further include a start time and an end time.
[0376] For example, taking a third script that includes two script blocks (i.e., first facial information and second facial information) and where both script blocks include time periods (i.e., start time and end time), the user interface can be as described above. Figure 8 As shown, Figure 8 For relevant descriptions, please refer to the aforementioned descriptions of the first or second script; they will not be repeated here. For example... Figure 36 As shown, the user can click on area 501, and the device responds to the user's click, determining that the script block to be edited belongs to the third script (i.e., the first facial information). Since the third script corresponds to facial expression type and expression attributes, the device can change the displayed "Type" to "Expression Type" and "Attributes" to "Expression Attributes". Furthermore, as... Figure 37 As shown, users can edit the fifth time period using the start and end time areas. They can also click on the expression type. The device responds to the user's click by displaying a drop-down menu containing multiple facial expression type options. The first facial expression type is then determined based on the user's selection; for example, clicking the "surprised" icon 601. At this point, the editing of the first facial information is complete, and the device can display the following: Figure 38 The interface shown has the fifth start time of the first facial information set to 00:00:0.5 and the fifth end time set to 00:00:02. The first facial expression type is surprise, and the first expression attribute is 0.8. Furthermore, users can slide and drag the cursor in the script timeline interface to adjust the fifth start time and / or the fifth end time of the first facial information; specific adjustments are not limited here.
[0377] Similarly, the operation of the second facial information is similar to that of the first facial information, and will not be repeated here. Figure 39As shown, the sixth start time of the second facial information is 00:00:02; the sixth end time is 00:00:08; the second facial expression type is smiling, and the second expression attribute is 0.6.
[0378] Understandable Figures 36 to 39 These are just a few examples of user interfaces for editing third-party scripts. In practical applications, there may be other forms of user interfaces, such as those excluding the end time or animation preview interface. No specific limitations are made here.
[0379] The two methods for obtaining the third script in the embodiments of this application are just examples. In practical applications, there may be other ways to obtain the third script, which are not limited here.
[0380] Furthermore, after acquiring facial information (i.e., the third script), the device can determine a facial expression sequence based on the facial information and a first association relationship. This first association relationship represents the correlation between the facial information and the facial expression sequence. This first association relationship can be understood as an expression dictionary. This expression dictionary includes multiple levels of expression bases. For example, multiple expression bases might include a level 5 smile over 120 frames and a level 8 surprise over 60 frames. The aforementioned first association relationship can be constructed by the device, or it can be obtained by the device from a database or by receiving data from other devices; the specific method is not limited here.
[0381] Optionally, if the first association is understood as a facial expression model, then the input of the model includes expression type, level, and duration. The output of the model includes a matrix of expression segments (i.e., expression sequences).
[0382] Optionally, a first expression fragment is obtained based on the first facial information, a second expression fragment is obtained based on the second facial information, and then the first and second expression fragments are concatenated to obtain an expression sequence. Further, if the first and second facial information are not sequentially continuous and there are idle time periods, the expressions within the idle time periods can be set as neutral expressions. When concatenating multiple expression fragments, an interpolation transition method is used to obtain the expression sequence. Additionally, a blendshape sequence corresponding to periodic blinking can be added to facial expressions other than closed-eye or wide-eyed expressions to update the expression sequence.
[0383] Optionally, after acquiring the facial expression sequence, the device can process the target object's limb movements (i.e., first limb movements) and facial movements based on the first action state and the facial expression sequence to obtain a first image. Specifically, it can first determine one first limb movement and one facial movement based on the joint information of the first action state and the facial expression sequence, and then redirect the one first limb movement and one facial movement to the target object's limb movements and facial movements to obtain the first image. Similarly, after acquiring the facial expression sequence, the device can process the target object's limb movements (i.e., second limb movements) and facial movements based on the fifth action state and the facial expression sequence to obtain a second image.
[0384] Optionally, after acquiring a facial expression sequence (which may include the aforementioned first expression segment and / or second expression segment), the device can process the target object's limb movements (i.e., first limb movements) and facial movements based on n sets of action states and the facial expression sequence to obtain n images. Specifically, it can first determine n first limb movements and facial movements based on the joint information of the n sets of action states and the facial expression sequence, and then redirect the n first limb movements and facial movements to the target object's limb movements and facial movements, thereby obtaining n images. Similarly, after acquiring a facial expression sequence, the device can process the target object's limb movements (i.e., second limb movements) and facial movements based on m sets of action states and the facial expression sequence to obtain m images.
[0385] Optionally, after acquiring the second image and m images, the device can generate a target video based on the time sequence of the predicted n sets of action states, the first image and n images, the time sequence of the predicted m sets of action states, and the second image and m images. This target video is related to the first action type, the first action attribute, the second action type, and the second action attribute.
[0386] The third script can include one or more facial information items. If it includes one, the third script includes either the first or second facial information item; if it includes two, the third script includes both the first and second facial information items. Of course, it can also include more facial information (including the facial expression type and expression attributes corresponding to the target object).
[0387] Of course, if each script includes its own time period, the target video can be generated based on the time period and images corresponding to each script. Additionally, the frame rate can be set according to actual needs, and the target video can be generated based on the frame rate, time period, and images. Specifically, the images corresponding to the first script include: a first image, n images, a second image, and m images. The images corresponding to the first and second scripts include: the first image updated using local limb movements, the n images updated using local limb movements, the second image updated using local limb movements, and the m images updated using local limb movements. The images corresponding to the first and third scripts include: the first image updated using facial expression sequences (e.g., including a first expression fragment and / or a second expression fragment), the n images updated using facial expression sequences, the second image updated using facial expression sequences, and the m images updated using facial expression sequences.
[0388] Optionally, if the n images are n images updated based on the first expression segment, then the target video is associated with the first action type, the first action attribute, the first facial expression type, and the first expression attribute. If the n images are n images updated based on the second expression segment, then the target video is associated with the first action type, the first action attribute, the second facial expression type, and the second expression attribute. If the n images are n images updated based on both the first and second expression segments, then the target video is associated with the first action type, the first action attribute, the first facial expression type, the first expression attribute, the second facial expression type, and the second expression attribute.
[0389] In another possible implementation, optionally, it may also include acquiring text information; generating speech segments based on the text information; generating lip-sync sequences based on the speech segments; in this case, step 605 may specifically be processing the limb and facial movements of the target object based on the first action state and the lip-sync sequence to obtain a first image. Step 606 may specifically be processing the limb and facial movements of the target object based on n sets of action states and lip-sync sequences to obtain n images; step 607 may specifically be generating a target video based on the first image, the n images, and the speech segments. These are described below:
[0390] It is understandable that the limb and facial movements of the target object can be processed based on m sets of action states and lip-sync sequences to obtain m images (i.e., m images are updated using lip-sync sequences). The following description only uses the example of updating n images using lip-sync sequences.
[0391] Optionally, the device can also acquire text information (also known as a fourth script). The amount of text information can be one or more; this embodiment only describes an example where the text information includes two text messages. The text information can include dialogue and the corresponding tone of voice. The tone of voice can be understood as an attribute of the dialogue, and can be set according to actual needs. For example, the tone of voice can include statements, exclamations, questions, etc. For example, the tone of voice can also correspond to the aforementioned facial expressions, including surprise, happiness, etc. This text information is used to acquire the corresponding speech segments and the lip-sync sequence of the target object.
[0392] The methods for obtaining text information are similar to those for obtaining the first information, and can involve several scenarios, which are described below:
[0393] The first method is to obtain the fourth script.
[0394] Optionally, the fourth script includes first text information and second text information. The first text information may include a first line of dialogue and a corresponding first tone of voice. The second text information may include a second line of dialogue and a corresponding second tone of voice. It is understood that the first text information may also include a seventh start time and a seventh end time, and the second text information may also include an eighth start time and an eighth end time. Of course, the seventh end time may also be determined based on the duration of the first speech segment corresponding to the first line of dialogue, and the eighth end time may also be determined based on the duration of the second speech segment corresponding to the second line of dialogue.
[0395] For example, taking a fourth script that includes two text messages as an example, the format of the fourth script can be as follows: Figure 40 As shown, the role in the configuration information is "Tom"; the seventh start time of the first text message is 00:00:01; the first line is: "Jerry, what are you doing here?", with the first tone being astonished. The eighth start time of the second text message is 00:00:06; the second line is: "It's really nice to see you again!", with the second tone being delighted.
[0396] The second method involves obtaining the fourth script through user interactions with the user interface.
[0397] Optionally, the device (terminal device or cloud server) displays a user interface to the user. This user interface may include a dialogue editing area and a tone editing area (e.g., a blank area or tone options). The user can perform operations on the user interface (e.g., filling in, clicking, etc.), that is, entering the first line in the dialogue editing area and entering the first tone corresponding to the first line in the tone editing area. Of course, the first tone can also be determined by the user selecting from the tone options; the specific method is not limited here.
[0398] For example, taking a fourth script that includes two pieces of text information and where both pieces of text information include a start time, the user interface can be as follows: Figure 8 As shown, Figure 8 For relevant descriptions, please refer to the aforementioned descriptions of the first or second script; they will not be repeated here. For example... Figure 41 As shown, the user can click on area 701. The device responds to the user's click, determining whether the script block to be edited belongs to the fourth script (i.e., the dialogue script) or the first text information. Since the fourth script corresponds to dialogue and tone, the device can change the displayed "Type" and "Attribute" to "Dialogue" and "Tone". Furthermore, the user can edit the start time and enter the first dialogue, "Jiri, what are you doing there?" and the first tone, "Surprised," in the script block editing interface. At this point, the editing operation of the first text information is complete, and the device can display as shown... Figure 42 The interface shown has the seventh start time of the first text information set to 00:00:01; the first line of dialogue is: "Jerry, what are you doing here?"; and the first tone is astonished. Furthermore, users can slide and drag the cursor in the script timeline interface to adjust the seventh start time of the first text information; specific adjustments are not limited here.
[0399] Similarly, the operation of the second text information is similar to that of the first text information, and will not be repeated here. Figure 43 As shown, the eighth start time of the second text information is 00:00:06; the second line is: It's really nice to see you again!, and the second tone is "delighted".
[0400] Understandable Figures 41 to 43 These are just a few examples of user interfaces for editing the fourth script. In practical applications, there may be other forms of user interfaces, such as those excluding animation preview interfaces, etc., which are not limited here.
[0401] The two methods for obtaining the fourth script in this application embodiment are just examples. In practical applications, there may be other methods to obtain the fourth script, which are not limited here.
[0402] In addition, after acquiring text information (i.e., the fourth script), the device can generate speech segments based on the text information, and then generate lip-sync sequences based on the speech segments. The lip-sync sequences are used to describe the lip movements of the target object. Optionally, dialogue from the text information can be input into the speech generation model to obtain speech segments.
[0403] For example, continuing the above example, if the fourth script includes first text information and second text information, then the audio segments of the fourth script can include a first audio segment and a second audio segment. The first audio segment is the audio segment corresponding to "Jerry, what are you doing here?", and the second audio segment is the audio segment corresponding to "It's really nice to see you again!".
[0404] Optionally, after acquiring a speech segment, the device can generate one or more blendshape sequences based on one or more speech segments.
[0405] Optionally, speech segments can be input into a lip-sync synthesis model to generate a blendshape sequence about the lip shapes.
[0406] For example, continuing the above example, if the fourth script includes first text information and second text information, then the blendshape sequence of the fourth script may include a first blendshape sequence and a second blendshape sequence. The first blendshape sequence corresponds to the lip movements of the first line, and the second blendshape sequence corresponds to the lip movements of the second line.
[0407] Optionally, after acquiring the lip-sync sequence, the device can process the target object's limb movements (i.e., first limb movements) and lip movements (or mouth movements) based on the first action state and the lip-sync sequence to obtain a first image. Specifically, it can first determine one first limb movement and one lip movement based on the joint information and lip-sync sequence of the first action state, and then redirect this first limb movement and lip movement to the target object's limb movements and lip movements to obtain the first image. Similarly, after acquiring the lip-sync sequence, the device can process the target object's limb movements (i.e., second limb movements) and lip movements based on the fifth action state and the lip-sync sequence to obtain a second image.
[0408] Understandably, when acquiring a lip-sync sequence, a facial expression sequence is usually also acquired. That is, the first image can be obtained by processing the target object's limb movements (i.e., the first limb movement), facial movements, and lip-sync movements based on the first action state, facial expression sequence, and lip-sync sequence. Specifically, a weighted average can be performed using the lip-sync sequence and the facial expression sequence corresponding to that time period, with interpolation used at the transition points to finally output a blendshape sequence of facial movements (including the face and mouth).
[0409] Optionally, after acquiring the lip-sync sequence, the device can process the target object's limb movements (i.e., first limb movements) and lip-sync movements (or mouth movements) based on n sets of action states and the lip-sync sequence to obtain n images. Specifically, n first limb movements and lip-sync movements can be determined based on the joint information and lip-sync sequence of the n sets of action states, and then redirected to the target object's limb movements and lip-sync movements to obtain n images. Similarly, after acquiring the lip-sync sequence, the device can process the target object's limb movements (i.e., second limb movements) and lip-sync movements based on m sets of action states and the lip-sync sequence to obtain m images.
[0410] It is understandable that when obtaining lip-sync sequences, facial expression sequences are also usually obtained. That is, based on n sets of action states, facial expression sequences, and lip-sync sequences, n images can be obtained by processing the target object's limb movements (i.e., first limb movements), facial movements, and lip-sync movements.
[0411] Furthermore, the device can play the target video. Of course, if a fourth script is included, meaning the audio clip was previously acquired, the device can play the audio clip while playing the target video. Alternatively, it can be understood that the target video can be obtained by first synthesizing the image and audio clips using image and audio tracks. This generates animations corresponding to actions, expressions, and lip movements, suitable for high-quality animation production scenarios.
[0412] The fourth script can include one or more pieces of text information. If it contains only one piece of text information, the fourth script includes either the first or second piece of text information. If it contains two pieces of text information, the fourth script includes both the first and second pieces of text information. Of course, it can also include more text information (including the dialogue corresponding to the target object and the tone of voice corresponding to the dialogue).
[0413] Of course, if each script includes its own time period, the target video can be generated based on the time period and images corresponding to each script. Additionally, the frame rate can be set according to actual needs, and the target video can be generated based on the frame rate, time period, and images. Specifically, the images corresponding to the first script include: a first image, n images, a second image, and m images. The images corresponding to the first and second scripts include: the first image updated using local body movements, n images updated using local body movements, the second image updated using local body movements, and m images updated using local body movements. The images corresponding to the first and third scripts include: the first image updated using facial expression sequences (e.g., including a first expression fragment and / or a second expression fragment), n images updated using facial expression sequences, the second image updated using facial expression sequences, and m images updated using facial expression sequences. The images corresponding to the first and fourth scripts include: the first image updated using lip-sync sequences, n images updated using facial expression sequences, the second image updated using lip-sync sequences, and m images updated using lip-sync sequences.
[0414] Optionally, if the n images are n images updated according to the lip-sync sequence, then the target video is related to the first action type, the first action attribute, the dialogue, and the tone corresponding to the dialogue.
[0415] In the above method, the first image acquired by the device contains body movements, facial movements, and lip movements, so as to facilitate the subsequent generation of complex character animations (i.e., videos about the target object). Optionally, if the first script (for acquiring full-body body movements), the second script (for acquiring partial body movements), the third script (for acquiring facial expression sequences), and the fourth script (for acquiring lip movement sequences) all include their respective time periods, an action sequence can be obtained by splicing them together according to their respective time periods and / or coverage relationships (e.g., the coverage relationship between partial body movements and full-body movements).
[0416] The scripts included in the embodiments of this application may include a first script, or a first script and a second script, or a first script and a third script, or a first script and a fourth script, or a first script, a second script and a third script, or a first script, a second script and a fourth script, or a first script, a second script and a fourth script, etc., and the specifics are not limited here.
[0417] In this application embodiment, on the one hand, the higher-level semantics of types and attributes in the script are more efficient, intuitive, and understandable. Compared to the prior art that requires users to input low-level control signals, this reduces the user's workload and improves the efficiency of subsequent target video generation. On the other hand, compared to the prior art that requires users to input control signals for each frame, this application can predict n sets of action states based on the first action state and the first information, which is equivalent to obtaining a sequence of actions over a period of time based on the script. This reduces the user's operation and technical requirements, improves the user experience, and increases the efficiency of animation generation. Furthermore, through interactive interfaces and other means, various types of scripts can be flexibly adjusted or set, making the generated animation more versatile and applicable to complex actions of the target object (e.g., limb movements, facial movements, lip movements).
[0418] Understandable Figure 6 The method shown can also be executed jointly by a terminal device and a cloud server. For example, the terminal device obtains first information and a first action state based on its interaction with the user, and sends the first information and the first action state to the cloud server. The cloud server predicts n sets of action states based on the first action state and the first information. The cloud server obtains the target object, processes the target object's limb movements based on the first action state to obtain a first image, and processes the target object's limb movements based on the n sets of action states to obtain n images. The cloud server generates a target video based on the first image and the n images. The cloud server sends the target video to the terminal device. The terminal device plays the target video. Alternatively, the cloud server can obtain the first image and the n images, and then send the first image and the n images to the terminal device. The terminal device then generates the target video based on the first image and the n images and plays the target video. Figure 6 The steps in the illustrated embodiments can be executed by the terminal device, the cloud server, or both the terminal device and the cloud server; no specific limitation is made here.
[0419] The data processing method in the embodiments of this application has been described above. The data processing device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 44 One embodiment of the data processing device in this application includes:
[0420] The acquisition unit 4401 is used to acquire first information, which includes a first action type and a first action attribute. The first action type is used to describe the first limb action, and the first action attribute is used to describe the process of the first limb action.
[0421] The acquisition unit 4401 is also used to acquire the first action state;
[0422] Prediction unit 4402 is used to predict n sets of action states based on the first action state and the first information, where n is a positive integer;
[0423] Acquisition unit 4401 is also used to acquire the target object;
[0424] Processing unit 4403 is used to process the limb movements of the target object based on the first action state to obtain a first image;
[0425] The processing unit 4403 is also used to process the limb movements of the target object based on n sets of action states to obtain n images;
[0426] The generation unit 4404 is used to generate a target video based on a first image and n images, wherein the target video is related to a first action type and a first action attribute.
[0427] In this embodiment, the operations performed by each unit in the data processing device are the same as those described above. Figure 3 or Figure 6 The embodiments shown are similar and will not be repeated here.
[0428] In this embodiment, on the one hand, using higher-level semantics such as types and attributes in the script is more efficient, intuitive, and understandable. Compared to the prior art that requires users to input low-level control signals, this reduces the user's workload and improves the efficiency of the subsequent generation unit 4404 in generating the target video. On the other hand, compared to the prior art where users need to input control signals for each frame, the prediction unit 4402 predicts n sets of action states based on the first action state and the first information, which is equivalent to obtaining a sequence of actions over a period of time. This reduces the user's operational and technical requirements, improves the user experience, and enhances the efficiency of animation generation. Furthermore, by flexibly adjusting the various scripts, the generated animation has strong versatility.
[0429] See Figure 45 This application provides a schematic diagram of another data processing device. The data processing device may include a processor 4501, a memory 4502, and a communication interface 4503. The processor 4501, memory 4502, and communication interface 4503 are interconnected via lines. The memory 4502 stores program instructions and data.
[0430] The aforementioned memory 4502 stores Figure 3 or Figure 6 In the corresponding implementation shown, the program instructions and data corresponding to the steps executed by the device are described.
[0431] Processor 4501, for performing the aforementioned Figure 3 or Figure 6 The steps performed by the device are shown in any of the embodiments illustrated.
[0432] Communication interface 4503 can be used for receiving and sending data, and for performing the aforementioned functions. Figure 3 or Figure 6 The steps related to acquiring, sending, and receiving in any of the embodiments shown.
[0433] In one implementation, the data processing device may include, relative to Figure 45 More or fewer components are merely illustrative in this application and are not intended to limit the scope of the application.
[0434] This application also provides another data processing device, such as... Figure 46 As shown, for ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. The data processing device can be any terminal device including mobile phones, tablets, etc. Taking a mobile phone as an example:
[0435] Figure 46 The diagram shown is a block diagram of a portion of the structure of a mobile phone, a data processing device provided in an embodiment of this application. (Reference) Figure 46 The mobile phone includes components such as a radio frequency (RF) circuit 4610, a memory 4620, an input unit 4630, a display unit 4640, a sensor 4650, an audio circuit 4660, a wireless fidelity (WiFi) module 4670, a processor 4680, and a power supply 4690. Those skilled in the art will understand that... Figure 46 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0436] The following is combined with Figure 46 A detailed introduction to each component of a mobile phone:
[0437] RF circuit 4610 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with processor 4680; additionally, it transmits uplink data to the base station. Typically, RF circuit 4610 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 4610 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0438] The memory 4620 can be used to store software programs and modules. The processor 4680 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 4620. The memory 4620 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 4620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0439] The input unit 4630 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 4630 may include a touch panel 4631 and other input devices 4632. The touch panel 4631, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 4631), and drive the corresponding connection devices according to a pre-set program. Optionally, the touch panel 4631 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 4680, and can also receive and execute commands sent by the processor 4680. In addition, the touch panel 4631 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 4631, the input unit 4630 may also include other input devices 4632. Specifically, other input devices 4632 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0440] Display unit 4640 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. Display unit 4640 may include display panel 4641, optionally configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel 4641. Further, touch panel 4631 may cover display panel 4641. When touch panel 4631 detects a touch operation on or near it, it transmits the information to processor 4680 to determine the type of touch event. Subsequently, processor 4680 provides corresponding visual output on display panel 4641 based on the type of touch event. Although in Figure 46 In this embodiment, the touch panel 4631 and the display panel 4641 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 4631 and the display panel 4641 can be integrated to realize the input and output functions of the mobile phone.
[0441] The mobile phone may also include at least one sensor 4650, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 4641 according to the ambient light level, and the proximity sensor can turn off the display panel 4641 and / or the backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, which can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, IMUs, and SLAM sensors, will not be described in detail here.
[0442] Audio circuit 4660, speaker 4661, and microphone 4662 provide an audio interface between the user and the mobile phone. Audio circuit 4660 converts received audio data into electrical signals and transmits them to speaker 4661, where speaker 4661 converts them into sound signals for output. On the other hand, microphone 4662 converts collected sound signals into electrical signals, which are received by audio circuit 4660, converted into audio data, and then processed by processor 4680 before being transmitted via RF circuit 4610 to, for example, another mobile phone, or the audio data can be output to memory 4620 for further processing.
[0443] WiFi is a short-range wireless transmission technology. Mobile phones using the WiFi module 4670 can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 46 The WiFi module 4670 is shown, but it is understandable that it is not an essential component of a mobile phone.
[0444] The processor 4680 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 4620, and calls data stored in the memory 4620 to perform various functions and process data, thereby providing overall monitoring of the phone. Optionally, the processor 4680 may include one or more processing units; preferably, the processor 4680 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 4680.
[0445] The mobile phone also includes a power supply 4690 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 4680 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0446] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0447] In this embodiment of the application, the processor 4680 included in the mobile phone can perform the aforementioned... Figure 3 or Figure 6 The functions of the data processing device in the illustrated embodiment will not be described in detail here.
[0448] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0449] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0450] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented wholly or partially through software, hardware, firmware, or any combination thereof.
[0451] When the integrated unit is implemented using software, it can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0452] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
Claims
1. A data processing method, characterized in that, The method includes: Edit the first information, which includes a first action type, a first action attribute, and a first start time and / or a first end time of the first action type. The first action type is used to describe a first limb action, and the first action attribute is used to describe the process in which the first limb action occurs. The editing of the first information includes: obtaining a first script input by the user, the first script including the first information; or, displaying a first user interface and responding to a first operation by the user on the first user interface. Obtain the first action state; Based on the first action state and the first information input, the trained first network obtains n sets of action states, where n is a positive integer; Obtain the target object; Based on the first action state, determine the limb action corresponding to the first action state, and redirect the limb action corresponding to the first action state to the target object to obtain the first image; Based on the n sets of action states, determine the limb movements corresponding to the n sets of action states, and redirect the limb movements corresponding to the n sets of action states to the target object to obtain n images; generate a target video based on the first image and the n images, wherein the target video is related to the first action type, the first action attribute, and the first start time and / or the first end time.
2. The method according to claim 1, characterized in that, The first user interface includes an action type option, an action attribute option, and a first start time option and / or a first end time option for the first action type; The response to the user's first operation on the first user interface includes: determining the first action type from the action type options, selecting the first action attribute from the action attribute options, and selecting the first start time from the first start time option of the first action type and / or selecting the first end time from the first end time option of the first action type.
3. The method according to claim 1 or 2, characterized in that, The process of obtaining the first action state includes: Obtain a second action state, which represents a historical action state prior to the first action state; The second action state and the first information are input into the trained first network to obtain the first action state; or, A second user interface is displayed to the user, the second user interface including an action status area; In response to a second user action on the second user interface, the first action state is determined from the action state area; or, Obtain the pre-set state of the first action.
4. The method according to claim 1, characterized in that, The first network is used to obtain a fourth action state based on a third action state, the first action type, and the first action attribute, wherein the change from the third action state to the fourth action state is related to the first action type.
5. The method according to claim 1, characterized in that, The trained first network is obtained by training the first network with the first training data as input and the first loss function value being less than a first threshold. The first training data includes a third action state, a first action type, and a first action attribute. The third action state includes at least one of the third coordinate, third rotation angle, and third velocity of the joint corresponding to the first limb action. The first loss function is used to indicate the difference between the fourth action state output by the first network and the first target action state. The fourth action state includes at least one of the fourth coordinate, fourth rotation angle, and fourth velocity of the joint corresponding to the first limb action. The first target action state includes at least one of the first target coordinate, first target rotation angle, and first target velocity. The first target action state and the third action state belong to the action states corresponding to two adjacent frames in the same action video.
6. The method according to claim 1, characterized in that, The target object is a three-dimensional model, and the target object is used to perform the first limb action.
7. The method according to claim 1, characterized in that, The method further includes: Get the value of n; The process of obtaining n sets of action states based on the first action state and the first information includes: Based on the first action state and the first information, the output is obtained through the first network and iterated to obtain the n sets of action states.
8. The method according to claim 7, characterized in that, The method further includes: Obtain an end condition, which includes at least one of the following: the completion progress of the first limb movement, the end time of the first movement type, and the parameters in the first movement attribute. The end condition is either preset or input by the user on a third user interface. Obtaining the value of n includes: The value of n is determined based on the termination condition.
9. The method according to claim 1, characterized in that, The method further includes: Obtain environmental information, which includes at least one of props and objects that interact with the target object; The process of obtaining n sets of action states based on the first action state and the first information includes: Based on the first action state, the first information, and the environmental information, the n sets of action states and n environmental contact information corresponding to the n sets of action states are obtained. One of the n environmental contact information is used to indicate whether the joint in the action state corresponding to the environmental contact information is in contact with the environmental information. The process of processing the limb movements of the target object based on the n sets of action states to obtain n images includes: The n images are obtained by processing the limb movements of the target object based on the n sets of action states and the environmental contact information.
10. The method according to claim 1, characterized in that, The method further includes: Obtain second information, which includes a second action type and a second action attribute. The second action type is used to describe a second limb action, and the second action attribute is used to describe the process in which the second limb action occurs. Obtain the state of the fifth action; Based on the fifth action state and the second information, m sets of action states are obtained, where m is a positive integer; The second image is obtained by processing the limb movements of the target object based on the fifth action state; Based on the m sets of action states, the limb movements of the target object are processed to obtain m images; The step of generating a target video based on the first image and the n images includes: The target video is generated based on the first image, the n images, the second image, and the m images.
11. The method according to claim 10, characterized in that, The first information also includes a first time period of the first action type, and the second information also includes a second time period of the second action type; The process of generating the target video based on the first image, the n images, the second image, and the m images includes: The target video is generated based on the first image, the n images, the second image, the m images, the first time period, and the second time period, wherein the first time period corresponds to the first image and the n images, and the second time period corresponds to the second image and the m images.
12. The method according to claim 11, characterized in that, The process of obtaining m sets of action states based on the fifth action state and the second information includes: The fifth action state and the second information are input into the trained second network to obtain the m sets of action states. The second network is used to obtain the seventh action state based on the sixth action state, the second action type and the second action attribute. The change from the sixth action state to the seventh action state is related to the second action type.
13. The method according to claim 12, characterized in that, The trained second network is obtained by training the second network with the second training data as input and the second loss function value being less than a second threshold. The second training data includes a sixth action state, a second action type, and attributes related to the second action. The sixth action state includes at least one of the sixth coordinate, sixth rotation angle, and sixth velocity of the joint corresponding to the second limb action. The second loss function is used to indicate the difference between the seventh action state output by the second network and the second target action state. The seventh action state includes at least one of the seventh coordinate, seventh angle, and seventh velocity of the joint corresponding to the second limb action. The second target action state includes at least one of the second target coordinate, second target angle, and second target velocity. The second target action state and the sixth action state belong to the action states corresponding to two adjacent frames in the same action video.
14. The method according to claim 1, characterized in that, The method further includes: Obtain third information, which includes a third action type and a third action attribute. The third action type is used to describe a third limb action, and the third action attribute is used to describe the process of the third limb action. The third limb corresponding to the third limb action is a part of the first limb corresponding to the first limb action. Obtain the state of the eighth action; Based on the eighth action state and the third information, p groups of action states are obtained, where p is a positive integer; The process of processing the limb movements of the target object based on the first action state to obtain the first image includes: Based on the overlap relationship between the third limb and the first limb, the first image is obtained by overlaying the limb action of the target object corresponding to the first action state with the limb action of the target object corresponding to the eighth action state. Based on the n sets of action states, the limb movements of the target object are processed to obtain n images, including: Based on the overlap relationship between the third limb and the first limb, the limb actions of the target object corresponding to the p groups of action states are used to overlap the limb actions of the target object corresponding to the n groups of action states to obtain the n images.
15. The method according to claim 1, characterized in that, The method further includes: Acquire facial information, which includes facial expression type and expression attribute corresponding to the facial expression type. The facial expression type is used to describe the facial movements of the target object, and the expression attribute is used to describe the amplitude of the facial movements. A facial expression sequence is obtained based on the facial information and a first association relationship, wherein the first association relationship is used to represent the association relationship between the facial information and the facial expression sequence; The process of processing the limb movements of the target object based on the first action state to obtain the first image includes: The first image is obtained by processing the limb and facial movements of the target object based on the first action state and the facial expression sequence. The process of processing the limb movements of the target object based on the n sets of action states to obtain n images includes: The n images are obtained by processing the limb and facial movements of the target object based on the n sets of action states and the facial expression sequence.
16. The method according to claim 1, characterized in that, The method further includes: Retrieve text information; Generate speech segments based on the text information; A lip-sync sequence is generated based on the speech segment, and the lip-sync sequence is used to describe the lip-sync of the target object; The process of processing the limb movements of the target object based on the first action state to obtain the first image includes: The first image is obtained by processing the limb and facial movements of the target object based on the first action state and the lip-sync sequence. The process of processing the limb movements of the target object based on the n sets of action states to obtain n images includes: The n images are obtained by processing the limb and facial movements of the target object based on the n sets of action states and the lip-sync sequence.
17. The method according to claim 16, characterized in that, The step of generating a target video based on the first image and the n images includes: The target video is generated based on the first image, the n images, and the audio segment.
18. The method according to claim 1, characterized in that, The first type of action includes at least one of walking, running, jumping, sitting down, standing up, squatting down, lying down, hugging, punching, swinging a sword, and dancing.
19. The method according to claim 1, characterized in that, The first action attribute includes at least one of displacement, travel path, action speed, frequency of action occurrence, action amplitude, and action orientation.
20. A data processing device, characterized in that, The device includes: The acquisition unit is used to edit first information, which includes a first action type, a first action attribute, and a first start time and / or a first end time of the first action type. The first action type is used to describe a first limb action, and the first action attribute is used to describe the process of the first limb action. Specifically, when the acquisition unit edits the first information, it is used to acquire a first script input by the user, the first script including the first information; or, to display a first user interface and respond to the user's first operation on the first user interface. The acquisition unit is further configured to acquire the first action state; The prediction unit is used to obtain n sets of action states based on the first action state and the first information input trained on the first network, where n is a positive integer; The acquisition unit is also used to acquire the target object; The processing unit is configured to determine the limb action corresponding to the first action state based on the first action state, redirect the limb action corresponding to the first action state to the target object, and obtain the first image; The processing unit is further configured to determine the limb movements corresponding to the n sets of action states based on the n sets of action states, and redirect the limb movements corresponding to the n sets of action states to the target object to obtain n images; The generation unit is configured to generate a target video based on the first image and the n images, wherein the target video is related to the first action type, the first action attribute, and the first start time and / or the first end time.
21. The device according to claim 20, characterized in that, The first user interface includes an action type option, an action attribute option, and a first start time option and / or a first end time option for the first action type; The acquisition unit is specifically used to display a first user interface, which includes action type options and / or action attribute options. When the acquisition unit responds to the user's first operation on the first user interface, it is specifically used to determine the first action type from the action type option, select the first action attribute from the action attribute option, and select the first start time from the first start time option of the first action type and / or select the first end time from the first end time option of the first action type.
22. The device according to claim 20 or 21, characterized in that, The acquisition unit is specifically used to acquire a second action state, which represents a historical action state prior to the first action state. The acquisition unit is specifically used to input the second action state and the first information into the trained first network to obtain the first action state; or, The acquisition unit is specifically used to display a second user interface to the user, the second user interface including an action state area; The acquisition unit is specifically used to respond to a second operation by the user in the second user interface and determine the first action state from the action state area. or, The acquisition unit is specifically used to acquire the pre-set first action state.
23. The device according to claim 20, characterized in that, The first network is used to obtain a fourth action state based on a third action state, the first action type, and the first action attribute, wherein the change from the third action state to the fourth action state is related to the first action type.
24. The device according to claim 23, characterized in that, The trained first network is obtained by training the first network with the first training data as input and the first loss function value being less than a first threshold. The first training data includes the third action state, the first action type, and the first action attribute. The third action state includes at least one of the third coordinate, third rotation angle, and third velocity of the joint corresponding to the first limb action. The first loss function is used to indicate the difference between the fourth action state output by the first network and the first target action state. The fourth action state includes at least one of the fourth coordinate, fourth rotation angle, and fourth velocity of the joint corresponding to the first limb action. The first target action state includes at least one of the first target coordinate, first target rotation angle, and first target velocity. The first target action state and the third action state belong to the action states corresponding to two adjacent frames in the same action video.
25. The device according to claim 20, characterized in that, The target object is a three-dimensional model, and the target object is used to perform the first limb action.
26. The device according to claim 20, characterized in that, The acquisition unit is also used to acquire the value of n; The prediction unit is specifically used to obtain the n sets of action states by iterating through the first network based on the first action state and the first information.
27. The device according to claim 26, characterized in that, The acquisition unit is further configured to acquire an end condition, which includes at least one of the following: the completion progress of the first limb movement, the end time of the first movement type, and the parameters in the first movement attribute. The end condition is either preset or input by the user on a third user interface. The acquisition unit is specifically used to determine the value of n based on the termination condition.
28. The device according to claim 20, characterized in that, The acquisition unit is further configured to acquire environmental information, which includes at least one of props and objects that interact with the target object. The prediction unit is specifically used to obtain the n sets of action states and n environmental contact information corresponding to the n sets of action states based on the first action state, the first information and the environmental information. One of the environmental contact information is used to indicate whether the joint in the action state corresponding to the environmental contact information is in contact with the environmental information. The processing unit is specifically used to process the limb movements of the target object based on the n sets of action states and the environmental contact information to obtain the n images.
29. The device according to claim 20, characterized in that, The acquisition unit is further configured to acquire second information, the second information including a second action type and a second action attribute, the second action type being used to describe a second limb action, and the second action attribute being used to describe the process in which the second limb action occurs; The acquisition unit is also used to acquire the fifth action state; The prediction unit is further configured to obtain m sets of action states based on the fifth action state and the second information, where m is a positive integer; The processing unit is further configured to process the limb movements of the target object based on the fifth action state to obtain a second image; The processing unit is also used to process the limb movements of the target object based on the m sets of action states to obtain m images; The generation unit is specifically used to generate the target video based on the first image, the n images, the second image, and the m images.
30. The device according to claim 29, characterized in that, The first information also includes a first time period of the first action type, and the second information also includes a second time period of the second action type; The generation unit is specifically used to generate the target video based on the first image, the n images, the second image, the m images, the first time period, and the second time period, wherein the first time period corresponds to the first image and the n images, and the second time period corresponds to the second image and the m images.
31. The device according to claim 30, characterized in that, The prediction unit is specifically used to input the fifth action state and the second information into the trained second network to obtain the m sets of action states. The second network is used to obtain the seventh action state based on the sixth action state, the second action type and the second action attribute. The change from the sixth action state to the seventh action state is related to the second action type.
32. The device according to claim 31, characterized in that, The trained second network is obtained by training the second network with the second training data as input and the second loss function value being less than a second threshold. The second training data includes the sixth action state, the second action type, and the second action attribute. The sixth action state includes at least one of the sixth coordinate, sixth rotation angle, and sixth velocity of the joint corresponding to the limb action. The second loss function is used to indicate the difference between the seventh action state output by the second network and the second target action state. The seventh action state includes at least one of the seventh coordinate, seventh angle, and seventh velocity of the joint corresponding to the second limb action. The second target action state includes at least one of the second target coordinate, second target angle, and second target velocity. The second target action state and the sixth action state belong to the action states corresponding to two adjacent frames in the same action video.
33. The device according to claim 20, characterized in that, The acquisition unit is further configured to acquire third information, the third information including a third action type and a third action attribute, the third action type being used to describe a third limb action, the third action attribute being used to describe the process of the third limb action occurring, and the third limb corresponding to the third limb action being a local limb in the first limb corresponding to the first limb action. The acquisition unit is also used to acquire the eighth action state; The prediction unit is further configured to obtain p groups of action states based on the eighth action state and the third information, where p is a positive integer; The processing unit is specifically used to cover the limb movements of the target object corresponding to the first action state with the limb movements of the target object corresponding to the eighth action state based on the coverage relationship between the third limb and the first limb to obtain the first image; The processing unit is specifically used to cover the limb actions of the target object corresponding to the n groups of action states with the limb actions of the target object corresponding to the p groups of action states based on the coverage relationship between the third limb and the first limb to obtain the n images.
34. The device according to claim 20, characterized in that, The acquisition unit is further configured to acquire facial information, the facial information including facial expression type and expression attribute corresponding to the facial expression type, the facial expression type being used to describe the facial movements of the target object, and the expression attribute being used to describe the amplitude of the facial movements; The acquisition unit is further configured to acquire a facial expression sequence based on the facial information and a first association relationship, wherein the first association relationship is used to represent the association relationship between the facial information and the facial expression sequence; The processing unit is specifically used to process the limb movements and facial movements of the target object based on the first action state and the facial expression sequence to obtain the first image; The processing unit is specifically used to process the limb movements and facial movements of the target object based on the n sets of action states and the facial expression sequence to obtain the n images.
35. The device according to claim 20, characterized in that, The acquisition unit is also used to acquire text information; The generation unit is also used to generate speech segments based on the text information; The generation unit is further configured to generate a lip-shape sequence based on the speech segment, the lip-shape sequence being used to describe the lip shape of the target object; The processing unit is specifically used to process the limb movements and facial movements of the target object based on the first action state and the lip-sync sequence to obtain the first image; The processing unit is specifically used to process the limb and facial movements of the target object based on the n sets of action states and the lip-sync sequence to obtain the n images.
36. The device according to claim 35, characterized in that, The generation unit is specifically used to generate the target video based on the first image, the n images, and the audio segment.
37. The device according to claim 20, characterized in that, The first type of action includes at least one of walking, running, jumping, sitting down, standing up, squatting down, lying down, hugging, punching, swinging a sword, and dancing.
38. The device according to claim 20, characterized in that, The first action attribute includes at least one of displacement, travel path, action speed, frequency of action occurrence, action amplitude, and action orientation.
39. A data processing device, characterized in that, include: The method includes a processor coupled to a memory storing a program, wherein the program instructions stored in the memory are executed by the processor to implement the method of any one of claims 1 to 19.
40. A computer-readable storage medium, characterized in that, Includes a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 19.
41. A computer program product, characterized in that, When the computer program product is executed on a computer, it causes the computer to perform the method as described in any one of claims 1 to 19.
Citation Information
Patent Citations
Animation generation method and device, electronic equipment and storage medium
CN111223170A
Three-dimensional animation attitude prediction method and system
CN111311714A