Method, device, readable medium and electronic device for generating an animation curve

By generating animation curves in continuous space, the problem of low accuracy of animation curves is solved, enabling flexible control and efficient performance of animation.

CN116309989BActive Publication Date: 2026-05-26BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2023-01-31
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of animation curves is low, especially in the inability to accurately perceive invisible parts (such as teeth), resulting in poor animation effects. Furthermore, sampling can only be performed at fixed moments, making it impossible to implement acceleration and deceleration operations.

Method used

By acquiring the text and audio information of the target animation, determining the phoneme information, and inputting pre-generated curve parameters to obtain the model, a target animation curve in continuous space is generated, allowing sampling at any time and flexible adjustment of the animation speed.

Benefits of technology

It improves the accuracy and effect of animation, and can sample the animation curve at any time to realize the acceleration and deceleration of the animation, thus enhancing the expressiveness of the animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309989B_ABST
    Figure CN116309989B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, readable medium, and electronic device for generating animation curves. The method includes: acquiring target animation text and target animation audio for a target animation to be generated; determining multiple target phoneme information corresponding to the target animation text based on the target animation text and the target animation audio, wherein the target phoneme information includes the target phoneme and its phoneme timing information; inputting the multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model; and generating a target animation curve based on the multiple target curve parameters, wherein the target animation curve is used to generate the target animation. In other words, the target animation curve generated by this disclosure is a curve in continuous space, and when generating an animation based on this animation curve, the animation curve can be sampled at any time, improving the animation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a method, apparatus, readable medium, and electronic device for generating animation curves. Background Technology

[0002] Voice-driven lip-sync refers to a technique that drives the lip movements of a person in a video based on input audio information, while keeping all other information in the base video unchanged except for lip information. Lip-sync driving relies on animation curves, which are currently obtained through facial capture. However, facial capture is generally image-based and cannot accurately perceive some invisible parts of the face (such as teeth), resulting in relatively low accuracy of the animation curves.

[0003] In related technologies, animators determine keyframes from animation curves and correct errors introduced by the animation curves by revising the values ​​of the keyframes. However, since the animation curves are revised based on keyframes, sampling can only be performed at fixed moments when generating animations based on these animation curves, resulting in relatively poor animation effects. Summary of the Invention

[0004] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, this disclosure provides a method for generating animation curves, including:

[0006] Obtain the target animation text and target animation audio of the target animation to be generated;

[0007] Based on the target animation text and the target animation audio, determine multiple target phoneme information corresponding to the target animation text, wherein the target phoneme information includes the target phoneme and the phoneme time information of the target phoneme;

[0008] Multiple target phoneme information are input into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model.

[0009] A target animation curve is generated based on multiple target curve parameters, and the target animation curve is used to generate the target animation.

[0010] Secondly, this disclosure provides an apparatus for generating animation curves, comprising:

[0011] The first acquisition module is used to acquire the target animation text and target animation audio of the target animation to be generated;

[0012] The determining module is used to determine multiple target phoneme information corresponding to the target animation text based on the target animation text and the target animation audio, wherein the target phoneme information includes the target phoneme and the phoneme time information of the target phoneme;

[0013] The second acquisition module is used to input multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model.

[0014] The first generation module is used to generate a target animation curve based on a plurality of target curve parameters, wherein the target animation curve is used to generate the target animation.

[0015] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of this disclosure.

[0016] Fourthly, this disclosure provides an electronic device, comprising:

[0017] A storage device having at least one computer program stored thereon;

[0018] At least one processing means is configured to execute the at least one computer program in the storage device to implement the steps of the method described in the first aspect of this disclosure.

[0019] The above technical solution obtains the target animation text and target animation audio of the target animation to be generated; based on the target animation text and target animation audio, determines multiple target phoneme information corresponding to the target animation text, the target phoneme information including the target phoneme and the phoneme time information of the target phoneme; inputs the multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model; and generates a target animation curve based on the multiple target curve parameters, the target animation curve being used to generate the target animation. In other words, this disclosure first generates multiple target curve parameters based on multiple target factor information corresponding to the target animation, and then generates a target animation curve based on the target curve parameters. This target animation curve is a curve in continuous space, and when generating the animation based on this animation curve, the animation curve can be sampled at any time, improving the animation effect.

[0020] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0022] Figure 1 This is a flowchart illustrating a method for generating animation curves according to an exemplary embodiment of the present disclosure;

[0023] Figure 2 This is a flowchart illustrating another method for generating animation curves according to an exemplary embodiment of the present disclosure;

[0024] Figure 3 It is based on Figure 2 The illustrated embodiment shows a flowchart of another method for generating animation curves;

[0025] Figure 4 It is based on Figure 3 The illustrated embodiment shows a flowchart of another method for generating animation curves;

[0026] Figure 5 This is a schematic diagram illustrating a curve parameter acquisition model according to an exemplary embodiment of the present disclosure;

[0027] Figure 6 This is a flowchart illustrating a model training method according to an exemplary embodiment of the present disclosure;

[0028] Figure 7 This is a block diagram illustrating an apparatus for generating animation curves according to an exemplary embodiment of the present disclosure;

[0029] Figure 8 This is a block diagram illustrating another apparatus for generating animation curves according to an exemplary embodiment of the present disclosure;

[0030] Figure 9 This is a block diagram illustrating another apparatus for generating animation curves according to an exemplary embodiment of the present disclosure;

[0031] Figure 10 This is a block diagram illustrating an apparatus for generating animation curves according to an exemplary embodiment of the present disclosure;

[0032] Figure 11 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0033] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0034] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0035] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0036] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0037] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0038] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0039] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0040] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0041] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0042] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0043] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0044] First, the application scenario of this disclosure will be explained. Currently, facial expression animation curves are implemented using facial capture devices and algorithms. The basic process of facial capture includes: 1. Capturing a face video; 2. Initializing facial expression animation parameters, calculating the vertex positions of the mesh within the algorithm under the current facial expression parameters for each frame of the face video; 3. Iteratively optimizing the facial expression animation parameters based on the difference between the key point positions of the face in the image and the vertex positions of the mesh, ultimately obtaining the animation curve. As can be seen from the above basic process of facial capture, facial capture is image-based and cannot perceive some invisible parts (such as teeth). Therefore, when the mouth is open but the teeth are clenched, facial capture generally cannot accurately estimate the state of the teeth, resulting in a relatively large error in the raw data obtained from facial capture. To solve the error in the raw data, animators need to refine the animation by finding keyframes in the facial expression animation curve and modifying the values ​​of the keyframes to obtain a refined animation curve. However, since this animation curve is based on keyframe revision, when generating animation based on this animation curve, sampling can only be performed at fixed moments, and operations such as acceleration and deceleration of the animation cannot be performed, resulting in a relatively poor animation effect.

[0045] To address the aforementioned problems, this disclosure provides a method, apparatus, readable medium, and electronic device for generating animation curves. First, multiple target curve parameters are generated based on multiple target factor information corresponding to the target animation. Then, a target animation curve is generated based on the target curve parameters. This target animation curve is a curve in continuous space. When generating animation based on this animation curve, the animation curve can be sampled at any time, improving the animation effect.

[0046] The present disclosure will now be described in conjunction with specific embodiments.

[0047] Figure 1 This is a flowchart illustrating a method for generating animation curves according to an exemplary embodiment of the present disclosure, such as... Figure 1 As shown, the method may include:

[0048] S101. Obtain the target animation text and target animation audio of the target animation to be generated.

[0049] The target animation audio can be the audio that the target animation needs to output, and the target animation text can be the text corresponding to the target animation audio.

[0050] S102. Based on the target animation text and the target animation audio, determine multiple target phoneme information corresponding to the target animation text.

[0051] The target phoneme information may include the target phoneme and the phoneme time information of the target phoneme. For example, the phoneme time information may include the start time and end time of the target phoneme. The phoneme time information may also be the intermediate time between the start time and the end time of the target phoneme. This disclosure does not limit this.

[0052] In this step, after obtaining the target animated text and the target animated audio, multiple target phonemes corresponding to the target animated text and the phoneme timing information of each target phoneme can be determined using existing techniques. For example, if the target animated text is "today", then the target phonemes may include: j, i, n, t, i, a, n.

[0053] It should be noted that the target phoneme can also be represented by IPA (International Phonetic Alphabet), and this disclosure does not limit it to that.

[0054] S103. Input multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model.

[0055] The target curve parameters may include parameters that can generate the curve. For example, the target curve parameters may include the current time, the left neighbor tangent parameter, and the right neighbor tangent parameter.

[0056] In this step, after determining multiple target phoneme information, the multiple target phoneme information can be input into the curve parameter acquisition model. The curve parameter acquisition model is used to perform splicing, fusion, feature acquisition and other processing on the multiple target phoneme information to obtain multiple target curve parameters corresponding to each target phoneme information.

[0057] S104. Generate the target animation curve based on multiple target curve parameters.

[0058] The target animation curve can be used to generate the target animation.

[0059] In this step, after obtaining multiple target curve parameters corresponding to each target phoneme information, a phoneme animation curve corresponding to that target phoneme information can be generated based on the multiple target curves. Then, the phoneme animation curves corresponding to each target phoneme information are connected to obtain the target animation curve.

[0060] In one possible implementation, after generating the target animation curve, the animation image of the target animation can be obtained; the target animation is then generated based on the target animation curve and the animation image. For example, after generating the target animation curve, the animation image can be rendered using an existing image rendering engine based on the target animation curve to obtain the target animation. If the target animation curve is represented as f(t) = at... 2 +b, where a and b are known constants and t represents time, then by adjusting t, the playback speed of the target animation can be arbitrarily adjusted.

[0061] Using the above method, multiple target curve parameters are first generated based on the target factor information corresponding to the target animation. Then, a target animation curve is generated based on these parameters. This target animation curve is a continuous spatial curve. When generating animation based on this curve, sampling of the animation curve can be performed at any time. When rendering video, a fixed frame rate is not required; the animation curve can be directly solved based on the rendering engine's real time, improving the animation effect. Furthermore, this target animation curve can be flexibly scaled on the timeline, enabling acceleration and deceleration operations in the animation, further enhancing the animation effect.

[0062] The curve parameter acquisition model can include multiple feature acquisition sub-models, with different target curve parameters corresponding to different feature acquisition sub-models. Figure 2 This is a flowchart illustrating another method for generating animation curves according to an exemplary embodiment of the present disclosure, such as... Figure 2 As shown, the implementation of step S103 may include:

[0063] S1031. Input multiple target phoneme information into the curve parameter acquisition model, and obtain multiple target phoneme features through multiple feature acquisition sub-models.

[0064] In this step, after obtaining multiple target phoneme information, these multiple target phoneme information can be input into multiple feature acquisition sub-models respectively. Taking the target curve parameters including the current time, left neighbor tangent parameters, and right neighbor tangent parameters as an example, the feature acquisition sub-model can include three. The multiple target phoneme information are input into three feature acquisition sub-models respectively. For each feature acquisition sub-model, the feature acquisition sub-model can determine the target phoneme feature corresponding to each target phoneme information.

[0065] S1032. Based on multiple target phoneme features, determine multiple target curve parameters.

[0066] In one possible implementation, the curve parameter acquisition model may further include multiple parameter generation sub-models, the input of which is coupled to the output of the feature acquisition sub-model.

[0067] In this step, for each feature acquisition sub-model, after the feature acquisition sub-model outputs multiple target phoneme features, the multiple target phoneme features can be input into the parameter generation sub-model coupled with the feature acquisition sub-model, and the target curve parameters can be output through the parameter generation sub-model.

[0068] The curve parameter acquisition model also includes a feature fusion sub-model, the output of which is coupled to the inputs of multiple feature acquisition sub-models. Figure 3 It is based on Figure 2 The illustrated embodiment shows a flowchart of another method for generating animation curves, as shown in the figure. Figure 3 As shown, the method may further include:

[0069] S1033. The target phoneme information is fused through the feature fusion sub-model to obtain the target fusion feature.

[0070] In this step, after obtaining multiple target phoneme information, the multiple target phoneme information can be input into the feature fusion sub-model. The feature fusion sub-model then fuses the features of the multiple target phoneme information to obtain the target fusion feature.

[0071] Step S1031 can be implemented as follows:

[0072] S1034. Input the target fusion feature into multiple feature acquisition sub-models respectively to obtain the target phoneme features output by each feature acquisition sub-model.

[0073] In this step, after obtaining the target fusion feature, the target fusion feature can be input into multiple feature acquisition sub-models respectively, and the target phoneme feature corresponding to each target phoneme information of each sub-model can be obtained through each feature.

[0074] The curve parameter acquisition model also includes a splicing sub-model, the output of which is coupled with the input of the feature fusion sub-model; Figure 4 It is based on Figure 3 The illustrated embodiment shows a flowchart of another method for generating animation curves, as shown in the figure. Figure 4 As shown, the method may further include:

[0075] S1035. The target phoneme information is spliced ​​through the splicing sub-model to obtain the target splicing feature.

[0076] In this step, after obtaining multiple target phoneme information, the multiple target phoneme information can be input into the splicing sub-model, and the splicing sub-model can be used to splice the multiple target phoneme information to obtain the target splicing feature.

[0077] Step S1033 can be implemented as follows:

[0078] S1036. The target splicing features are fused through the feature fusion sub-model to obtain the target fused features.

[0079] In this step, after obtaining the target splicing features, the target splicing features can be input into the feature fusion sub-model. The feature fusion sub-model performs feature fusion processing on the target splicing features to obtain the target fused features.

[0080] Figure 5 This is a schematic diagram illustrating a curve parameter acquisition model according to an exemplary embodiment of the present disclosure, such as... Figure 5As shown, the curve parameter model includes a splicing sub-model, a feature fusion sub-model, and three feature acquisition sub-models. Each feature acquisition sub-model includes a self-attention mechanism and a long short-term memory network. The output of the splicing sub-model is coupled to the input of the feature fusion sub-model, and the output of the feature fusion sub-model is coupled to the input of each feature acquisition sub-model. After obtaining multiple target phoneme information, the splicing sub-model processes the multiple target phoneme information by splicing them to obtain target splicing features. These target splicing features are then input to the feature fusion sub-model, which performs feature fusion processing to obtain target fused features. These target fused features are then input to the three feature acquisition sub-models, which obtain three target phoneme features corresponding to each target phoneme information. Finally, based on each target phoneme feature, the target curve parameters corresponding to each target phoneme information are determined, ultimately yielding the three target curve parameters for each target phoneme information.

[0081] Figure 6 This is a flowchart illustrating a model training method according to an exemplary embodiment of the present disclosure, such as... Figure 6 As shown, the method may include:

[0082] S601. Obtain multiple sample sets.

[0083] The sample set includes sample animation curves corresponding to sample animations and multiple sample phoneme information. The sample phoneme information may include the sample phoneme and its phoneme timing information. For example, the phoneme timing information may include the start time and end time of the sample phoneme, or it may be the intermediate time between the start and end times of the sample phoneme. This disclosure does not limit this.

[0084] In this step, multiple sample animations can be acquired. For each sample animation, its sample animation text and audio can be obtained. Based on the sample animation text and audio, multiple sample phonemes and their timing information are determined, resulting in multiple sample phoneme information. Additionally, for each sample animation, the animator can first refine it to obtain the corresponding target sample animation. Then, the multiple sample phoneme information corresponding to the target sample animation is acquired. Based on this information, multiple curve sampling points for facial capture are determined. Curve fitting is then performed on these sampling points to obtain the sample animation curve corresponding to the target sample animation.

[0085] S602. Train the target neural network model using multiple sample sets to obtain the curve parameter acquisition model.

[0086] In one possible implementation, after obtaining multiple sample sets, the current sample set can be determined from the multiple sample sets, and the model training steps can be executed iteratively based on the current sample set until the trained target neural network model meets the preset stopping iteration condition. The trained target neural network model is then used as the curve parameter to obtain the model.

[0087] The model training steps may include:

[0088] S1. Input the multiple phoneme information of the current sample set into the target neural network model, and determine the multiple predicted curve parameters corresponding to the sample animation of the current sample set through the target neural network model.

[0089] S2. Generate a predicted animation curve based on multiple parameters of the predicted curve.

[0090] S3. Determine the target loss value based on the predicted animation curve and the sample animation curve.

[0091] S4. If the target neural network does not meet the preset stopping iteration condition based on the target loss value, update the parameters of the target neural network model based on the target loss value to obtain the trained target neural network model. Use the trained target neural network model as the new target neural network model and determine a new current sample set from multiple sample sets.

[0092] For example, after acquiring multiple sample sets, one sample set can be selected as the current sample set. Multiple sample phoneme information corresponding to this sample set is input into the target neural network to obtain multiple predicted curve parameters corresponding to each sample phoneme information output by the target neural network. For each sample phoneme information, an animation curve corresponding to that sample phoneme information is generated based on the multiple predicted curve parameters. Then, the animation curves corresponding to the multiple sample phoneme information are concatenated to obtain the predicted animation curve. Based on the predicted animation curve and the sample animation curve, the target loss value is determined using a preset loss function. This preset loss function can be a loss function from the prior art, and this disclosure does not limit it.

[0093] After determining the target loss value, a preset loss threshold can be obtained. If the target loss threshold is less than or equal to the preset loss threshold, the target neural network model is determined to meet the preset stopping iteration condition, and this target neural network model is used as the curve parameter acquisition model. If the target loss value is greater than the preset loss threshold, the target neural network model is determined to not meet the preset stopping iteration condition. The parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model. This trained target neural network model is used as the new target neural network model, and another sample set is randomly selected from multiple sample sets as the new current sample set. The above model training steps are then performed again based on this new current sample set until the trained target neural network model is determined to meet the preset stopping iteration condition, and this trained target neural network model is used as the curve parameter acquisition model.

[0094] It should be noted that the preset stopping iteration condition can also be other iteration conditions in the prior art, and this disclosure does not limit it.

[0095] Figure 7 This is a block diagram illustrating an apparatus for generating animation curves according to an exemplary embodiment of the present disclosure, such as... Figure 7 As shown, the device may include:

[0096] The first acquisition module 701 is used to acquire the target animation text and target animation audio of the target animation to be generated;

[0097] The determining module 702 is used to determine multiple target phoneme information corresponding to the target animation text based on the target animation text and the target animation audio, wherein the target phoneme information includes the target phoneme and the phoneme time information of the target phoneme;

[0098] The second acquisition module 703 is used to input multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model.

[0099] The first generation module 704 is used to generate a target animation curve based on multiple target curve parameters, and the target animation curve is used to generate the target animation.

[0100] Optionally, the curve parameter acquisition model includes multiple feature acquisition sub-models, with different target curve parameters corresponding to different feature acquisition sub-models; the second acquisition module 703 is also used for:

[0101] Multiple target phoneme information are input into the curve parameters to obtain the model, and multiple target phoneme features are obtained through multiple feature acquisition sub-models.

[0102] Based on multiple characteristics of the target phoneme, multiple curve parameters of the target are determined.

[0103] Optionally, the curve parameter acquisition model further includes a feature fusion sub-model, the output of which is coupled to the inputs of multiple feature acquisition sub-models respectively; Figure 8 This is a block diagram illustrating another apparatus for generating animation curves according to an exemplary embodiment of the present disclosure, such as... Figure 8 As shown, the device also includes:

[0104] The fusion module 705 is used to fuse multiple target phoneme information through the feature fusion sub-model to obtain target fusion features;

[0105] The second acquisition module 703 is also used for:

[0106] The target fusion feature is input into multiple feature acquisition sub-models to obtain the target phoneme features output by each feature acquisition sub-model.

[0107] Optionally, the curve parameter acquisition model also includes a splicing sub-model, the output of which is coupled to the input of the feature fusion sub-model; Figure 9 This is a block diagram illustrating another apparatus for generating animation curves according to an exemplary embodiment of the present disclosure, such as... Figure 9 As shown, the device also includes:

[0108] The splicing module 706 is used to splice multiple target phoneme information through the splicing sub-model to obtain target splicing features;

[0109] The fusion module 705 is also used for:

[0110] The feature fusion sub-model is used to fuse multiple spliced ​​features of the target to obtain the target fused feature.

[0111] Optionally, the curve parameter acquisition model is pre-generated using the following method:

[0112] Acquire multiple sample sets, which include the sample animation curves corresponding to the sample animations and multiple sample phoneme information, which includes the sample phonemes and the phoneme time information of the sample phonemes;

[0113] The target neural network model is trained using multiple samples to obtain the curve parameter acquisition model.

[0114] Optionally, the model for obtaining the curve parameters can be obtained by training the target neural network model using multiple sample sets, including:

[0115] The current sample set is determined from multiple such sample sets;

[0116] The model training steps are executed iteratively based on the current sample set until the trained target neural network model meets the preset stopping iteration condition. The trained target neural network model is then used as the curve parameter to obtain the model.

[0117] The training steps for this model include:

[0118] The multiple phoneme information corresponding to the current sample set is input into the target neural network model, and the target neural network model is used to determine multiple predicted curve parameters corresponding to the sample animation of the current sample set.

[0119] Generate a predicted animated curve based on multiple parameters of the predicted curve;

[0120] Based on the predicted animation curve and the sample animation curve, determine the target loss value;

[0121] If the target neural network does not meet the preset stopping iteration condition based on the target loss value, the parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model. The trained target neural network model is then used as the new target neural network model, and a new current sample set is determined from multiple sample sets.

[0122] Optionally, Figure 10 This is a block diagram illustrating an apparatus for generating animation curves according to an exemplary embodiment of the present disclosure, such as... Figure 10 As shown, the device also includes:

[0123] The third acquisition module 707 is used to acquire the animation image of the target animation;

[0124] The second generation module 708 is used to generate the target animation based on the target animation curve and the animation image.

[0125] The aforementioned device first generates multiple target curve parameters based on various target factor information corresponding to the target animation. Then, it generates a target animation curve based on these parameters. This target animation curve is a continuous spatial curve. When generating animation based on this curve, sampling can be performed at any time. During video rendering, a fixed frame rate is not required; the animation curve can be directly solved based on the rendering engine's real-time calculations, thus improving the animation effect. Furthermore, this target animation curve can be flexibly scaled along the timeline, enabling acceleration and deceleration operations in the animation, further enhancing its quality.

[0126] The following is for reference. Figure 11This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0127] like Figure 11 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0128] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 11 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0129] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0130] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0131] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0132] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0133] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire target animation text and target animation audio of the target animation to be generated; determine multiple target phoneme information corresponding to the target animation text based on the target animation text and the target animation audio, the target phoneme information including target phonemes and phoneme time information of the target phonemes; input the multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model; and generate a target animation curve based on the multiple target curve parameters, the target animation curve being used to generate the target animation.

[0134] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0136] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the module itself; for example, the first acquisition module can also be described as "a module for acquiring the target animation text and target animation audio of the target animation to be generated".

[0137] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0138] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0139] According to one or more embodiments of this disclosure, Example 1 provides a method for generating an animation curve, comprising: acquiring target animation text and target animation audio of a target animation to be generated; determining multiple target phoneme information corresponding to the target animation text based on the target animation text and the target animation audio, wherein the target phoneme information includes target phonemes and phoneme time information of the target phonemes; inputting the multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model; and generating a target animation curve based on the multiple target curve parameters, wherein the target animation curve is used to generate the target animation.

[0140] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the curve parameter acquisition model includes multiple feature acquisition sub-models, and different target curve parameters correspond to different feature acquisition sub-models; the step of inputting multiple target phoneme information into the pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model includes: inputting multiple target phoneme information into the curve parameter acquisition model, obtaining multiple target phoneme features through the multiple feature acquisition sub-models; and determining multiple target curve parameters based on the multiple target phoneme features.

[0141] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein the curve parameter acquisition model further includes a feature fusion sub-model, the output of the feature fusion sub-model being coupled to the inputs of multiple feature acquisition sub-models respectively; the method further includes: fusing multiple target phoneme information through the feature fusion sub-model to obtain target fusion features; the step of obtaining multiple target phoneme features through multiple feature acquisition sub-models includes: inputting the target fusion features into multiple feature acquisition sub-models respectively to obtain the target phoneme features output by each feature acquisition sub-model.

[0142] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, wherein the curve parameter acquisition model further includes a splicing sub-model, the output of the splicing sub-model being coupled to the input of the feature fusion sub-model; the method further includes: splicing multiple target phoneme information through the splicing sub-model to obtain target splicing features; the step of fusing multiple target phoneme information through the feature fusion sub-model to obtain target fusion features includes: fusing multiple target splicing features through the feature fusion sub-model to obtain the target fusion features.

[0143] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 1, wherein the curve parameter acquisition model is pre-generated by the following method: acquiring multiple sample sets, the sample sets including sample animation curves corresponding to sample animations and multiple sample phoneme information, the sample phoneme information including sample phonemes and phoneme time information of the sample phonemes; training a target neural network model through the multiple sample sets to obtain the curve parameter acquisition model.

[0144] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein training a target neural network model using multiple sample sets to obtain the curve parameter acquisition model includes: determining a current sample set from the multiple sample sets; iteratively executing model training steps based on the current sample set until the trained target neural network model meets a preset stopping iteration condition, and using the trained target neural network model as the curve parameter acquisition model; the model training steps include: inputting multiple sample phoneme information corresponding to the current sample set into the target neural network model, determining multiple estimated curve parameters corresponding to the sample animation of the current sample set through the target neural network model; generating an estimated animation curve based on the multiple estimated curve parameters; determining a target loss value based on the estimated animation curve and the sample animation curve; if the target neural network does not meet the preset stopping iteration condition based on the target loss value, updating the parameters of the target neural network model based on the target loss value to obtain a trained target neural network model, using the trained target neural network model as a new target neural network model, and determining a new current sample set from the multiple sample sets.

[0145] According to one or more embodiments of this disclosure, Example 7 provides a method of any one of Examples 1-6, the method further comprising: acquiring an animation image of the target animation; and generating the target animation based on the target animation curve and the animation image.

[0146] According to one or more embodiments of this disclosure, Example 8 provides an apparatus for generating an animation curve, comprising: a first acquisition module, configured to acquire target animation text and target animation audio of a target animation to be generated; a determination module, configured to determine, based on the target animation text and the target animation audio, a plurality of target phoneme information corresponding to the target animation text, the target phoneme information including target phonemes and phoneme time information of the target phonemes; a second acquisition module, configured to input the plurality of target phoneme information into a pre-generated curve parameter acquisition model to acquire a plurality of target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model; and a first generation module, configured to generate a target animation curve based on the plurality of target curve parameters, the target animation curve being used to generate the target animation.

[0147] According to one or more embodiments of this disclosure, Example 9 provides the apparatus of Example 8, wherein the curve parameter acquisition model includes multiple feature acquisition sub-models, and different target curve parameters correspond to different feature acquisition sub-models; the second acquisition module is further configured to: input multiple target phoneme information into the curve parameter acquisition model, acquire multiple target phoneme features through the multiple feature acquisition sub-models; and determine multiple target curve parameters based on the multiple target phoneme features.

[0148] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 9, wherein the curve parameter acquisition model further includes a feature fusion sub-model, the output of which is coupled to the input of a plurality of feature acquisition sub-models respectively; the apparatus further includes: a fusion module, configured to perform fusion processing on the plurality of target phoneme information through the feature fusion sub-model to obtain target fusion features; the second acquisition module is further configured to: input the target fusion features into the plurality of feature acquisition sub-models respectively to obtain the target phoneme features output by each feature acquisition sub-model.

[0149] According to one or more embodiments of this disclosure, Example 11 provides the apparatus of Example 10, wherein the curve parameter acquisition model further includes a splicing sub-model, the output of the splicing sub-model being coupled to the input of the feature fusion sub-model; the apparatus further includes: a splicing module, configured to splice multiple target phoneme information through the splicing sub-model to obtain target spliced ​​features; the fusion module is further configured to: fuse multiple target spliced ​​features through the feature fusion sub-model to obtain target fused features.

[0150] According to one or more embodiments of this disclosure, Example 12 provides the apparatus of Example 8, wherein the curve parameter acquisition model is pre-generated by: acquiring multiple sample sets, the sample sets including sample animation curves corresponding to sample animations and multiple sample phoneme information, the sample phoneme information including sample phonemes and phoneme time information of the sample phonemes; training a target neural network model using the multiple sample sets to obtain the curve parameter acquisition model.

[0151] According to one or more embodiments of this disclosure, Example 13 provides an apparatus of Example 12, wherein training a target neural network model using multiple sample sets to obtain the curve parameter acquisition model includes: determining a current sample set from the multiple sample sets; iteratively executing model training steps based on the current sample set until the trained target neural network model meets a preset stopping iteration condition, and using the trained target neural network model as the curve parameter acquisition model; the model training steps include: inputting multiple sample phoneme information corresponding to the current sample set into the target neural network model, determining multiple estimated curve parameters corresponding to the sample animation of the current sample set through the target neural network model; generating an estimated animation curve based on the multiple estimated curve parameters; determining a target loss value based on the estimated animation curve and the sample animation curve; if the target neural network does not meet the preset stopping iteration condition based on the target loss value, updating the parameters of the target neural network model based on the target loss value to obtain a trained target neural network model, using the trained target neural network model as a new target neural network model, and determining a new current sample set from the multiple sample sets.

[0152] According to one or more embodiments of the present disclosure, Example 14 provides an apparatus of any one of Examples 8-13, the apparatus further comprising: a third acquisition module for acquiring an animation image of the target animation; and a second generation module for generating the target animation based on the target animation curve and the animation image.

[0153] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0154] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0155] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A method for generating animation curves, characterized in that, include: Obtain the target animation text and target animation audio of the target animation to be generated; Based on the target animation text and the target animation audio, determine multiple target phoneme information corresponding to the target animation text, wherein the target phoneme information includes the target phoneme and the phoneme time information of the target phoneme; Multiple target phoneme information are input into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model; wherein, the curve parameter acquisition model includes a splicing sub-model, a feature fusion sub-model, multiple feature acquisition sub-models, and multiple parameter generation sub-models; the output of the splicing sub-model is coupled to the input of the feature fusion sub-model; the output of the feature fusion sub-model is coupled to the input of each of the multiple feature acquisition sub-models; the output of each feature acquisition sub-model is coupled to the input of the corresponding parameter generation sub-model; the target curve parameters are output through the parameter generation sub-model; For each target phoneme information, a phoneme animation curve corresponding to that target phoneme information is generated based on multiple target curve parameters corresponding to that target phoneme information. Connect the phoneme animation curves corresponding to each target phoneme information to obtain the target animation curve, which is used to generate the target animation.

2. The method according to claim 1, characterized in that, Different target curve parameters correspond to different feature acquisition sub-models; the step of inputting multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model includes: The target phoneme information is input into the curve parameter acquisition model, and multiple target phoneme features are obtained through multiple feature acquisition sub-models. Based on the multiple target phoneme features, determine the multiple target curve parameters corresponding to each target phoneme information.

3. The method according to claim 2, characterized in that, The method further includes: The target phoneme information is fused using the feature fusion sub-model to obtain the target fusion feature; The step of obtaining multiple target phoneme features through multiple feature acquisition sub-models includes: The target fusion features are respectively input into multiple feature acquisition sub-models to obtain the target phoneme features output by each feature acquisition sub-model.

4. The method according to claim 3, characterized in that, The method further includes: The target phoneme information is spliced ​​together using the splicing sub-model to obtain the target splicing feature. The process of fusing multiple target phoneme information through the feature fusion sub-model to obtain target fusion features includes: The target fusion feature is obtained by fusing multiple target splicing features through the feature fusion sub-model.

5. The method according to claim 1, characterized in that, The curve parameter acquisition model is pre-generated using the following method: Multiple sample sets are acquired, the sample sets including sample animation curves corresponding to sample animations and multiple sample phoneme information, the sample phoneme information including sample phonemes and phoneme time information of the sample phonemes; The target neural network model is trained using multiple sample sets to obtain the curve parameter acquisition model.

6. The method according to claim 5, characterized in that, The step of training the target neural network model using multiple sample sets to obtain the curve parameter acquisition model includes: Determine the current sample set from the multiple sample sets; The model training steps are executed iteratively according to the current sample set until the trained target neural network model meets the preset stopping iteration condition. The trained target neural network model is then used as the curve parameter to obtain the model. The model training steps include: The sample phoneme information corresponding to the current sample set is input into the target neural network model, and the target neural network model is used to determine the multiple predicted curve parameters corresponding to the sample animation of the current sample set. Generate a predicted animation curve based on multiple predicted curve parameters; The target loss value is determined based on the predicted animation curve and the sample animation curve; If the target neural network does not meet the preset stopping iteration condition based on the target loss value, the parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model. The trained target neural network model is then used as the new target neural network model, and a new current sample set is determined from the multiple sample sets.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtain the animated image of the target animation; The target animation is generated based on the target animation curve and the animation image.

8. An apparatus for generating animation curves, characterized in that, include: The first acquisition module is used to acquire the target animation text and target animation audio of the target animation to be generated; The determining module is used to determine multiple target phoneme information corresponding to the target animation text based on the target animation text and the target animation audio, wherein the target phoneme information includes the target phoneme and the phoneme time information of the target phoneme; The second acquisition module is used to input multiple target phoneme information into a pre-generated curve parameter acquisition model to obtain multiple target curve parameters corresponding to each target phoneme information output by the curve parameter acquisition model; wherein, the curve parameter acquisition model includes a splicing sub-model, a feature fusion sub-model, multiple feature acquisition sub-models, and multiple parameter generation sub-models; the output of the splicing sub-model is coupled to the input of the feature fusion sub-model; the output of the feature fusion sub-model is coupled to the input of each of the multiple feature acquisition sub-models; the output of each feature acquisition sub-model is coupled to the input of the corresponding parameter generation sub-model; and the target curve parameters are output through the parameter generation sub-model. The first generation module is used to generate a phoneme animation curve corresponding to each target phoneme information based on multiple target curve parameters corresponding to the target phoneme information; and to connect the phoneme animation curves corresponding to each target phoneme information to obtain a target animation curve, wherein the target animation curve is used to generate the target animation.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method described in any one of claims 1-7.

10. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-7.