Method, apparatus, medium and electronic device for determining lip animation parameters

CN116168124BActive Publication Date: 2026-08-07BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2023-01-29
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

其中,第一种是基于规则的方式,在基于规则得到唇部动画参数对应的原始曲线后,通常会进行平滑处理,而平滑处理是针对原始曲线的全局,即针对原始曲线的每个部分都做一样的平滑处理,因此,在局部细节上会存在损失

Benefits of technology

[0008]通过上述技术方案,由于将基于预设的规则数据库确定的第一唇部动画参数作为预测模型的一部分输入,使得模型的训练难度降低,从而提升模型泛化性;此外,模型可以基于国际音标序列的上下文信息,利用训练数据学习到的更好的卷积参数对第一唇部参数进行优化,得到第二唇部动画参数,从而解决相关技术中采用平滑处理时导致的局部细节上存在损失的问题,提升了基于第二唇部动画参数渲染目标虚拟人物所得到的动画的真实效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168124B_ABST
    Figure CN116168124B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, medium and electronic equipment for determining lip animation parameters, the method comprising: obtaining a phoneme sequence corresponding to a target text and a phoneme duration corresponding to each phoneme in the phoneme sequence; determining an international phonetic alphabet sequence according to the phoneme sequence, the phoneme duration corresponding to each phoneme and a preset phoneme phonetic alphabet correspondence relationship; determining first lip animation parameters according to the international phonetic alphabet sequence and a preset rule database; and processing the international phonetic alphabet sequence and the first lip animation parameters through a prediction model to obtain second lip animation parameters, wherein the second lip animation parameters are used to render a target virtual character, improve the generalization of the model, solve the problem of loss of local details in smoothing processing in related technologies, and improve the realistic effect of the animation obtained by rendering the target virtual character based on the second lip animation parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of electronic information technology, and more specifically, to a method, apparatus, medium, and electronic device for determining lip animation parameters. Background Technology

[0002] In related technologies, there are generally two methods to determine lip animation parameters, which can be represented by curves. The first is a rule-based approach. After obtaining the original curves corresponding to the lip animation parameters based on rules, smoothing is usually performed. However, this smoothing is applied globally to every part of the original curve, resulting in a loss of local detail. The second method is based on neural networks. However, neural networks require a large amount of training data; otherwise, they are prone to overfitting, leading to poor stability and potentially unreasonable predictions. Summary of the Invention

[0003] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, this disclosure provides a method for determining lip animation parameters, including: Obtain the phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; The International Phonetic Alphabet (IPA) sequence is determined based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and the preset phoneme-phoneme correspondence. The first lip animation parameters are determined based on the International Phonetic Alphabet sequence and a preset rule database; The International Phonetic Alphabet sequence and the first lip animation parameters are processed by a prediction model to obtain the second lip animation parameters, which are used to render the target virtual character.

[0005] Secondly, this disclosure provides an apparatus for determining lip animation parameters, comprising: The acquisition module is used to acquire the phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; The first determining module is used to determine the International Phonetic Alphabet sequence based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and the preset phoneme phonetic symbol correspondence. The second determining module is used to determine the first lip animation parameters based on the International Phonetic Alphabet sequence and a preset rule database; The third determining module is used to process the International Phonetic Alphabet sequence and the first lip animation parameters through a prediction model to obtain the second lip animation parameters, wherein the second lip animation parameters are used to render the target virtual character.

[0006] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.

[0007] Fourthly, this disclosure provides an electronic device, comprising: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method described in the first aspect.

[0008] By using the above technical solution, the training difficulty of the model is reduced and the generalization ability of the model is improved by using the first lip animation parameters determined based on the preset rule database as part of the input of the prediction model. In addition, the model can optimize the first lip parameters based on the contextual information of the International Phonetic Alphabet sequence and the better convolution parameters learned from the training data to obtain the second lip animation parameters. This solves the problem of loss of local details caused by smoothing in related technologies and improves the realism of the animation obtained by rendering the target virtual character based on the second lip animation parameters.

[0009] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating a method for determining lip animation parameters according to an exemplary embodiment.

[0011] Figure 2 This is a schematic diagram illustrating a rule database according to an exemplary embodiment.

[0012] Figure 3 This is a schematic diagram illustrating a process for determining second lip animation parameters according to an exemplary embodiment.

[0013] Figure 4 This is a block diagram illustrating an apparatus for determining lip animation parameters according to an exemplary embodiment.

[0014] Figure 5 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0017] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] All actions involving the acquisition of signals, information, or data in this disclosure are carried out in accordance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.

[0022] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0023] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0026] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0027] As mentioned in the background section, smoothing in rule-based prediction of lip animation parameters relies on filtering algorithms, such as mean filtering. However, these algorithms are global, applying the same filtering to every part of each original curve, ignoring the contextual differences between different texts. This leads to the use of the same filtering parameters for the same phoneme or IPA symbol in different original curves, resulting in loss of local detail. In contrast, neural networks can be viewed as a set of superimposed filter banks that consider the contextual information of the text. By learning from the data, the parameters of the filter banks can be obtained, thus solving the problem of loss of local detail in rule-based prediction. However, neural networks require a large amount of data to avoid overfitting, and their stability is relatively poor, easily leading to unreasonable predictions.

[0028] In view of this, embodiments of the present disclosure provide a method, apparatus, medium, and electronic device for determining lip animation parameters, which combines the advantages of determining lip animation parameters based on rules and determining lip animation parameters based on neural networks, improves the generalization of the model, and solves the problem of loss in local details, thereby being able to improve the real effect of the animation obtained by rendering the target virtual character based on the lip animation parameters.

[0029] Figure 1 is a flowchart of a method for determining lip animation parameters shown according to an exemplary embodiment. This method can be applied to an electronic device, such as a mobile terminal like a smart phone or a tablet computer, or a fixed terminal like a server or a desktop computer. Referring to Figure 1 , the method includes the following steps: Step S101, obtain a phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence.

[0030] Among them, the target text can be, for example, the text corresponding to the subtitles extracted from a video, or the text corresponding to the speech in an audio, etc.

[0031] Among them, the target text can be stored locally by the electronic device, or obtained by the electronic device from other devices.

[0032] In some embodiments, TTS (Text To Speech) technology can be used to obtain the phoneme sequence corresponding to the target text and the phoneme duration corresponding to each phoneme in the phoneme sequence. Exemplarily, a text-phoneme conversion model can be obtained by training a neural network. The text-phoneme conversion model can take the target text as input and obtain the phoneme sequence output by the text-phoneme conversion model and the phoneme duration corresponding to each phoneme in the phoneme sequence. The training method of the text-phoneme conversion model can refer to related technologies, and this embodiment will not elaborate here.

[0033] It should be noted that the phoneme duration of a phoneme is used to represent the pronunciation duration of the phoneme.

[0034] Exemplarily, taking the text "Hello everyone" as an example, its corresponding phoneme sequence is ["d", "a", "j", "ia", "h", "ao"].

[0035] Step S102, determine an international phonetic alphabet sequence according to the phoneme sequence, the phoneme duration corresponding to each phoneme, and a preset phoneme-phonetic symbol correspondence.

[0036] It is worth noting that the preset phoneme-phoneme correspondence is used to maintain the correspondence between phonemes and national standard phonetic symbols. Among them, the International Phonetic Alphabet (IPA) is a system for phonetic transcription, based on the Latin alphabet, designed by the International Phonetic Association as a standardized representation of spoken sounds. Phonemes correspond one-to-one with IPA symbols, and IPA symbols are smaller units of speech than phonemes.

[0037] It is worth noting that in the International Phonetic Alphabet (IPA) sequence, the total number of all IPA symbols corresponding to a phoneme is the same as the number of frames corresponding to that phoneme. The number of frames for a phoneme can be determined by the phoneme duration and the transmission time per second.

[0038] Step S103: Determine the first lip animation parameters based on the International Phonetic Alphabet sequence and a preset rule database.

[0039] It's worth noting that the preset rule database includes each International Phonetic Alphabet (IPA) symbol and its corresponding lip animation parameters. These parameters represent the weight of different preset expressions, and based on these weights, the lip shape used to pronounce the corresponding IPA symbol can be depicted. (See reference...) Figure 2 The rule database shown establishes a mapping relationship between each International Phonetic Alphabet (IPA) symbol and its corresponding lip animation parameters. For example, using... Figure 2 For the International Phonetic Alphabet symbol "t", the corresponding lip animation parameters are [0.1,0.2,0.1,0.4], where "0.1", "0.2", "0.1", and "0.4" in the lip animation parameters are used to represent the weight of different expressions.

[0040] Here, the first lip animation parameter can be represented by a tensor. This first lip animation parameter is used to characterize the combination of lip animation parameters corresponding to all International Phonetic Symbols in the International Phonetic Symbol sequence. The preset expression can be a crying expression, a laughing expression, etc.

[0041] For example, step S103 above can be implemented in the following way: based on a preset rule database, find the lip animation parameters corresponding to each International Phonetic Alphabet symbol in the preset rule database, and construct the first lip animation parameters based on the lip animation parameters corresponding to all International Phonetic Alphabet symbols in the found International Phonetic Alphabet sequence.

[0042] Step S104: The International Phonetic Alphabet sequence and the first lip animation parameters are processed by the prediction model to obtain the second lip animation parameters, which are used to render the target virtual character.

[0043] The second lip animation parameter is used to render the target virtual character, thereby driving the lips of the target virtual character to move according to the second lip animation parameter, thus generating the corresponding animation.

[0044] It is worth noting that the second lip animation parameter can be represented by a curve, which is used to represent the facial expression changes in the sequence frames. Here, the sequence frames can represent the International Phonetic Alphabet sequence.

[0045] The training process of the prediction model can be referred to in the following relevant embodiments, which will not be elaborated here.

[0046] By using the above technical solution, the training difficulty of the model is reduced and the generalization ability of the model is improved by using the first lip animation parameters determined based on the preset rule database as part of the input of the prediction model. In addition, the model can optimize the first lip parameters based on the contextual information of the International Phonetic Alphabet sequence and the better convolution parameters learned from the training data to obtain the second lip animation parameters. This solves the problem of loss of local details caused by smoothing in related technologies and improves the realism of the animation obtained by rendering the target virtual character based on the second lip animation parameters.

[0047] In some embodiments, step S102 described above can be implemented in the following manner: determining the International Phonetic Alphabet (IPA) symbol corresponding to each phoneme according to a preset phoneme-phoneme correspondence; determining an initial IPA symbol sequence based on the IPA symbols corresponding to all phonemes; determining the number of frames for each IPA symbol corresponding to each phoneme based on the phoneme duration and the number of frames transmitted per second; and expanding the number of IPA symbols corresponding to each phoneme in the initial IPA symbol sequence based on the number of frames for each IPA symbol corresponding to each phoneme to obtain an IPA symbol sequence.

[0048] The phoneme-phoneme correspondence is used to maintain the mapping between phonemes and the International Phonetic Alphabet (IPA). This correspondence can be stored in a database as key-value pairs, with phonemes as keys and IPA symbols as values.

[0049] Continuing with the example of the phoneme sequence above, the corresponding IPA symbol for the phoneme "d" is "t"; for the phoneme "a", it is "a"; and for the phoneme "j", it is "t". For the phoneme “ia”, its corresponding IPA symbols are “i” and “a”; for the phoneme “h”, its corresponding IPA symbol is “x”; for the phoneme “ao”, its corresponding IPA symbols are “ɑ” and “ʊ”.

[0050] It is worth noting that in the initial IPA sequence, the order of the IPA symbols corresponding to each phoneme is the same as the order of the phonemes in the phoneme sequence. For example, if the first phoneme precedes the second phoneme, in the initial IPA sequence, the IPA symbol corresponding to the first phoneme precedes the IPA symbol corresponding to the second phoneme. Both the first and second phonemes are phonemes within the phoneme sequence.

[0051] In the initial IPA sequence, when a phoneme corresponds to multiple IPA symbols, the order of these symbols is the same as the order in which they are stored in the phoneme-symbol correspondence. For example, taking the phoneme "ao," whose corresponding IPA symbols in the phoneme-symbol correspondence are "ɑ" and "ʊ," the IPA symbol "ɑ" will also precede the IPA symbol "ʊ" in the initial IPA sequence.

[0052] Following the example of the phoneme sequence above, the corresponding initial International Phonetic Alphabet sequence is: ["t", "a", "t"] ”, “i”, “a”, “x”, “ɑ”, “ʊ”].

[0053] This process involves first allocating the phoneme duration corresponding to each phoneme to each corresponding International Phonetic Alphabet (IPA). Then, the number of frames for each IPA is determined based on the allocated duration and the frames per second (FPS). FPS describes the number of frames in an animation or video, and the product of the allocated duration and FPS is the number of frames for each IPA. For example, FPS could be 60 frames, 120 frames, etc. The implementation process of allocating duration to each IPA based on the phoneme duration can be found in the following related embodiments, which will not be elaborated upon here.

[0054] Specifically, the number of each type of IPA symbol corresponding to a phoneme in the initial IPA symbol sequence is expanded according to the number of frames corresponding to each IPA symbol for that phoneme. Finally, the IPA symbol sequence is determined based on the expansion results of all phonemes, and the number of each type of IPA symbol corresponding to a phoneme in the IPA symbol sequence is the same as the number of frames corresponding to that type of IPA symbol for that phoneme.

[0055] It is worth noting that, depending on the number of frames for each IPA symbol corresponding to a phoneme, the corresponding IPA symbol can be increased by 0 compared to the initial IPA symbol sequence, or the corresponding IPA symbol can be increased by at least 1 compared to the initial IPA symbol sequence. Whether or not an increase is made and by how much, depends on the number of frames for each IPA symbol corresponding to the phoneme.

[0056] Continuing with the example of the phoneme sequence above, the number of frames for each IPA symbol corresponding to each phoneme in the phoneme sequence is 1, 2, 1, 4, 1, 6 respectively. For the IPA symbol "t" corresponding to the phoneme "d", the number of IPA symbols "t" in the IPA sequence is 1, which is 0 more than in the initial IPA sequence. For the IPA symbols "i" and "a" corresponding to the phonemes "ia", the total number of IPA symbols "i" and "a" in the IPA sequence is 4, which is 1 more than in the initial IPA sequence. The IPA symbols corresponding to other phonemes are similar. Based on the principle that the total number of IPA symbols corresponding to each phoneme in the IPA sequence is the same as the number of frames for that type of IPA symbol for that phoneme, the number of IPA symbols corresponding to each phoneme in the initial IPA sequence is expanded to obtain the IPA sequence. Taking the above phoneme sequence and the number of frames for each IPA symbol corresponding to each phoneme as an example, the final IPA symbol sequence is ["t", "a", "a", "t"]. ", "i", "i", "a", "a", "x", "ɑ", "ɑ", "ɑ", "ʊ", "ʊ", "ʊ"].

[0057] The above method yields an IPA sequence including the IPA symbols for each frame, which allows the prediction model to obtain the second lip animation parameters for rendering the target virtual character based on the IPA symbols of consecutive frames and the first lip animation parameters.

[0058] In some embodiments, an extended model can be pre-trained to determine the duration allocated to each IPA symbol corresponding to each phoneme. The following explanation uses the first allocated duration as the duration allocated to each IPA symbol corresponding to each phoneme. For example, the step of determining the number of frames for each IPA symbol corresponding to each phoneme based on the phoneme duration and the number of frames transmitted per second can be implemented as follows: input the phoneme durations of all phonemes and the initial IPA symbol sequence into the extended model to obtain the first allocated duration for each IPA symbol corresponding to each phoneme output by the extended model; determine the number of frames for each IPA symbol corresponding to each phoneme based on the first allocated duration for each IPA symbol corresponding to each phoneme and the number of frames transmitted per second.

[0059] The extended model predicts the first allocated duration of each phoneme based on the phoneme duration and the corresponding IPA symbol in the initial IPA sequence.

[0060] It is understandable that the sum of the first allocation durations of all the IPA symbols corresponding to a phoneme is equal to the duration of the phoneme itself.

[0061] In this way, the model can learn to pay attention to information from different contexts, thereby learning better network parameters, improving the allocation effect of the first allocation duration for each phoneme corresponding to each IPA symbol, and thus improving the effect of the prediction model's prediction of IPA symbol sequences. This, in turn, improves the animation effect of the target virtual character rendered using the second lip animation parameters obtained from the IPA symbol sequence.

[0062] In some embodiments, other methods may be used to determine the duration allocated to each IPA symbol corresponding to each phoneme. The following explanation uses the second allocated duration as the basis for describing this embodiment. For example, the step of determining the number of frames for each IPA symbol corresponding to each phoneme based on the phoneme duration and the number of frames transmitted per second can be implemented as follows: For a target phoneme, the phoneme duration is allocated to each IPA symbol corresponding to the target phoneme according to a preset allocation method, and the number of frames for each IPA symbol corresponding to the target phoneme is determined based on the second allocated duration and the number of frames transmitted per second; for non-target phonemes, the number of frames for the IPA symbol corresponding to the non-target phoneme is determined based on the phoneme duration and the number of frames transmitted per second.

[0063] The target phonemes are phonemes that correspond to multiple International Phonetic Alphabet symbols, such as “ia” and “ao” in the above phoneme sequence; the non-target phonemes are phonemes that correspond to one International Phonetic Alphabet symbol, such as “d”, “a”, “j” and “h” in the above phoneme sequence.

[0064] The preset allocation method can be a preset average allocation method or a preset statistical rule. The specific implementation process of allocating the phoneme duration of the target phoneme to each International Phonetic Alphabet symbol corresponding to the target phoneme based on the preset average allocation method or preset statistical rule can be referred to the following embodiment, which will not be elaborated here.

[0065] As mentioned earlier, the sum of the second allocated durations of all IPA symbols corresponding to a phoneme is equal to the phoneme duration itself. Here, a phoneme can be either a target phoneme or a non-target phoneme. The phoneme duration of the target phoneme is allocated to each IPA symbol corresponding to it according to a preset allocation method to obtain the second allocated duration for each IPA symbol corresponding to the target phoneme.

[0066] It is worth noting that when a phoneme corresponds to multiple International Phonetic Alphabet (IPA) symbols, the allocation of frames for each IPA symbol is involved, thus enabling the determination of the number of each IPA symbol based on the frame allocation. In this embodiment, the duration of the phoneme is first allocated to the corresponding IPA symbol, and then combined with the number of frames transmitted per second, the number of frames for each IPA symbol is obtained, thereby realizing the frame allocation for each IPA symbol.

[0067] It is worth noting that when a phoneme corresponds to an International Phonetic Alphabet (IPA), the allocation of phoneme duration or frame count is not involved. Instead, the number of IPA frames corresponding to that phoneme is determined directly by the phoneme duration and the number of frames transmitted per second. This type of phoneme is the non-target factor.

[0068] The above method provides a flexible duration allocation method. The phoneme duration of the target phoneme is allocated to the corresponding International Phonetic Alphabet (IPA) symbol using a preset allocation method. Then, the number of frames for each IPA symbol is determined based on the second allocation duration allocated to each IPA symbol, thereby improving the flexibility of IPA symbol frame allocation.

[0069] In some embodiments, the preset allocation method may be a preset average allocation method. For example, the phoneme duration of the target phoneme can be allocated using a preset average allocation method in the following way: that is, the step described above, for the target phoneme, allocating the phoneme duration of the target phoneme to each International Phonetic Alphabet symbol corresponding to the target phoneme according to the preset allocation method, may include: for the target phoneme, evenly allocating the phoneme duration of the target phoneme to each International Phonetic Alphabet symbol corresponding to the target phoneme according to the preset average allocation method.

[0070] Following the examples of phoneme sequences, initial IPA sequences, and corresponding frame numbers for the phoneme "ia," if the phoneme duration of "ia" is T1, then according to the preset average allocation method, the second allocation durations of IPA symbols "i" and "a" are T1 / 2 respectively. That is, the first allocation durations for the two IPA symbols corresponding to the phoneme "ia" are the same and equally distributed. Similarly, for the phoneme "ao," if the phoneme duration of "ao" is T2, then according to the preset average allocation method, the second allocation durations of IPA symbols "a" and "ʊ" are T2 / 2 respectively. That is, the first allocation durations for the two IPA symbols corresponding to the phoneme "ao" are the same and equally distributed.

[0071] Following the example above, according to the allocation result of T2 / 2 for the second allocation duration of the IPA symbol “ɑ” and IPA symbol “ʊ”, if the number of frames transmitted per second is P1, for the IPA symbols “ɑ” and “ʊ” corresponding to the phoneme “ao”, the number of frames for the IPA symbols “ɑ” and “ʊ” is the product of T2 / 2 and P1.

[0072] In some embodiments, the preset allocation method may be a preset statistical rule. For example, the phoneme duration of the target phoneme can be allocated using a preset statistical rule in the following manner: that is, the step described above, for the target phoneme, allocating the phoneme duration of the target phoneme to each International Phonetic Alphabet symbol corresponding to the target phoneme according to the preset allocation method, may include: for the target phoneme, allocating the phoneme duration of the target phoneme to each International Phonetic Alphabet symbol corresponding to the target phoneme according to the preset statistical rule.

[0073] The preset statistical rules include the duration ratio of different IPA symbols corresponding to each phoneme, and the ratio of the second allocation duration of each IPA symbol corresponding to the target phoneme to the phoneme duration of the target phoneme is the same as the duration ratio of that IPA symbol.

[0074] Following the example above, if the preset statistical rules represent that the duration of the phoneme “ao” corresponding to the IPA symbols “ɑ” and “ʊ” is 30% and 70% respectively, and if the phoneme duration of “ao” is T2, then the second allocation duration of the IPA symbol “ɑ” is 30% of T2, and the second allocation duration of the IPA symbol “ʊ” is 70% of T2.

[0075] It's worth noting that the preset statistical rules can be obtained through big data analysis. For example, using multiple data points corresponding to the same phoneme obtained from big data analysis, each data point represents the duration percentage of different International Phonetic Alphabet (IPA) symbols corresponding to that phoneme. The preset statistical rules determine the duration percentage of different IPA symbols corresponding to that phoneme based on all the statistically analyzed data. For instance, for all the statistically analyzed data points corresponding to that phoneme, the duration percentage with the largest number of occurrences is selected as the preset statistical rule for the duration percentage of different IPA symbols corresponding to that phoneme.

[0076] In some embodiments, step S104 above can be implemented in the following manner: the International Phonetic Alphabet sequence is processed by the feature extraction network of the prediction model to obtain the text feature tensor corresponding to the International Phonetic Alphabet sequence; the text feature tensor is processed by the convolutional network of the prediction model to obtain the text latent features corresponding to the text feature tensor; the text latent features and the first lip animation parameters are concatenated by the concatenation network of the prediction model to obtain the concatenated features; the concatenated features are processed by the long short-term memory network of the prediction model to obtain the target features; and the target features are processed by the fully connected layer of the prediction model to obtain the second lip animation parameters.

[0077] For example, following the example of the above International Phonetic Alphabet sequence, refer to... Figure 3As shown, the prediction model includes a feature extraction network, a convolutional network, a concatenation network, a long short-term memory (LSTM) network, and a fully connected layer. The feature extraction network and convolutional network can be collectively referred to as the encoder network. Several long short-term memory (LSTM) networks can be included. The output of the last LSTM layer is the target feature, the output of the previous LSTM layer is the input of the next LSTM layer, and the input of the first LSTM layer is the output of the concatenation network, i.e., the concatenated feature.

[0078] exist Figure 3 In this process, the International Phonetic Alphabet (IPA) sequence is input into a feature extraction network to obtain a text feature tensor output by the feature extraction network. This text feature tensor is then used as input to a convolutional network, which processes it to obtain the latent text features output by the convolutional network. Figure 2 The rule database shown is used to find the lip animation parameters corresponding to each International Phonetic Alphabet (IPA) symbol in the sequence. The lip animation parameters corresponding to all IPA symbols are combined to obtain the first lip animation parameters. The first lip animation parameters and textual latent features are concatenated into a concatenation network to obtain the concatenated features output by the concatenation network. The concatenated features are then input into a Long Short-Term Memory (LSTM) network to obtain the target features output by the LSM network. Finally, the target features are processed through a fully connected layer to obtain the second lip animation parameters.

[0079] By combining the advantages of rule-based and neural network-based lip animation parameter determination, the generalization ability of the model is improved, and the problem of loss in local details is solved. This can enhance the realism of the animation obtained by rendering the target virtual character based on lip animation parameters.

[0080] The loss function of the prediction model can be represented by the following equation (1): (1); Where y represents the lip animation parameters corresponding to the training samples. The loss is used to predict lip animation parameters.

[0081] Among them, it can be characterized by the following formula (2). : (2); in, Characterizing the International Phonetic Alphabet sequence of the sample, To find matching rules from the pre-defined rules database The corresponding sample's first lip animation parameters, Characterization feature extraction networks and convolutional networks, Characterizes long short-term memory networks and fully connected layers.

[0082] By using the above method, training samples and setting stopping model iteration conditions, Figure 3 The model structure shown is used for training to obtain a trained prediction model for rendering target virtual characters in practical applications.

[0083] The conditions for stopping model iteration can be that the loss value is less than a preset value, or that the number of model iterations reaches a preset threshold, etc.

[0084] The acquisition and interpretation of parameters involved in training the prediction model that have the same physical meaning as the prediction model in practical applications can be found above. Figure 1 and Figure 2 The embodiments involved are not described in detail here.

[0085] This disclosure also provides an apparatus for determining lip animation parameters, referring to... Figure 4 The apparatus for determining lip animation parameters may include: The acquisition module 401 is used to acquire the phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; The first determining module 402 is used to determine the International Phonetic Alphabet sequence based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and the preset phoneme phonetic symbol correspondence relationship; The second determining module 403 is used to determine the first lip animation parameters based on the International Phonetic Alphabet sequence and a preset rule database; The third determining module 404 is used to process the International Phonetic Alphabet sequence and the first lip animation parameters through a prediction model to obtain the second lip animation parameters, wherein the second lip animation parameters are used to render the target virtual character.

[0086] Optionally, the first determining module 402 includes: The first determining submodule is used to determine the International Phonetic Alphabet (IPA) symbol corresponding to each phoneme according to a preset phoneme-phoneme correspondence relationship, wherein the phoneme-phoneme correspondence relationship is used to maintain the correspondence between phonemes and national standard phonetic symbols; The second determining submodule is used to determine the initial international phonetic symbol sequence based on the international phonetic symbols corresponding to all the phonemes; The third determining submodule is used to determine the number of frames for each phoneme corresponding to each international phonetic symbol based on the phoneme duration and the number of frames transmitted per second for each phoneme. The fourth determining submodule is used to expand the number of IPA symbols corresponding to each phoneme in the initial IPA symbol sequence according to the number of frames of each IPA symbol corresponding to each phoneme, so as to obtain an IPA symbol sequence, wherein the number of IPA symbols corresponding to each phoneme in the IPA symbol sequence is the same as the number of frames of the IPA symbol corresponding to each phoneme.

[0087] Optionally, the third determining submodule is specifically used for: The phoneme durations of all the phonemes and the initial IPA sequence are input into the extended model to obtain the first allocated duration of each IPA corresponding to each phoneme output by the extended model. The number of frames for each phoneme corresponding to each international phoneme is determined based on the first allocated duration and the number of frames transmitted per second for each international phoneme.

[0088] Optionally, the third determining submodule is specifically used for: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset allocation method, and the number of frames corresponding to each IPA symbol corresponding to the target phoneme is determined according to the second allocation duration allocated to each IPA symbol corresponding to the target phoneme and the number of frames transmitted per second, wherein the target phoneme is a phoneme corresponding to multiple IPA symbols; For non-target phonemes, the number of frames corresponding to the International Phonetic Alphabet (IPA) is determined based on the phoneme duration and the number of frames transmitted per second of the non-target phoneme, wherein the non-target phoneme is a phoneme corresponding to one IPA symbol.

[0089] Optionally, the third determining submodule is further specifically used for: For a target phoneme, the phoneme duration is evenly distributed to each International Phonetic Alphabet symbol corresponding to the target phoneme according to a preset average distribution method.

[0090] Optionally, the third determining submodule is further specifically used for: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset statistical rule. The preset statistical rule includes the duration ratio of different IPA symbols corresponding to each phoneme. The ratio of the second allocated duration of each IPA symbol corresponding to the target phoneme to the phoneme duration of the target phoneme is the same as the duration ratio of that IPA symbol.

[0091] Optionally, the third determining module 404 includes: The first processing submodule is used to process the International Phonetic Alphabet sequence through the feature extraction network of the prediction model to obtain the text feature tensor corresponding to the International Phonetic Alphabet sequence; The second processing submodule is used to process the text feature tensor through the convolutional network of the prediction model to obtain the text latent features corresponding to the text feature tensor. The third processing submodule is used to concatenate the text latent features and the first lip animation parameters through the concatenation network of the prediction model to obtain concatenated features; The fourth processing submodule is used to process the spliced ​​features through the long short-term memory network of the prediction model to obtain the target features; The fifth processing submodule is used to process the target features through the fully connected layer of the prediction model to obtain the second lip animation parameters.

[0092] The implementation methods of each module of the above-mentioned device can be referred to the above-mentioned related embodiments, and will not be repeated here.

[0093] This disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps in the method embodiments described above.

[0094] This disclosure also provides an electronic device, including: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps in the above-described method embodiments.

[0095] The following is for reference. Figure 5 This diagram illustrates a structural schematic of an electronic device 500 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0096] like Figure 5As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data required for the operation of electronic device 500. Processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0097] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0098] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0099] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0100] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0101] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0102] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a phoneme sequence corresponding to a target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; determine an International Phonetic Alphabet (IPA) sequence based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and a preset phoneme-phoneme correspondence; determine first lip animation parameters based on the IPA sequence and a preset rule database; and process the IPA sequence and the first lip animation parameters using a prediction model to obtain second lip animation parameters, wherein the second lip animation parameters are used to render a target virtual character.

[0103] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0105] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the module itself.

[0106] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0107] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0108] According to one or more embodiments of this disclosure, Example 1 provides a method for determining lip animation parameters, including: Obtain the phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; The International Phonetic Alphabet (IPA) sequence is determined based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and the preset phoneme-phoneme correspondence. The first lip animation parameters are determined based on the International Phonetic Alphabet sequence and a preset rule database; The International Phonetic Alphabet sequence and the first lip animation parameters are processed by a prediction model to obtain the second lip animation parameters, which are used to render the target virtual character.

[0109] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein determining the International Phonetic Alphabet sequence based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and a preset phoneme-phoneme correspondence includes: Based on the preset phoneme-phoneme correspondence, the corresponding International Phonetic Alphabet (IPA) symbol for each phoneme is determined, wherein the phoneme-phoneme correspondence is used to maintain the correspondence between phonemes and national standard phonetic symbols; Determine the initial IPA sequence based on the IPA symbols corresponding to all the phonemes; The number of frames for each phoneme corresponding to each International Phonetic Alphabet is determined based on the phoneme duration and the number of frames transmitted per second for each phoneme. Based on the number of frames for each type of IPA symbol corresponding to each phoneme, the number of IPA symbols corresponding to that phoneme in the initial IPA symbol sequence is expanded to obtain an IPA symbol sequence, wherein the number of IPA symbols corresponding to each phoneme in the IPA symbol sequence is the same as the number of frames for that type of IPA symbol corresponding to that phoneme.

[0110] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein determining the number of frames for each phoneme corresponding to each International Phonetic Alphabet based on the phoneme duration and the number of frames transmitted per second includes: The phoneme durations of all the phonemes and the initial IPA sequence are input into the extended model to obtain the first allocated duration of each IPA corresponding to each phoneme output by the extended model. The number of frames for each phoneme corresponding to each international phoneme is determined based on the first allocated duration and the number of frames transmitted per second for each international phoneme.

[0111] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2, wherein determining the number of frames for each phoneme corresponding to each International Phonetic Alphabet based on the phoneme duration and the number of frames transmitted per second includes: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset allocation method, and the number of frames corresponding to each IPA symbol corresponding to the target phoneme is determined according to the second allocation duration allocated to each IPA symbol corresponding to the target phoneme and the number of frames transmitted per second, wherein the target phoneme is a phoneme corresponding to multiple IPA symbols; For non-target phonemes, the number of frames corresponding to the International Phonetic Alphabet (IPA) is determined based on the phoneme duration and the number of frames transmitted per second of the non-target phoneme, wherein the non-target phoneme is a phoneme corresponding to one IPA symbol.

[0112] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, wherein, for a target phoneme, the phoneme duration is allocated to each International Phonetic Alphabet symbol corresponding to the target phoneme according to a preset allocation method, including: For a target phoneme, the phoneme duration is evenly distributed to each International Phonetic Alphabet symbol corresponding to the target phoneme according to a preset average distribution method.

[0113] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 4, wherein, for a target phoneme, the phoneme duration is allocated to each International Phonetic Alphabet symbol corresponding to the target phoneme according to a preset allocation method, including: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset statistical rule. The preset statistical rule includes the duration ratio of different IPA symbols corresponding to each phoneme. The ratio of the second allocated duration of each IPA symbol corresponding to the target phoneme to the phoneme duration of the target phoneme is the same as the duration ratio of that IPA symbol.

[0114] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein processing the International Phonetic Alphabet sequence and the first lip animation parameters through a prediction model to obtain second lip animation parameters includes: The International Phonetic Alphabet (IPA) sequence is processed by the feature extraction network of the prediction model to obtain the text feature tensor corresponding to the IPA sequence; The text feature tensor is processed by the convolutional network of the prediction model to obtain the text latent features corresponding to the text feature tensor. The textual latent features and the first lip animation parameters are concatenated through the concatenation network of the prediction model to obtain concatenated features; The spliced ​​features are processed by the long short-term memory network of the prediction model to obtain the target features; The target features are processed by the fully connected layer of the prediction model to obtain the second lip animation parameters.

[0115] According to one or more embodiments of this disclosure, Example 8 provides a method for determining lip animation parameters, including: The acquisition module is used to acquire the phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; The first determining module is used to determine the International Phonetic Alphabet sequence based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and the preset phoneme phonetic symbol correspondence. The second determining module is used to determine the first lip animation parameters based on the International Phonetic Alphabet sequence and a preset rule database; The third determining module is used to process the International Phonetic Alphabet sequence and the first lip animation parameters through a prediction model to obtain the second lip animation parameters, wherein the second lip animation parameters are used to render the target virtual character.

[0116] According to one or more embodiments of this disclosure, Example 9 provides the apparatus of Example 8, wherein the first determining module includes: The first determining submodule is used to determine the International Phonetic Alphabet (IPA) symbol corresponding to each phoneme according to a preset phoneme-phoneme correspondence relationship, wherein the phoneme-phoneme correspondence relationship is used to maintain the correspondence between phonemes and national standard phonetic symbols; The second determining submodule is used to determine the initial international phonetic symbol sequence based on the international phonetic symbols corresponding to all the phonemes; The third determining submodule is used to determine the number of frames for each phoneme corresponding to each international phonetic symbol based on the phoneme duration and the number of frames transmitted per second for each phoneme. The fourth determining submodule is used to expand the number of IPA symbols corresponding to each phoneme in the initial IPA symbol sequence according to the number of frames of each IPA symbol corresponding to each phoneme, so as to obtain an IPA symbol sequence, wherein the number of IPA symbols corresponding to each phoneme in the IPA symbol sequence is the same as the number of frames of the IPA symbol corresponding to each phoneme.

[0117] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 9, wherein the third determining submodule is specifically used for: The phoneme durations of all the phonemes and the initial IPA sequence are input into the extended model to obtain the first allocated duration of each IPA corresponding to each phoneme output by the extended model. The number of frames for each phoneme corresponding to each international phoneme is determined based on the first allocated duration and the number of frames transmitted per second for each international phoneme.

[0118] According to one or more embodiments of this disclosure, Example 11 provides the apparatus of Example 9, wherein the third determining submodule is specifically used for: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset allocation method, and the number of frames corresponding to each IPA symbol corresponding to the target phoneme is determined according to the second allocation duration allocated to each IPA symbol corresponding to the target phoneme and the number of frames transmitted per second, wherein the target phoneme is a phoneme corresponding to multiple IPA symbols; For non-target phonemes, the number of frames corresponding to the International Phonetic Alphabet (IPA) is determined based on the phoneme duration and the number of frames transmitted per second of the non-target phoneme, wherein the non-target phoneme is a phoneme corresponding to one IPA symbol.

[0119] According to one or more embodiments of this disclosure, Example 12 provides the apparatus of Example 11, wherein the third determining submodule is further specifically configured to: For a target phoneme, the phoneme duration is evenly distributed to each International Phonetic Alphabet symbol corresponding to the target phoneme according to a preset average distribution method.

[0120] According to one or more embodiments of this disclosure, Example 13 provides the apparatus of Example 11, wherein the third determining submodule is further specifically configured to: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset statistical rule. The preset statistical rule includes the duration ratio of different IPA symbols corresponding to each phoneme. The ratio of the second allocated duration of each IPA symbol corresponding to the target phoneme to the phoneme duration of the target phoneme is the same as the duration ratio of that IPA symbol.

[0121] According to one or more embodiments of this disclosure, Example 14 provides the apparatus of Example 8, wherein the third determining module includes: The first processing submodule is used to process the International Phonetic Alphabet sequence through the feature extraction network of the prediction model to obtain the text feature tensor corresponding to the International Phonetic Alphabet sequence; The second processing submodule is used to process the text feature tensor through the convolutional network of the prediction model to obtain the text latent features corresponding to the text feature tensor. The third processing submodule is used to concatenate the text latent features and the first lip animation parameters through the concatenation network of the prediction model to obtain concatenated features; The fourth processing submodule is used to process the spliced ​​features through the long short-term memory network of the prediction model to obtain the target features; The fifth processing submodule is used to process the target features through the fully connected layer of the prediction model to obtain the second lip animation parameters.

[0122] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-7.

[0123] According to one or more embodiments of this disclosure, Example 16 provides an electronic device comprising: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of any one of the methods described in Examples 1-7.

[0124] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0125] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0126] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A method for determining lip animation parameters, characterized in that, include: Obtain the phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; The International Phonetic Alphabet (IPA) sequence is determined based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and the preset phoneme-phoneme correspondence. The first lip animation parameters are determined based on the International Phonetic Alphabet sequence and a preset rule database; The International Phonetic Alphabet sequence and the first lip animation parameters are processed by a prediction model to obtain the second lip animation parameters, wherein the second lip animation parameters are used to render the target virtual character; The step of determining the International Phonetic Alphabet (IPA) sequence based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and a preset phoneme-phoneme correspondence includes: determining the IPA corresponding to each phoneme based on the preset phoneme-phoneme correspondence, wherein the phoneme-phoneme correspondence is used to maintain the correspondence between phonemes and GBP (GBP); determining an initial IPA sequence based on the IPA corresponding to all phonemes; determining the number of frames for each IPA corresponding to each phoneme based on the phoneme duration and the number of frames transmitted per second; and expanding the number of IPAs of that type corresponding to the phoneme in the initial IPA sequence based on the number of frames for each IPA corresponding to the phoneme to obtain the IPA sequence, wherein the number of IPAs of that type corresponding to the phoneme in the IPA sequence is the same as the determined number of frames for that type of IPA corresponding to the phoneme.

2. The method according to claim 1, characterized in that, The step of determining the number of frames for each phoneme corresponding to each International Phonetic Alphabet based on the phoneme duration and the number of frames transmitted per second includes: The phoneme durations of all the phonemes and the initial IPA sequence are input into the extended model to obtain the first allocated duration of each IPA corresponding to each phoneme output by the extended model. The number of frames for each phoneme corresponding to each international phoneme is determined based on the first allocated duration and the number of frames transmitted per second for each international phoneme.

3. The method according to claim 1, characterized in that, The step of determining the number of frames for each phoneme corresponding to each International Phonetic Alphabet based on the phoneme duration and the number of frames transmitted per second includes: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset allocation method, and the number of frames corresponding to each IPA symbol corresponding to the target phoneme is determined according to the second allocation duration allocated to each IPA symbol corresponding to the target phoneme and the number of frames transmitted per second, wherein the target phoneme is a phoneme corresponding to multiple IPA symbols; For non-target phonemes, the number of frames corresponding to the International Phonetic Alphabet (IPA) is determined based on the phoneme duration and the number of frames transmitted per second of the non-target phoneme, wherein the non-target phoneme is a phoneme corresponding to one IPA symbol.

4. The method according to claim 3, characterized in that, The step of allocating the phoneme duration of the target phoneme to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset allocation method includes: For a target phoneme, the phoneme duration is evenly distributed to each International Phonetic Alphabet symbol corresponding to the target phoneme according to a preset average distribution method.

5. The method according to claim 3, characterized in that, The step of allocating the phoneme duration of the target phoneme to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset allocation method includes: For a target phoneme, the phoneme duration of the target phoneme is allocated to each International Phonetic Alphabet (IPA) symbol corresponding to the target phoneme according to a preset statistical rule. The preset statistical rule includes the duration ratio of different IPA symbols corresponding to each phoneme. The ratio of the second allocated duration of each IPA symbol corresponding to the target phoneme to the phoneme duration of the target phoneme is the same as the duration ratio of that IPA symbol.

6. The method according to claim 1, characterized in that, The step of processing the International Phonetic Alphabet sequence and the first lip animation parameters using a prediction model to obtain the second lip animation parameters includes: The International Phonetic Alphabet (IPA) sequence is processed by the feature extraction network of the prediction model to obtain the text feature tensor corresponding to the IPA sequence; The text feature tensor is processed by the convolutional network of the prediction model to obtain the text latent features corresponding to the text feature tensor. The textual latent features and the first lip animation parameters are concatenated through the concatenation network of the prediction model to obtain concatenated features; The spliced ​​features are processed by the long short-term memory network of the prediction model to obtain the target features; The target features are processed by the fully connected layer of the prediction model to obtain the second lip animation parameters.

7. A device for determining lip animation parameters, characterized in that, include: The acquisition module is used to acquire the phoneme sequence corresponding to the target text, and the phoneme duration corresponding to each phoneme in the phoneme sequence; The first determining module is used to determine the International Phonetic Alphabet sequence based on the phoneme sequence, the phoneme duration corresponding to each phoneme, and the preset phoneme phonetic symbol correspondence. The second determining module is used to determine the first lip animation parameters based on the International Phonetic Alphabet sequence and a preset rule database; The third determining module is used to process the International Phonetic Alphabet sequence and the first lip animation parameters through a prediction model to obtain the second lip animation parameters, wherein the second lip animation parameters are used to render the target virtual character; The first determining module includes: The first determining submodule is used to determine the International Phonetic Alphabet (IPA) symbol corresponding to each phoneme according to a preset phoneme-phoneme correspondence relationship, wherein the phoneme-phoneme correspondence relationship is used to maintain the correspondence between phonemes and national standard phonetic symbols; The second determining submodule is used to determine the initial international phonetic symbol sequence based on the international phonetic symbols corresponding to all the phonemes; The third determining submodule is used to determine the number of frames for each phoneme corresponding to each international phonetic symbol based on the phoneme duration and the number of frames transmitted per second for each phoneme. The fourth determining submodule is used to expand the number of IPA symbols corresponding to each phoneme in the initial IPA symbol sequence according to the number of frames of each IPA symbol corresponding to each phoneme, so as to obtain an IPA symbol sequence, wherein the number of IPA symbols corresponding to each phoneme in the IPA symbol sequence is the same as the number of frames of the IPA symbol corresponding to each phoneme.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-6.

9. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Animation generation method and device, storage medium and electronic equipment

    CN113902838A

  • KR20220112422A