Method, device and computer program product for predicting pronunciation duration of lyrics phonemes

By combining the trained phoneme pronunciation duration prediction model with the coding information, the pronunciation duration of each phoneme in the cover song is adaptively predicted, which solves the problem of inaccurate phoneme duration in cover songs and improves the accuracy of phoneme duration prediction and the quality of cover songs.

CN115171661BActive Publication Date: 2025-09-12TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210723111.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-09-12
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

Existing phoneme duration prediction methods cannot accurately predict the duration of each phoneme in a fixed beat in cover song synthesis, resulting in poor song cover effects.

Method used

By determining the number of words in the lyrics to be replaced and the target lyrics in the original song, and using the trained phoneme pronunciation duration prediction model, combined with the encoding information of phonemes, phoneme types and actual pronunciation duration of syllables, the pronunciation duration of each phoneme is adaptively predicted to meet the duration constraints of the cover song.

Benefits of technology

The accuracy of predicting the phoneme duration of each word in the cover lyrics has been improved, ensuring the accuracy of the song's duration in the original beat and improving the quality of the cover songs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171661B_ABST
    Figure CN115171661B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device and computer program product for predicting the pronunciation duration of phonemes in lyrics. The present application can predict the duration of phonemes in each word in new lyrics under the duration constraint of the original lyrics, thereby improving the accuracy of duration prediction for cover lyrics. The method comprises: determining the original lyrics and obtaining the number of words in the target lyrics; determining the actual pronunciation duration of the syllables of each word in the target lyrics based on the timestamp of each word in the original lyrics and the number of words in the target lyrics; wherein the syllable corresponding to each word in the target lyrics includes at least one phoneme; encoding the phonemes, phoneme types and actual pronunciation duration of the syllables of each word in the target lyrics respectively to obtain phoneme encoding information, phoneme type encoding information and actual pronunciation duration encoding information of the syllables; inputting the phoneme encoding information, phoneme type encoding information and actual pronunciation duration encoding information of the syllables into a phoneme pronunciation duration prediction model to obtain a pronunciation duration prediction result for each phoneme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing technology, and in particular to a method for predicting the pronunciation duration of lyric phonemes, a computer device, and a computer program product. Background Art

[0002] A song cover is a re-sing of someone else's original song in their own style, often with new lyrics or a new arrangement. Since the number of words in a cover often differs from the original, it's necessary to re-estimate the duration of each syllable within the beat during the lyric recording process.

[0003] Syllables are composed of phonemes, and predicting the duration of a syllable essentially means predicting the duration of a phoneme. Currently, the most commonly used phoneme duration prediction method is primarily used in speech synthesis. This method is unconstrained in duration, meaning that the pronunciation duration of each phoneme is unrestricted. The accuracy of pronunciation duration prediction only affects the naturalness of speech synthesis. However, in cover song synthesis, the song's tempo determines the pronunciation duration of each word. The total duration of the phonemes is fixed, and the approximate position of each phoneme within the tempo is also determined. Applying the same phoneme duration prediction method used in speech synthesis to cover song synthesis would result in inaccurate lyrics within the original tempo, resulting in poor cover performance. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, computer device and computer program product for predicting the pronunciation duration of lyrics phonemes to address the above technical problems.

[0005] In a first aspect, the present application provides a method for predicting the pronunciation duration of lyrics phonemes. The method comprises:

[0006] Determining the original lyrics to be replaced in the original song, and obtaining the word count of the target lyrics used to replace the original lyrics;

[0007] Determining the actual pronunciation duration of a syllable corresponding to each word in the target lyrics based on the timestamp of each word in the original lyrics and the number of words in the target lyrics; wherein the syllable corresponding to each word in the target lyrics includes at least one phoneme;

[0008] Encoding the phoneme of each word in the target lyrics, the phoneme type of the phoneme, and the actual pronunciation duration of the syllable to obtain phoneme encoding information, phoneme type encoding information, and actual pronunciation duration encoding information of the syllable for each word;

[0009] The phoneme coding information, phoneme type coding information and syllable actual pronunciation duration coding information are input into a trained phoneme pronunciation duration prediction model to obtain a pronunciation duration prediction result of each phoneme output by the phoneme pronunciation duration prediction model.

[0010] In a second aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned method for predicting the pronunciation duration of lyrics phonemes when executing the computer program.

[0011] In a third aspect, the present application further provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the above-mentioned method for predicting the pronunciation duration of lyrics phonemes.

[0012] The above-mentioned method, computer equipment and computer program product for predicting the pronunciation duration of lyrics phonemes determine the original lyrics to be replaced in the original song and obtain the number of words in the target lyrics used to replace the original lyrics; determine the actual pronunciation duration of the syllable corresponding to each word in the target lyrics based on the timestamp of each word in the original lyrics and the number of words in the target lyrics; wherein the syllable corresponding to each word in the target lyrics includes at least one phoneme; encode the phoneme, phoneme type and actual pronunciation duration of the syllable of each word in the target lyrics respectively to obtain the phoneme encoding information, phoneme type encoding information and actual pronunciation duration encoding information of the syllable for each word; input the phoneme encoding information, phoneme type encoding information and actual pronunciation duration encoding information of the syllable into a trained phoneme pronunciation duration prediction model to obtain the pronunciation duration prediction result of each phoneme output by the phoneme pronunciation duration prediction model. This application uses a trained phoneme pronunciation duration prediction model to predict the phoneme duration of each word (or each syllable) in the new lyrics under the duration constraint of the original lyrics, and adaptively obtains the pronunciation duration of each phoneme. Compared with the phoneme duration prediction method used in traditional speech synthesis, it can further improve the accuracy of phoneme duration prediction for each word in the cover lyrics. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A diagram illustrating an application environment of a method for predicting the pronunciation duration of phonemes in lyrics according to an embodiment;

[0014] Figure 2 1 is a flow chart of a method for predicting the pronunciation duration of lyrics phonemes in one embodiment;

[0015] Figure 3 Schematic diagram of a phoneme pronunciation duration prediction model in one embodiment;

[0016] Figure 4Schematic diagram of a training method for a phoneme pronunciation duration prediction model in one embodiment;

[0017] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0019] The intelligent scheduling method for streaming media files provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or it can be placed on the cloud or other network servers. Among them, the terminal 101 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 102 can be implemented as a server cluster consisting of multiple servers.

[0020] In one embodiment, Figure 2 As shown, a method for predicting the pronunciation duration of lyrics phonemes is provided, and the method is applied to Figure 1 Taking the server 102 in the example as an example, the following steps are included:

[0021] Step S201, determining the original lyrics to be replaced in the original song, and obtaining the number of words in the target lyrics used to replace the original lyrics.

[0022] The target lyrics refer to the new lyrics.

[0023] Specifically, first, the server 102 obtains the lyrics file of the original song, and obtains the original lyrics and the timestamp of each word in the original lyrics. The server 102 also needs to obtain the target lyrics input by the user and determine the number of words in the target lyrics based on the target lyrics.

[0024] Step S202, determining the actual pronunciation duration of the syllable corresponding to each word in the target lyrics based on the timestamp of each word in the original lyrics and the number of words in the target lyrics; wherein the syllable corresponding to each word in the target lyrics includes at least one phoneme;

[0025] Among them, the timestamp refers to the time marker information of each word in the lyrics, which is used to display the lyrics at the corresponding music beats and show them to the user. A syllable refers to the pronunciation structure corresponding to each word. Each syllable can include at least one phoneme, that is, a syllable can be composed of one phoneme or a combination of multiple phonemes. For example, in the lyrics "Olympic athletes win championships and fulfill their dreams", the syllable corresponding to "奥" is "ao", and the corresponding phoneme is "ao". The syllable corresponding to "夺" is "duo", and the corresponding phonemes are "d" and "uo".

[0026] Specifically, extract the timestamps in the lyrics file of the original song lyrics. According to the timestamps of each word in the original song lyrics and the number of words in the target lyrics, the actual pronunciation duration of the syllable corresponding to each word in the target lyrics can be determined. The actual pronunciation duration of the syllable can be understood as the pronunciation duration of the syllable of this word in the original song. For example, the first few lyrics of the well-known song "Happy New Year" are "Happy New Year, Happy New Year, Wish you all a happy new year. We sing, we dance, Wish you all a happy new year", and the target lyrics are "Beijing Winter Olympics, various competitions, are being staged passionately. Inside and outside the stadium, cheer and shout. Olympic athletes win championships and fulfill their dreams". Among them, the target lyrics "Beijing Winter Olympics" correspond to "Happy New Year" in the original song lyrics. Therefore, according to the timestamps of the four words "Happy New Year", the timestamps of each word in "Beijing Winter Olympics" can be obtained, that is, the actual pronunciation duration of the syllable corresponding to each word in the target lyrics can be obtained.

[0027] Step S203, encode the phonemes, phoneme types, and actual pronunciation durations of the syllables of each word in the target lyrics respectively to obtain the phoneme encoding information, phoneme type encoding information, and syllable actual pronunciation duration encoding information of each word.

[0028] Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech; phonemes are divided into two major categories: vowels and consonants; generally speaking, a syllable (word) is composed of a consonant and a vowel, but there may also be no consonant part. For example, the phoneme of "奥" above is "ao", and the phoneme of "运" is "vn", and they only have vowels and no consonants. The phoneme type is represented by an enumerated value in this application. For example, 1, 2, and 3 are used to represent consonants, vowels, and syllables without consonants respectively.

[0029] Specifically, as shown in Table 1 below, where PhI represents the phoneme, PhT represents the phoneme type, SyD represents the actual pronunciation duration of the syllable, and PhD represents the phoneme duration; taking the four words "Olympic athletes" in the above target lyrics as an example, it is represented by Table 1 as follows:

[0030] PhI ao vn j ian er PhT 3 3 1 2 3 SYD 200 150 250 250 160 PhD 200 150 50 200 160

[0031] Table 1

[0032] In one example, the phoneme PhI and phoneme type PhT of each of the above characters can be encoded using a one-hot encoding method to obtain the phoneme encoding information and phoneme type encoding information of each character. For example, for the phonemes of each character in the target lyrics, each phoneme can be encoded as an N*D vector, where N refers to the number of register bits in the one-hot encoding technology, and D represents the vector dimension of each phoneme. For another example, for the phoneme type of each phoneme, the phoneme type can be encoded as an N*3 vector, and the actual pronunciation duration of the syllable SyD can be encoded as a vector with a data dimension of N*1.

[0033] Step S204: input the phoneme coding information, phoneme type coding information and syllable actual pronunciation duration coding information into the trained phoneme pronunciation duration prediction model to obtain a pronunciation duration prediction result of each phoneme output by the phoneme pronunciation duration prediction model.

[0034] Specifically, this application designs a phoneme pronunciation duration prediction model, and inputs the above-mentioned phoneme encoding information PhI, phoneme type encoding information PhT and syllable actual pronunciation duration encoding information SyD into the trained phoneme pronunciation duration prediction model. After model calculation, the pronunciation duration prediction result of each phoneme PhD is obtained, as shown in Table 1 above.

[0035] The above embodiment determines the original lyrics to be replaced in the original song and obtains the number of words in the target lyrics used to replace the original lyrics; determines the actual pronunciation duration of the syllable corresponding to each word in the target lyrics according to the timestamp of each word in the original lyrics and the number of words in the target lyrics; wherein the syllable corresponding to each word in the target lyrics includes at least one phoneme; respectively encodes the phoneme, the phoneme type of the phoneme and the actual pronunciation duration of the syllable of each word in the target lyrics to obtain the phoneme encoding information, phoneme type encoding information and the actual pronunciation duration encoding information of the syllable for each word; inputs the phoneme encoding information, the phoneme type encoding information and the actual pronunciation duration encoding information of the syllable into a trained phoneme pronunciation duration prediction model to obtain the pronunciation duration prediction result of each phoneme output by the phoneme pronunciation duration prediction model. This embodiment uses a trained phoneme pronunciation duration prediction model to predict the phoneme duration of each word (or each syllable) in the new lyrics under the duration constraint of the original lyrics, and adaptively obtains the pronunciation duration of each phoneme. Compared with the phoneme duration prediction method used in traditional speech synthesis, it can further improve the accuracy of phoneme duration prediction for each word in the cover lyrics.

[0036] In one embodiment, the above-mentioned step S202 includes: if the number of words in the target lyrics is equal to the number of words in the original lyrics, then the timestamp of each word in the original lyrics is used as the timestamp of each word in the target lyrics; and the actual pronunciation duration of each syllable in the target lyrics is obtained according to the timestamp of each word in the target lyrics.

[0037] Specifically, if the number of characters in the target lyrics (i.e., the new lyrics) is equal to the number of characters in the original lyrics, then use the timestamp of each character in the original lyrics as the timestamp of each character in the new lyrics; obtain the actual pronunciation duration of each syllable of each character in the new lyrics according to the timestamp of each character in the new lyrics. For example, if "Happy New Year" is changed to "Beijing Winter Olympics", then the start and end timestamps of the character "北" are the same as those of the character "新", and so on for each character.

[0038] In the above embodiment, when the number of characters in the new lyrics is the same as that in the original lyrics, directly use the timestamps of the original lyrics, which can quickly determine the actual pronunciation duration of each syllable of each character in the new lyrics.

[0039] In one embodiment, the above step S202 includes: if the number of characters in the target lyrics is less than the number of characters in the original lyrics, then perform at least one round of merging on the timestamps of the characters in the original lyrics until the number of timestamps in the original lyrics is the same as the number of characters in the target lyrics; where, in each round of merging, obtain the superimposed duration of the timestamps of any two adjacent characters in the original lyrics, and merge the two timestamps with the smallest superimposed duration, and update the number of timestamps in the original lyrics according to the merged timestamps; based on the order of the current timestamps in the original lyrics, map the current timestamps to each character in the target lyrics to obtain the actual pronunciation duration of each syllable corresponding to each character in the target lyrics.

[0040] Specifically, when the number of characters in the target lyrics (i.e., the new lyrics) is less than the number of characters in the original lyrics, then perform at least one round of merging on the timestamps corresponding to the characters in the original lyrics until the number of timestamps in the original lyrics is the same as the number of characters in the new lyrics. In each round of merging, first calculate the superimposed duration corresponding to the timestamps of any two adjacent characters in the original lyrics, that is, obtain the sum of the timestamps of every two characters (adjacent characters) in the original lyrics as the superimposed duration, and obtain the two timestamps with the smallest superimposed duration, and merge them. Update the number of timestamps in the original lyrics according to the merged timestamps. If it is still more than the number of characters in the new lyrics, then enter the next round of merging, and so on, until the number of timestamps in the original lyrics is the same as the number of characters in the new lyrics. Then, based on the order of the current timestamps in the original lyrics, map each timestamp to each character in the new lyrics to obtain the actual pronunciation duration of each syllable corresponding to each character in the target lyrics.

[0041] For example, the number of characters in the target lyrics "is being staged passionately" (six characters) is less than that in the original lyrics "Wish everyone a happy new year" (seven characters). The cumulative duration of the timestamps of the two characters "new year" in the original lyrics is shorter than that of the adjacent two characters such as "wish" and "fu da". Therefore, the timestamps of the characters "new" and "year" can be merged, and the duration corresponding to the merged timestamp is used as the actual pronunciation duration of the syllable of the character "on". If the difference in the number of characters between the new lyrics and the original lyrics is greater than 1, this rule is reused for adjustment until the number of characters in the new lyrics is the same as the number of timestamps after merging the original lyrics, and so on.

[0042] In the above embodiment, when the number of characters in the new lyrics is less than that in the original lyrics, the timestamps in the original lyrics can be merged according to the preset rules, so as to determine the actual pronunciation duration of each syllable of the characters in the new lyrics, providing a data basis for determining the pronunciation duration of phonemes in the subsequent process.

[0043] In one embodiment, the above step S202 includes: if the number of characters in the target lyrics is more than that in the original lyrics, at least one round of segmentation is performed on the timestamps of the characters in the original lyrics until the number of timestamps of the segmented original lyrics is the same as the number of characters in the target lyrics; wherein, in each round of segmentation, the timestamp corresponding to the character with the longest syllable pronunciation duration is obtained, and the timestamp corresponding to the character with the longest syllable pronunciation duration is segmented to obtain two segmented timestamps, and the number of timestamps of the original lyrics is updated according to the segmented timestamps; based on the order of the current timestamps of the original lyrics, each current timestamp is mapped to each character of the target lyrics to obtain the actual pronunciation duration of each syllable corresponding to each character in the target lyrics.

[0044] Specifically, when the number of characters in the target lyrics (i.e., the new lyrics) is more than that in the original lyrics, at least one round of segmentation is performed on the timestamps of the characters in the original lyrics until the number of timestamps of the original lyrics is the same as the number of characters in the new lyrics. In each round of segmentation, first, the timestamp corresponding to the character with the longest syllable pronunciation duration in the original lyrics is counted, the timestamp is segmented to obtain two timestamps, and the number of timestamps in the original lyrics is updated. If the current number of timestamps of the original lyrics is still not equal to the number of characters in the new lyrics, enter the next round of segmentation, and reuse the above segmentation rule until the number of timestamps of the original lyrics is equal to the number of characters in the new lyrics. Furthermore, based on the order of the current timestamps of the original lyrics, each timestamp can be mapped to each character of the target lyrics to obtain the actual pronunciation duration of each syllable corresponding to each character in the target lyrics. For example, for example, the eight characters "Olympic athletes win the championship and fulfill their dreams" correspond to the seven characters "Wish everyone a happy new year". The syllable pronunciation duration of the character "good" in the original lyrics is the longest, so it is evenly divided as the total duration of the two characters "fulfill their dreams", and so on.

[0045] In the above embodiment, when the number of words in the new lyrics is more than that in the original lyrics, the timestamps in the original lyrics can be divided according to preset rules, thereby determining the actual pronunciation duration of the syllables of each word in the new lyrics, providing a data basis for the subsequent determination of the pronunciation duration of the phonemes.

[0046] In one embodiment, if Figure 3 As shown, Figure 3 The flowchart of predicting the pronunciation duration of a phoneme using a phoneme pronunciation duration prediction model is shown. The phoneme pronunciation duration prediction model includes a bi-gated recurrent unit (Bi-GRU). Step S204 includes:

[0047] The phoneme coding information, phoneme type coding information and syllable actual pronunciation duration coding information are spliced ​​together to obtain splicing features; feature extraction is performed on the splicing features to obtain high-level features; the high-level features are input into a bidirectional gated recurrent unit for bidirectional calculation to obtain forward calculation results and reverse calculation results respectively; the forward calculation results and reverse calculation results are linearly converted to obtain the pronunciation duration prediction result of each phoneme.

[0048] Specifically, the above-mentioned phoneme encoding information (PhI, for example, an N*D vector), phoneme type encoding information (PhT, for example, an N*3 vector) and syllable actual pronunciation duration encoding information (SyD, for example, an N*1 vector) are spliced ​​to obtain a splicing feature (for example, a splicing feature of N*(D+3+1)), and the splicing feature is used as the input of the phoneme pronunciation duration prediction model.

[0049] In this model, the model first uses convolution kernels of different sizes for feature extraction, such as Figure 3 As shown, the sizes of the convolution kernels are 3, 5, 7, and 9, respectively, and the number of channels corresponding to each convolution kernel is 64. These convolution kernels are used to perform convolution operations on the spliced ​​features of the above inputs to obtain high-level features of the operation output, and then the high-level features are input into the above-mentioned bidirectional gated recurrent unit for bidirectional calculation to obtain forward calculation results and reverse calculation results, respectively. The forward calculation results and reverse calculation results can then be linearly converted, and the pronunciation duration prediction results of each phoneme can be obtained based on the linear conversion results (for example, an N*1 PhD vector can be output).

[0050] Optionally, the above-mentioned step of extracting features from the splicing features to obtain high-level features specifically includes: using multiple convolution kernels of different sizes to perform convolution operations on the splicing features to obtain the convolution results output by each convolution kernel; and splicing the convolution results output by each convolution kernel to obtain high-level features.

[0051] Specifically, the sizes of the convolution kernels can be 3, 5, 7, and 9 respectively, and the number of channels corresponding to each convolution kernel is 64. After using these convolution kernels to perform convolution operations on the splicing features of the above input, the convolution results output by each channel are spliced ​​to obtain high-level features (the high-level features can be, for example, N*(64*4) features).

[0052] In the above embodiment, the splicing features of the above input are calculated through a trained phoneme pronunciation duration prediction model to obtain the pronunciation duration prediction result of each phoneme. Compared with traditional speech synthesis technology, it can accurately calculate the duration of each phoneme in the new lyrics according to the duration constraints of the original lyrics, improve the accuracy of the duration prediction of each syllable, and thus improve the user experience.

[0053] In one embodiment, the above method also includes: obtaining the predicted pronunciation duration of the syllable based on the pronunciation duration prediction results of multiple phonemes in the target lyrics; if there is a word in the target lyrics whose predicted pronunciation duration is not equal to the actual pronunciation duration of the syllable, then based on the pronunciation duration prediction results of each phoneme of the word, determining the pronunciation duration weight of each phoneme of the word; based on the actual pronunciation duration of the syllable and the pronunciation duration weight of each phoneme, determining the updated pronunciation duration of each phoneme of the word.

[0054] Specifically, the present application also includes a duration post-processing step. After obtaining the pronunciation duration prediction results of each phoneme, it may happen that the sum of the durations of two phonemes is not equal to the corresponding syllable durations. For example, the pronunciation duration of "jian" is 250 milliseconds, but the model predicts that the pronunciation duration of "j" is 40 and the pronunciation duration of "ian" is 180, with a total duration of 220 milliseconds.

[0055] Therefore, in actual applications, after obtaining the pronunciation duration prediction results of each phoneme output by the phoneme pronunciation duration prediction model, the syllable predicted pronunciation duration of each word in the target lyrics can be obtained based on the pronunciation duration prediction results. Specifically, for each word in the target lyrics, the pronunciation duration prediction results of each phoneme of the word can be summed to obtain the syllable predicted pronunciation duration of the word.

[0056] Then, it is possible to determine whether the predicted pronunciation duration of the syllable of each word in the target lyrics is equal to the actual pronunciation duration of the syllable. If the predicted pronunciation duration of the syllable is equal to the actual pronunciation duration of the syllable, it can be determined that the pronunciation duration prediction results of each phoneme of the word are feasible, thereby obtaining the pronunciation duration corresponding to the phoneme of each word in the target lyrics; if there is at least one word in the target lyrics whose syllable predicted pronunciation duration is not equal to the actual pronunciation duration of the syllable, the pronunciation duration weight of each phoneme can be determined according to the pronunciation duration prediction results of each phoneme of the word, and the pronunciation duration of each phoneme of the word can be re-corrected according to the pronunciation duration weight and the actual pronunciation duration of the syllable to obtain an updated pronunciation duration.

[0057] For example, in the case where the pronunciation duration of the above-mentioned "jian" is 250 milliseconds, and the total duration of the phoneme prediction result is only 220 milliseconds, it is necessary to retain only the ratio of consonants and vowels, which is 40:180. Therefore, the actual pronunciation duration of "j" is set to 40 / (40+180)*250=45 milliseconds, and "ian" is 205 milliseconds.

[0058] In the above embodiment, when the predicted pronunciation duration of a syllable does not correspond to the actual pronunciation duration of the syllable, the phoneme pronunciation duration is further corrected, which can further improve the user experience.

[0059] In one embodiment, if Figure 4 As shown, Figure 4 A flow chart of the model training method is shown, which includes the following steps:

[0060] Step S401, obtaining the training phoneme of each word in the training lyrics, the training phoneme type of the training phoneme, and the actual pronunciation duration of the training syllable corresponding to each word in the training lyrics;

[0061] Specifically, the training lyrics are obtained in advance, and the syllables of each word in the training lyrics are broken down into training phonemes, each training phoneme corresponding to a training phoneme type. For the training lyrics, the actual pronunciation duration of the training syllable corresponding to each word also needs to be known in advance.

[0062] Step S402 : Encode the training phonemes, training phoneme types, and actual pronunciation duration of the training syllables into training phoneme encoding information, training phoneme type encoding information, and actual pronunciation duration encoding information of the training syllables.

[0063] Specifically, the training phonemes, training phoneme types and actual pronunciation duration of training syllables are encoded into training phoneme encoding information, training phoneme type encoding information and actual pronunciation duration encoding information of training syllables using a one-hot encoding method.

[0064] Step S403: input the training phoneme coding information, the training phoneme type coding information and the training syllable actual pronunciation duration coding information into the phoneme pronunciation duration prediction model to be trained to obtain the pronunciation duration prediction result of the training phoneme.

[0065] Specifically, the training phoneme encoding information, the training phoneme type encoding information, and the training syllable actual pronunciation duration encoding information are input into the phoneme pronunciation duration prediction model to be trained to obtain the pronunciation duration prediction result of the training phoneme. The phoneme pronunciation duration prediction model to be trained includes a bi-gated recurrent unit (Bi-GRU), which can capture the dependencies between time series.

[0066] Step S404 : determining a loss value based on a difference between the predicted pronunciation duration of the training phoneme and the actual pronunciation duration of the pre-marked training phoneme.

[0067] Specifically, based on the difference between the predicted pronunciation duration of the training phoneme and the actual pronunciation duration of the pre-labeled training phoneme, the minimum mean square error function is used as the loss function to determine the loss value.

[0068] Step S405 , adjusting the model parameters of the phoneme pronunciation duration prediction model to be trained according to the loss value until the training end condition is met, thereby obtaining a trained phoneme pronunciation duration prediction model.

[0069] Specifically, the model parameters are adjusted. When the loss value between the model's output value and the target value reaches a stable value or other preset conditions are met, the model converges and the training is completed, and a trained phoneme pronunciation duration prediction model is obtained.

[0070] In the above embodiment, a trained phoneme pronunciation duration prediction model can be quickly obtained by using easily accessible lyrics information and a small amount of annotation information as a training set.

[0071] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0072] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an intelligent scheduling method for streaming media files is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0073] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0074] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps of the embodiment of the intelligent scheduling method for streaming media files.

[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0076] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0077] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0078] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for predicting the pronunciation duration of lyrics phonemes, characterized in that: The method comprises: Determining the original lyrics to be replaced in the original song, and obtaining the number of words in the target lyrics used to replace the original lyrics; Determining the actual pronunciation duration of a syllable corresponding to each word in the target lyrics based on the timestamp of each word in the original lyrics and the number of words in the target lyrics; wherein the syllable corresponding to each word in the target lyrics includes at least one phoneme; Encoding the phoneme of each word in the target lyrics, the phoneme type of the phoneme, and the actual pronunciation duration of the syllable to obtain phoneme encoding information, phoneme type encoding information, and actual pronunciation duration encoding information of the syllable for each word; The phoneme coding information, phoneme type coding information and syllable actual pronunciation duration coding information are spliced ​​together, and the obtained splicing features of each word are input into a trained phoneme pronunciation duration prediction model. The feature extraction result of the splicing features is obtained and input into the bidirectional gated recurrent unit in the phoneme pronunciation duration prediction model, and bidirectional calculation is performed to obtain the forward calculation result and the reverse calculation result of each word respectively. The forward calculation result and the reverse calculation result of each word are linearly converted to obtain the pronunciation duration prediction result of each phoneme output by the phoneme pronunciation duration prediction model.

2. The method according to claim 1, characterized in that The step of determining the actual pronunciation duration of a syllable corresponding to each word in the target lyrics according to the timestamp of each word in the original lyrics and the number of words in the target lyrics includes: If the number of words in the target lyrics is equal to the number of words in the original lyrics, the timestamp of each word in the original lyrics is used as the timestamp of each word in the target lyrics; The actual pronunciation duration of the syllable of each word in the target lyrics is obtained according to the timestamp of each word in the target lyrics.

3. The method according to claim 1, characterized in that The step of determining the actual pronunciation duration of a syllable corresponding to each word in the target lyrics according to the timestamp of each word in the original lyrics and the number of words in the target lyrics includes: If the target lyrics have fewer words than the original lyrics, the timestamps of the words in the original lyrics are merged for at least one round until the number of timestamps of the original lyrics is the same as the number of words in the target lyrics; wherein, in each round of merging, the superposition time of the timestamps of any two adjacent words in the original lyrics is obtained, and the two timestamps with the smallest superposition time are merged, and the number of timestamps of the original lyrics is updated according to the merged timestamps; Based on the order of the current timestamps of the original lyrics, the current timestamps are mapped to the characters of the target lyrics to obtain the actual pronunciation duration of the syllable corresponding to each character in the target lyrics.

4. The method according to claim 1, wherein The step of determining the actual pronunciation duration of a syllable corresponding to each word in the target lyrics according to the timestamp of each word in the original lyrics and the number of words in the target lyrics includes: If the target lyrics have more words than the original lyrics, the timestamps of the words in the original lyrics are segmented for at least one round until the number of timestamps in the segmented original lyrics is the same as the number of words in the target lyrics; wherein, in each round of segmentation, the timestamp corresponding to the word with the longest syllable pronunciation duration is obtained, and the timestamp corresponding to the word with the longest syllable pronunciation duration is segmented to obtain two segmented timestamps, and the number of timestamps in the original lyrics is updated according to the segmented timestamps; Based on the order of the current timestamps of the original lyrics, the current timestamps are mapped to the characters of the target lyrics to obtain the actual pronunciation duration of the syllable corresponding to each character in the target lyrics.

5. The method according to claim 1, characterized in that The obtaining of the feature extraction result of the splicing feature includes: Using multiple convolution kernels of different sizes to perform convolution operation on the splicing features, and obtaining the convolution result output by each convolution kernel; The convolution results output by each convolution kernel are spliced ​​to obtain high-level features, and the high-level features are used as feature extraction results.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Based on the pronunciation duration prediction results of multiple phonemes of the target lyrics, obtaining the predicted pronunciation duration of each syllable in the target lyrics; If there is a word in the target lyrics whose predicted syllable pronunciation duration is not equal to the actual syllable pronunciation duration, determining the pronunciation duration weight of each phoneme of the word based on the pronunciation duration prediction results of each phoneme of the word; Based on the actual pronunciation duration of the syllable and the pronunciation duration weight of each phoneme, an updated pronunciation duration of each phoneme of the character is determined.

7. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Obtaining the training phoneme of each word in the training lyrics, the training phoneme type of the training phoneme, and the actual pronunciation duration of the training syllable corresponding to each word in the training lyrics; Encoding the training phoneme, the training phoneme type and the actual pronunciation duration of the training syllable into training phoneme encoding information, training phoneme type encoding information and training syllable actual pronunciation duration encoding information; Inputting the training phoneme encoding information, the training phoneme type encoding information and the training syllable actual pronunciation duration encoding information into the phoneme pronunciation duration prediction model to be trained to obtain the pronunciation duration prediction result of the training phoneme; Determining a loss value based on a difference between a predicted pronunciation duration of the training phoneme and an actual pronunciation duration of the training phoneme that has been marked in advance; The model parameters of the phoneme pronunciation duration prediction model to be trained are adjusted according to the loss value until the training end condition is met, thereby obtaining the trained phoneme pronunciation duration prediction model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Voice training method and device based on deep learning, equipment and storage medium

    CN112735389A

  • Speech recognition method and apparatus, device, and storage medium

    WO2022078146A1