A method for generating lyrics timestamps, a generating device, and a storage medium
By determining the phoneme state sequence and state transition probability, combining the allocation probability of the acoustic model, and calculating the confidence of the lyric timestamp, the problem of lyric timestamp offset in the prior art is solved, and a more accurate lyric timestamp generation is achieved.
Patent Information
- Application Number
- CN202310050590.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-01
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-02-01
AI Technical Summary
The existing lyric timestamp generation technology can easily lead to local lyrics timestamp offsets through global decoding, resulting in uneven quality of generated lyrics timestamps.
By obtaining the lyric text of the target song and the target stem sound audio, determining the phoneme state sequence and state transition probability, input a pre-trained acoustic model to obtain the allocation probability and state transition path, compute the confidence of the word timestamp and generate a more accurate lyric timestamp.
It effectively avoids the offset of lyric timestamps, and the generated lyric timestamps are more accurate, improving the performance of the automatic alignment model, and reducing the cost of manually labeling lyrics.
Smart Images

Figure CN116092515B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data processing, and in particular, to a method for generating lyrics timestamps, a generating device, and a storage medium. Background Art
[0002] Lyrics timestamps generally refer to the time when lyrics appear in a song audio. The timestamp consists of a start time and an end time. When a publisher releases a created song to a music platform, a QRC file needs to be produced. The QRC file is the word-by-word lyrics of the song and is a lyrics file containing timestamp information for each word. Since it is easy to have errors and requires a large amount of labor cost when manually producing the QRC file, the existing technologies for generating lyrics timestamps generally adopt an automatic alignment model with a DNN-HMM / GMM-HMM structure, where DNN / GMM is an acoustic model and HMM is a language model.
[0003] The working process of this automatic alignment model is as follows: The acoustic model identifies the state of the song audio features in the language model, and the language model calculates the transition probability between states to obtain the final decoding result. Align the whole song and the lyrics text. The features of the whole song are decoded in the lyrics text space, and the decoding result is globally optimal. Global optimality means that when performing alignment, it is generally based on the start time period and the end time period of the song to align with the lyrics text, and it is difficult to determine whether the local part between the two time periods is aligned and the specific position of the unaligned local part; it is easy to cause mutual influence between local parts, resulting in a deterioration of the local effect. If the situation is serious, the entire timestamp may shift.
[0004] It can be seen that the existing technologies for generating lyrics timestamps obtain the globally optimal decoding result of the whole song through global decoding, which easily causes the situation that local lyrics shift for global optimality. If the situation is serious, it may lead to large segments of lyrics shifting. If the shift is not processed, the quality of the generated lyrics timestamps will be uneven. Summary of the Invention
[0005] The embodiments of the present application provide a method for generating lyrics timestamps, a generating device, and a storage medium, which can effectively avoid the shift of word timestamps in the current lyrics text and generate more accurate lyrics timestamps.
[0006] The embodiments of the present application provide a method for generating lyrics timestamps, including:
[0007] Obtain the lyrics text corresponding to the target song and the target dry audio of the target song;
[0008] Determine a phoneme state sequence corresponding to the lyrics text according to the phonemes corresponding to each character in the lyrics text; the phoneme state sequence includes a plurality of phoneme states corresponding to the audio frames of the target dry audio;
[0009] Determine the state transition probability between adjacent two-frame phoneme states in the phoneme state sequence;
[0010] Input the target dry audio and the phoneme state sequence into a pre-trained acoustic model to obtain a first allocation probability that each frame of audio corresponding to the lyrics text is allocated to each phoneme state in the phoneme state sequence, and a second allocation probability that each frame of audio corresponding to the current lyrics text is allocated to each phoneme state corresponding to the current lyrics text; the current lyrics text is a segment of text in the lyrics text;
[0011] According to the second allocation probability and the state transition probability, determine the maximum probability of phoneme state transition in multiple frames of audio corresponding to the current lyrics text and the state transition path corresponding to the maximum probability;
[0012] According to the first allocation probability corresponding to each frame of phoneme state in the state transition path and the corresponding state transition probability, determine the target probability of each frame of phoneme state corresponding to the current lyrics text;
[0013] Determine the confidence of the word timestamp in the current lyrics text according to the state transition path and the target probability, and generate the lyrics timestamp of the current lyrics text according to the confidence of the word timestamp.
[0014] Further, the obtaining the lyrics text corresponding to the target song includes:
[0015] Obtain an initial lyrics text corresponding to the target song;
[0016] Perform non-lyrics information filtering processing on the initial lyrics text to obtain the lyrics text.
[0017] Further, the determining the phoneme state sequence corresponding to the lyrics text according to the phonemes corresponding to each character in the lyrics text includes:
[0018] Obtain the phonemes corresponding to the audio frames in the target dry audio to obtain a phoneme sequence corresponding to the target dry audio;
[0019] Match the phonemes corresponding to each character in the lyrics text with the phoneme sequence and input them into a pre-set language model to obtain the phonemes corresponding to each character in the lyrics text in the phoneme sequence;
[0020] Convert the phonemes corresponding to each character in the lyric text into phoneme states in the phoneme sequence, obtaining the phoneme state sequence corresponding to the lyric text.
[0021] Further, the determination of the state transition probability between adjacent two-frame phoneme states in the phoneme state sequence includes:
[0022] Take multiple phoneme states of adjacent preset frames as a multi-phoneme state group, determine the total number of phoneme states in the phoneme state sequence and the number of occurrences of the multi-phoneme state group in the phoneme state sequence;
[0023] Divide the number of occurrences by the total number of phoneme states, obtaining the occurrence probability of the multi-phoneme state group in the phoneme state sequence;
[0024] Take the occurrence probability as the state transition probability between adjacent two-frame phoneme states in the multi-phoneme state group.
[0025] Further, the input of the target dry audio and the phoneme state sequence into the pre-trained acoustic model to obtain the first allocation probability and the second allocation probability includes:
[0026] Extract the audio features of the target dry audio, input the audio features of the target dry audio and the phoneme states in the phoneme state sequence into the pre-trained acoustic model to obtain the first allocation probability;
[0027] Extract the audio features of the dry audio corresponding to the current lyric text, input the audio features of the dry audio and the phoneme states corresponding to the current lyric text into the pre-trained acoustic model to obtain the second allocation probability.
[0028] Further, the determination of the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability in the multi-frame audio corresponding to the current lyric text according to the second allocation probability and the state transition probability includes:
[0029] In the multi-frame audio corresponding to the current lyric text, calculate the sum of the products of the second allocation probability of the phoneme state of the current frame and the state transition probability from the phoneme state of the current frame to the phoneme state of the next frame;
[0030] Take the maximum value in the sum of the products as the maximum probability of the phoneme state transition;
[0031] Determine the target phoneme state of each frame of audio corresponding to the current lyric text among the maximum probabilities, and take the multiple target phoneme states corresponding to the multi-frame audio as the state transition path corresponding to the maximum probability.
[0032] Further, determining the target probability of each phoneme state corresponding to the current lyric text according to the first allocation probability and the corresponding state transition probability of each phoneme state in the state transition path includes:
[0033] Calculating the product of the first allocation probability and the corresponding state transition probability of each phoneme state in the state transition path to obtain the target probability of each phoneme state corresponding to the current lyric text.
[0034] Further, determining the confidence level of the word timestamp in the current lyric text according to the state transition path and the target probability includes:
[0035] Dividing the target probability of each phoneme state in the current lyric text by the probability of the corresponding phoneme state in the state transition path to obtain a quotient, taking the logarithm of the quotient and taking the negative value to obtain the pronunciation quality score of each audio frame in the current lyric text;
[0036] Taking the average value of the pronunciation quality scores of the audio frames corresponding to the words in the current lyric text to obtain the pronunciation quality score of the word;
[0037] If the pronunciation quality score of the target word is greater than a preset threshold, determining that the timestamp of the target word is a word timestamp with high confidence;
[0038] If the pronunciation quality score of the target word is less than a preset threshold, determining that the timestamp of the target word is a word timestamp with low confidence.
[0039] Further, generating the lyric timestamp of the current lyric text according to the confidence level of the word timestamp includes:
[0040] Determining the start time and the end time of the sub-timestamp with low confidence;
[0041] If the first signal energy of a preset frame before the start time is less than a preset energy threshold, shifting the start time backward by one frame to correct the sub-timestamp with low confidence until the first signal energy is greater than the preset energy threshold;
[0042] If the second signal energy of a preset frame after the end time is greater than the preset energy threshold, shifting the end time forward by one frame to correct the sub-timestamp with low confidence until the second signal energy is less than the preset energy threshold;
[0043] Taking the corrected sub-timestamp with low confidence and the word timestamp with high confidence as the lyric timestamp of the current lyric text.
[0044] The embodiment of the present application further provides a device for generating a lyric timestamp, including:
[0045] An acquisition unit, configured to acquire the lyric text corresponding to a target song and the target dry audio of the target song;
[0046] A first determination unit, configured to determine a phoneme state sequence corresponding to the lyric text according to the phonemes corresponding to each character in the lyric text; the phoneme state sequence includes a plurality of phoneme states corresponding to the audio frames of the target dry audio;
[0047] A second determination unit, configured to determine the state transition probability between adjacent two-frame phoneme states in the phoneme state sequence;
[0048] An input unit, configured to input the target dry audio and the phoneme state sequence into a pre-trained acoustic model, to obtain a first allocation probability that each frame of audio corresponding to the lyric text is allocated to each phoneme state in the phoneme state sequence, and a second allocation probability that each frame of audio corresponding to the current lyric text is allocated to each phoneme state corresponding to the current lyric text; the current lyric text is a segment of text in the lyric text;
[0049] A third determination unit, configured to determine the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability in multiple frames of audio corresponding to the current lyric text according to the second allocation probability and the state transition probability;
[0050] An execution unit, configured to determine the target probability of phoneme state transition in multiple frames of audio corresponding to the current lyric text according to the first allocation probability corresponding to each frame of phoneme state in the state transition path and the state transition probability between adjacent two-frame phoneme states in the state transition path;
[0051] A fourth determination unit, configured to determine the confidence of the word timestamp in the current lyric text according to the maximum probability and the target probability;
[0052] A generation unit, configured to generate a lyric timestamp of the current lyric text according to the confidence of the word timestamp.
[0053] An embodiment of the present application further provides a device for generating lyric timestamps, including:
[0054] A central processing unit, a memory, and an input / output interface;
[0055] The memory is a transient storage memory or a persistent storage memory;
[0056] The central processing unit is configured to communicate with the memory and execute the instructions in the memory to perform the above method.
[0057] The embodiments of the present application also provide a computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the above-mentioned method.
[0058] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0059] The embodiments of the present application include: determining the state transition probability between adjacent two-frame phoneme states in a phoneme state sequence; determining the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability according to the second allocation probability and the state transition probability; determining the target probability of the phoneme state in each frame of audio corresponding to the current lyric text according to the first allocation probability and the corresponding state transition probability corresponding to each frame of phoneme state in the state transition path; determining the confidence of the word timestamp in the current lyric text according to the state transition path and the target probability, and generating the lyric timestamp of the current lyric text according to the confidence of the word timestamp; by determining the confidence of the word timestamp in the current lyric text and generating the lyric timestamp of the current lyric text according to the confidence of the word timestamp, it can effectively avoid the offset of the word timestamp in the current lyric text and make the generated lyric timestamp of the current lyric text more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.
[0061] Figure 1 It is a communication network architecture diagram of a lyric timestamp generation device disclosed in the embodiments of the present application;
[0062] Figure 2 It is a flowchart for generating a lyric timestamp disclosed in the embodiments of the present application;
[0063] Figure 3 It is another flowchart for generating a lyric timestamp disclosed in the embodiments of the present application;
[0064] Figure 4 It is another flowchart for generating a lyric timestamp disclosed in the embodiments of the present application;
[0065] Figure 5 It is a state transition probability diagram disclosed in the embodiments of the present application;
[0066] Figure 6 It is a diagram of a lyric timestamp generation device disclosed in the embodiments of the present application;
[0067] Figure 7Another diagram of the device for generating lyrics timestamps disclosed in the embodiments of the present application. Detailed implementation manners
[0068] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.
[0069] In the description of the embodiments of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the embodiments of the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the embodiments of the present application.
[0070] In the description of the embodiments of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific situations.
[0071] When an existing audio player plays a song, it can display the lyrics corresponding to the current song playing progress on the song playing page according to the current playing progress of the song and the lyrics timestamp information included in the lyrics file. Therefore, before playing a song, an existing audio player needs a device for generating lyrics timestamps to generate the lyrics timestamps corresponding to the lyrics text. For example Figure 1As shown, the existing lyric timestamp generation device 101 can obtain the song audio of the audio player 102 and the lyric text corresponding to the song, generate the lyric timestamps corresponding to the lyric text based on the song audio and the lyric text, and send the lyric timestamp information to the audio player 102. When playing the song, the audio player 102 uses the lyric timestamps to display the lyrics corresponding to the current song playing progress. It can be understood that the audio player 102 can be a mobile phone, a tablet computer, a smart wearable device or a laptop computer, and specific limitations are not made here. The lyric timestamp generation device 101 can be connected to one or more audio players 102, and specific limitations are not made here; and the lyric timestamp generation device 101 can be set inside the audio player 102, that is, physically integrated with the audio player 102, or can be set outside the audio player 102, that is, physically relatively independent from the audio player 102, and specific limitations are not made here.
[0072] When the existing lyric timestamp generation device generates lyric timestamps, it generally aligns the whole song and the lyric text, that is, aligns the start time period and the end time period of the song with the lyric text, and it is difficult to determine whether the local part between the two time periods is aligned and the specific position of the local misalignment, which easily leads to the situation of local offset. Therefore, the embodiment of the present application provides a method for generating lyric timestamps, which can effectively avoid the offset of the word timestamps in the current lyric text and generate more accurate lyric timestamps. As Figure 2 shown, the specific steps are as follows:
[0073] 201. Obtain the lyric text corresponding to the target song and the target dry voice audio of the target song.
[0074] In the embodiment of the present application, the lyric timestamp generation device can obtain the lyric text corresponding to the target song and the target dry voice audio of the target song. Specifically, it can obtain the song audio corresponding to the target song (the current song to be predicted) for which lyric timestamps are to be generated, and perform dry voice extraction on the song audio, that is, extract the human voice from the song audio to obtain the target dry voice audio. For example, the dry voice in the song audio is extracted through the spleeter model, and the target dry voice audio includes the human voice singing audio corresponding to the target song and does not include the accompaniment audio of the song. Among them, the duration of the obtained target dry voice audio is the same as the duration of the song audio. And the lyric text is the text corresponding to the target song. It can be understood that the song audio and the lyric text of the target song can be obtained from the audio database corresponding to the audio player.
[0075] Optionally, since the lyrics text may include some non-lyric information such as songwriting and composition that is not related to singing, the lyrics text can be obtained by filtering the non-lyric information from the initial lyrics text of the target song. That is, obtain the initial lyrics text corresponding to the target song; perform non-lyric information filtering on the initial lyrics text to obtain the lyrics text. Non-lyric information generally appears at the beginning of the lyrics text. The non-lyric information filtering process can set corresponding filtering rules according to the information characteristics of the non-lyric information. For example, by detecting keywords, key symbols, etc., perform non-lyric information filtering on the initial lyrics text, delete the non-lyric information in the initial lyrics text, and only retain the text corresponding to the lyrics of the target song, that is, retain the lyrics text related to singing.
[0076] 202. Determine the phoneme state sequence corresponding to the lyrics text according to the phonemes corresponding to each character in the lyrics text.
[0077] After obtaining the lyrics text of the target song, the phoneme state sequence corresponding to the lyrics text can be determined according to the phonemes corresponding to each character in the lyrics text. Among them, the phoneme state sequence includes multiple phoneme states corresponding to the audio frames of the target dry audio, that is, each phoneme state corresponds to one frame of audio in the target dry audio. It can be understood that a phoneme is the smallest speech unit divided according to the natural attributes of a language, and one pronunciation action forms one phoneme. For example, the phonemes corresponding to the pronunciation of Mandarin can be represented by pinyin. For example, the phonemes corresponding to the pronunciation of the three characters "wo, he, ni" can be represented by "w, o, h, e, n, i" respectively.
[0078] A phoneme state is a more detailed speech unit obtained by dividing each phoneme. Generally, one phoneme can correspond to three phoneme states, that is, the starting sound, the sustained sound, and the ending sound of the phoneme pronunciation can be determined as the three phoneme states corresponding to the phoneme. The pronunciation dictionary generally records the mapping relationship between characters (words) and phonemes, that is, the phonemes corresponding to each character (word). The lyrics timestamp generation device can determine the phonemes corresponding to each character in the lyrics text according to the mapping relationship between characters and phonemes in the pronunciation dictionary, and obtain the phoneme set corresponding to the lyrics text. Then, according to the text order of each character in the lyrics text, or according to the pronunciation order of the phonemes in the target dry audio, sort the phonemes in the phoneme set to obtain the phoneme sequence corresponding to the lyrics text, and then convert the phonemes in the phoneme sequence into states to obtain the corresponding phoneme state sequence. It can be understood that the lyrics text corresponds to the target dry audio, that is, the text order and the pronunciation order are the same, and the corresponding relationship between each phoneme and the phoneme state can be set artificially.
[0079] 203. Determine the state transition probability between two adjacent frames of phoneme states in the phoneme state sequence.
[0080] After obtaining the phoneme state sequence, the state transition probability between adjacent two frames of phoneme states in the phoneme state sequence can be obtained. It can be understood that the phoneme state sequence includes multiple phoneme states, and one phoneme state corresponds to one frame of audio. The multiple phoneme states are sequentially distributed on one frame of audio respectively. The state transition probability between adjacent two frames of phoneme states can be obtained according to the adjacent multiple frames of phoneme states in the phoneme state sequence. The state transition probability refers to the probability that in the phoneme state sequence, the phoneme state corresponding to the current adjacent frame transitions to the phoneme state corresponding to the next frame. It can be understood that the transition of phoneme states can be the transition between two different phoneme states, that is, the phoneme states corresponding to the current frame and the next frame are different phoneme states, or it can be the transition between the same phoneme states, that is, the phoneme states corresponding to the current frame and the next frame are the same phoneme states. Specifically, it is not limited here.
[0081] 204. Input the target dry audio and the phoneme state sequence into the pre-trained acoustic model to obtain the first allocation probability of the lyric text and the second allocation probability of the current lyric text.
[0082] The device for generating lyric timestamps can input the target dry audio and the phoneme state sequence into the pre-trained acoustic model to obtain the first allocation probability of the lyric text and the second allocation probability of the current lyric text. Among them, the current lyric text is a segment of text in the lyric text. The first allocation probability is the probability that each frame of audio corresponding to the lyric text is allocated to each phoneme state in the phoneme state sequence. The second allocation probability is the probability that each frame of audio corresponding to the current lyric text is allocated to each phoneme state corresponding to the current lyric text. The acoustic model can be a GMM acoustic model or a DNN acoustic model. Specifically, it is not limited here.
[0083] It can be understood that the lyric text corresponds to the entire target dry audio and the entire phoneme state sequence. That is, the entire target dry audio and the entire phoneme state sequence can be input into the acoustic model, and then the acoustic model outputs the probability values of each phoneme state corresponding to each audio frame in the target dry audio, which is the first allocation probability. And the current lyric text is a segment of text in the lyric text. The current lyric text corresponds to a segment of dry audio in the target dry audio and a part of the phoneme state sequence. That is, the dry audio corresponding to the current lyric text and a part of the phoneme state sequence can be input into the acoustic model, and then the acoustic model outputs the probability values of each phoneme state corresponding to each audio frame in the dry audio corresponding to the current lyric text. And the phoneme state is a phoneme state in a part of the phoneme state sequence corresponding to the current lyric text, then the second allocation probability is obtained.
[0084] 205. Determine the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability according to the second allocation probability and the state transition probability.
[0085] The device for generating the lyric timestamp can determine the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability according to the second distribution probability and the state transition probability, that is, determine the maximum probability of phoneme state transition in the multi-frame audio corresponding to the current lyric text and the state transition path corresponding to the maximum probability. Specifically, in the dry audio corresponding to the current lyric text, each audio frame can be set with each phoneme state in the phoneme sequence corresponding to the current lyric text, and each phoneme state is associated with a corresponding second distribution probability (i.e., the probability that the phoneme state is located in this audio frame); when determining the maximum probability of phoneme state transition, all possible state transition paths in the dry audio corresponding to the current lyric text can be determined first, and then the probability of each state transition path can be calculated according to the second distribution probability and the state transition probability, and the maximum value is the maximum probability of the phoneme state transition. The maximum probability of phoneme state transition in the dry audio corresponding to the current lyric text can also be determined according to the Viterbi decoding. The state transition path is the phoneme state corresponding to each audio frame in the dry audio corresponding to the current lyric text when the maximum probability is obtained, and the obtained phoneme state sequence is the state transition path.
[0086] 206. Determine the target probability of each phoneme state according to the first distribution probability corresponding to each phoneme state in the state transition path and the state transition probability.
[0087] After determining the state transition path of the current lyric text, the target probability of each phoneme state can be determined according to the first distribution probability corresponding to each phoneme state in the state transition path and the state transition probability, that is, determine the target probability of the phoneme state in each audio frame corresponding to the current lyric text according to the first distribution probability corresponding to each phoneme state in the state transition path and the state transition probability of the phoneme state in the state transition path. Specifically, each audio frame in the dry audio corresponding to the current lyric text has a determined phoneme state in the state transition path. According to the first distribution probability of the determined phoneme state in the current audio frame and the state transition probability of the determined phoneme state in the current audio frame, the target probability of phoneme state transition in each audio frame corresponding to the current lyric text can be obtained.
[0088] 207. Determine the confidence of the word timestamp in the current lyric text according to the state transition path and the target probability, and generate the lyric timestamp of the current lyric text according to the confidence of the word timestamp.
[0089] The device for generating lyric timestamps can determine the confidence level of the word timestamps in the current lyric text according to the state transition path and the target probability, and generate the lyric timestamps of the current lyric text according to the confidence level of the word timestamps. Specifically, the probability of each frame of phoneme state can be determined in the state transition path. Generally, the second assignment probability of the phoneme state is multiplied by the state transition probability of the phoneme state to obtain the probability of the phoneme state. According to the probability of the phoneme state and the target probability corresponding to the phoneme state, the pronunciation quality score of the audio frame corresponding to the phoneme state can be calculated, and the pronunciation quality score of each frame corresponding to the current lyric text can be obtained. The pronunciation quality score of a word can be obtained by taking the average value of the pronunciation quality scores of the audio frames to which the word belongs. The confidence level of the word timestamp can be determined according to the pronunciation quality score of the word, and the confidence levels of all the word timestamps in the current lyric text can be determined. Then, the lyric timestamps of the current lyric text are generated according to the confidence level of the word timestamps, that is, in the current lyric text, it is determined whether the confidence level of the word timestamp is a high confidence level or a low confidence level, the word timestamps with low confidence levels are corrected, and the word timestamps with high confidence levels and the corrected word timestamps with low confidence levels are combined into the lyric timestamps of the current lyric text.
[0090] It can be seen that the embodiments of the present application include: determining the state transition probability between two adjacent frames of phoneme states in the phoneme state sequence; determining the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability according to the second assignment probability and the state transition probability; determining the target probability of the phoneme state in each frame of audio corresponding to the current lyric text according to the first assignment probability and the corresponding state transition probability of each frame of phoneme state in the state transition path; determining the confidence level of the word timestamps in the current lyric text according to the state transition path and the target probability, and generating the lyric timestamps of the current lyric text according to the confidence level of the word timestamps; by determining the confidence level of the word timestamps in the current lyric text and generating the lyric timestamps of the current lyric text according to the confidence level of the word timestamps, the offset of the word timestamps in the current lyric text can be effectively avoided, and the lyric timestamps of the generated current lyric text are more accurate.
[0091] Next, in combination with Figure 3 , the method for generating lyric timestamps in the embodiments of the present application will be described in detail as follows:
[0092] 301. Obtain the lyric text corresponding to the target song and the target dry audio of the target song.
[0093] The device for generating lyric timestamps can obtain the lyric text corresponding to the target song and the target dry audio of the target song. It can be understood that step 301 is similar to step 201 above, and will not be elaborated here specifically.
[0094] 302. Convert the phonemes corresponding to each character in the lyric text into phoneme states in the phoneme sequence, obtaining the phoneme state sequence corresponding to the lyric text.
[0095] The device for generating lyric timestamps can convert the phonemes corresponding to each character in the lyric text into phoneme states in the phoneme sequence, obtaining the phoneme state sequence corresponding to the lyric text, where the phoneme sequence is the phoneme sequence corresponding to the target dry audio. Specifically, the phonemes corresponding to the audio frames in the target dry audio can be obtained to get the phoneme sequence corresponding to the target dry audio, where the phoneme sequence is obtained by arranging multiple phonemes according to the pronunciation order of the phonemes in the target dry audio. Then, the phonemes corresponding to each character in the lyric text and the phoneme sequence can be input into a pre-set language model (HMM language model) for matching to obtain the phonemes corresponding to each character in the lyric text in the phoneme sequence. Specifically, the phoneme sequence can be input into the pre-set language model to obtain the recognized text corresponding to the phoneme sequence; the phonemes corresponding to each character in the recognized text in the phoneme sequence are determined as the phonemes corresponding to each character in the lyric text in the phoneme sequence.
[0096] Among them, the pre-set language model is a statistical model, such as the n-gram model. As Figure 4 shown, all the lyric texts corresponding to the songs in the song library can be obtained in advance, and then the number of occurrences of each character (word) in the obtained lyric texts can be counted to determine the possible probability of each character's occurrence. For the n-gram model, the probability of n characters appearing simultaneously can also be determined. For example, when n = 3, the 3-gram model can obtain the probability of each character in the lyric appearing in the form of a triple (for example: assuming that the currently counted lyric has 10,000 characters, and the triple "I and you" has appeared 5 times, then the statistical probability of this triple is 5 / 10,000, indicating that the probability of "I" transitioning to "and" and "and" transitioning to "you" is 5 / 10,000). Since the recognized text is recognized based on the target dry audio, the recognized text is the lyric text recognized based on the target dry audio. Each phoneme corresponding to each character in the recognized lyric text in the phoneme sequence is the phoneme corresponding to each character in the lyric text in the phoneme sequence. To correspond one-to-one with the audio frames, the unit of the characters in the lyric text can be refined, the characters can be converted into phonemes, the lyric text can be converted into the corresponding phoneme sequence, and the phonemes corresponding to each character in the lyric text in the phoneme sequence are converted into phoneme states (HMM states). Generally, one phoneme corresponds to 3 phoneme states, the frames and the phoneme states correspond one-to-one, and one frame of audio features corresponds to one phoneme state, thus obtaining the phoneme state sequence (HMM state sequence) corresponding to the lyric text.
[0097] 303. Calculate the occurrence probability of the multi-phoneme state group in the phoneme state sequence to obtain the state transition probability between adjacent two-frame phoneme states.
[0098] After obtaining the phoneme state sequence corresponding to the lyric text, the occurrence probability of the multi-phoneme state group in the phoneme state sequence can be calculated to obtain the state transition probability between adjacent two-frame phoneme states. Specifically, multiple phoneme states of adjacent preset frames can be used as the multi-phoneme state group to determine the total number of phoneme states in the phoneme state sequence and the number of occurrences of the multi-phoneme state group in the phoneme state sequence; for example, one letter can be used to represent one phoneme state, and the adjacent "ABC" in the phoneme state sequence can be used as the tri-phoneme state group, and the number of occurrences of "ABC" in the phoneme state sequence and the total number of phoneme states in the phoneme state sequence can be determined. Dividing the number of occurrences by the total number of phoneme states can obtain the occurrence probability of the multi-phoneme state group in the phoneme state sequence, and taking the occurrence probability as the state transition probability between adjacent two-frame phoneme states in the multi-phoneme state group, that is, the transition probability of the tri-phoneme state group. The transition probabilities between all phoneme states can form a state transition probability graph. As Figure 5 shown, in the figure, the horizontal axis represents the time frames (time observations) x1,…x7 (7 frames in this example), the vertical axis represents m phoneme states corresponding to the current lyric text, S refers to the phoneme states (states), a is the state transition probability, and the arrow indicates the transition of the phoneme state.
[0099] 304. Input the audio features of the target dry audio and the phoneme states in the phoneme state sequence into the pre-trained acoustic model to obtain the first assignment probability, and input the audio features of the dry audio and the phoneme states corresponding to the current lyric text into the pre-trained acoustic model to obtain the second assignment probability.
[0100] The generating device for lyric timestamps can input the audio features of the target dry audio and the phoneme states in the phoneme state sequence into a pre-trained acoustic model to obtain a first assignment probability, and input the audio features of the dry audio and the phoneme states corresponding to the current lyric text into the pre-trained acoustic model to obtain a second assignment probability. Specifically, the audio features of the target dry audio can be extracted. The audio features can be MFCC features. Then, the audio features of the target dry audio and the phoneme states in the phoneme state sequence are input into the pre-trained acoustic model to obtain a first assignment probability. It can be understood that the acoustic model can be used as a classification model. The target dry audio is the audio corresponding to the lyric text, and the phoneme state sequence is the state sequence corresponding to the lyric text. That is, the first assignment probability output by the acoustic model is the probability value that each audio frame in the target dry audio corresponding to the lyric text belongs to each phoneme state in the phoneme state sequence. Here, the phoneme state sequence can be understood as the entire state space, and obtaining the first assignment probability can be understood as the emission probability under the entire state space. When there are n states in the entire state space, the probability value that each frame belongs to each of the n states can be calculated. For example, the first assignment probability of the phoneme state in each frame corresponding to the lyric text can be calculated as 1 / n.
[0101] Meanwhile, the audio features of the dry audio corresponding to the current lyric text can also be extracted, and the audio features of the dry audio and the phoneme states corresponding to the current lyric text are input into the pre-trained acoustic model to obtain a second assignment probability. The second assignment probability can be understood as the emission probability under the current lyric state space, that is, it is calculated according to the phoneme states corresponding to the current lyric text without considering the phoneme states corresponding to the other lyric texts outside the current lyric text. Generally, the phoneme states corresponding to the current lyric text are fewer than the states in the phoneme state sequence corresponding to the lyric text. For example, if the phoneme states corresponding to the current lyric text are m phoneme states (m < n), the probability that each frame belongs to each of the m phoneme states can be calculated. For example, the second assignment probability of the phoneme state in each frame corresponding to the current lyric text can be calculated as 1 / m. It can be understood that Figure 5 is the state transition probability graph corresponding to the current lyric text, and the assignment probability of the phoneme state in the graph is the second assignment probability.
[0102] Among them, the training process of the acoustic model can be as follows: Obtain the sample dry audio of the sample song and the sample lyric text corresponding to the sample song; determine the phonemes corresponding to each character in the sample lyric text; extract the audio features of the sample dry audio, and use the audio features and the phonemes corresponding to each character in the sample lyric text as the first training sample, and perform monophone training on the acoustic model based on the first training sample to obtain a first acoustic model; perform triphone training on the first acoustic model based on the first training sample to obtain the trained acoustic model.
[0103] 305. Calculate the sum of the products of the second assignment probability of the phoneme state of the current frame and the state transition probability of the phoneme state of the current frame transitioning to the phoneme state of the next frame to obtain the maximum probability of phoneme state transition.
[0104] The device for generating the lyric timestamp can calculate the sum of the products of the second assignment probability of the phoneme state of the current frame and the state transition probability of the phoneme state of the current frame transitioning to the phoneme state of the next frame to obtain the maximum probability of phoneme state transition. That is, through forced alignment decoding, in the state transition probability graph corresponding to the current lyric text, the maximum probability path can be found, that is, the overall probability value is the largest. Here, the maximum probability path (optimal decoding path) can be obtained through Viterbi decoding. Specifically, the state transition probability graph can also be used to calculate the sum of the products of the second assignment probability of the phoneme state of the current frame and the state transition probability of the phoneme state of the current frame transitioning to the phoneme state of the next frame in the multi-frame audio corresponding to the current lyric text; take the maximum value in the sum of the products as the maximum probability of phoneme state transition; determine the target phoneme state of each frame of audio corresponding to the current lyric text in the maximum probability, and use the multiple target phoneme states corresponding to the multi-frame audio as the state transition path (maximum probability path) corresponding to the maximum probability.
[0105] 306. Calculate the product of the first assignment probability corresponding to the phoneme state of each frame in the state transition path and the corresponding state transition probability to obtain the target probability of the phoneme state of each frame corresponding to the current lyric text.
[0106] After obtaining the state transition path, the audio state corresponding to each audio frame can be determined. Then, the product of the first assignment probability corresponding to the phoneme state of each frame in the state transition path and the corresponding state transition probability can be calculated to obtain the target probability of the phoneme state of each frame corresponding to the current lyric text. That is, through free decoding, the probability value of each phoneme state in the forced alignment decoding path under free decoding can be obtained. It can be understood that the target probability is obtained by multiplying the first assignment rate corresponding to the phoneme state on the state transition path by the corresponding state transition probability, and the first assignment probability is the assignment probability of the phoneme state corresponding to the lyric text. That is, the target probability can be understood as the probability value corresponding to the maximum probability path corresponding to the current lyric text in the lyric text.
[0107] 307. Calculate the pronunciation quality score of the target word according to the state transition path and the target probability, determine the confidence of the timestamp of the target word, and correct the timestamp of the word with low confidence to obtain the lyric timestamp of the current lyric text.
[0108] The device for generating lyric timestamps can calculate the pronunciation quality score of a target word based on the state transition path and the target probability, determine the confidence level of the timestamp of the target word, and correct the timestamps of words with low confidence levels to obtain the lyric timestamps of the current lyric text. Specifically, the quotient can be obtained by dividing the target probability of each frame of phoneme state in the current lyric text by the probability of the corresponding phoneme state in the state transition path, and then taking the logarithm and the negative value of the quotient to obtain the pronunciation quality score of each audio frame in the current lyric text. This pronunciation quality score can be understood as the GOP score. That is, the probability value of each frame in the free decoding can be divided by the probability value of the corresponding forced alignment decoding, and then taking the logarithm and the negative value to obtain the GOP score of each frame. The average value of the pronunciation quality scores of the audio frames corresponding to the words in the current lyric text is taken to obtain the pronunciation quality score of the word; that is, the word GOP score is obtained by taking the average value of the frames it belongs to. If the pronunciation quality score of the target word is greater than the preset threshold, it is determined that the timestamp of the target word is a high-confidence word timestamp; if the pronunciation quality score of the target word is less than the preset threshold, it is determined that the timestamp of the target word is a low-confidence word timestamp. This preset threshold can be 0.4 or 0.5, and specific details are not limited here. That is, a threshold k is set, and the words with a word GOP score greater than k are filtered out as high-confidence word timestamps, and those lower than k are filtered out as low-confidence word timestamps.
[0109] Meanwhile, for the timestamps of words with low confidence levels, voice activity detection (VAD) can be performed on each word. The start time and end time of the sub-timestamps with low confidence levels can be determined; considering the possibility of early entry (from music to vocals, and the vocal timestamp enters early), if the first signal energy of the preset frames before the start time is less than the preset energy threshold, the start time is shifted back one frame to correct the sub-timestamps with low confidence levels until the first signal energy is greater than the preset energy threshold. This preset energy threshold can be 60HZ or 70HZ, and specific details are not limited here. For example, if the start time of the current low-confidence word timestamp is t1 and the end time is t2. Calculate the average energy of the 3 frames before t1. This average energy refers to the square of the time-domain value of the audio signal, representing the signal energy size; if the average energy is less than the energy threshold p, then t1 is moved backward one frame, and this is repeated until the average energy is greater than p. And considering the possibility of early truncation (from vocals to music, and the vocal timestamp ends early), if the second signal energy of the preset frames after the end time is greater than the preset energy threshold, the end time is shifted forward one frame to correct the sub-timestamps with low confidence levels until the second signal energy is less than the preset energy threshold. For example, if the average energy of the 3 frames after t2 is calculated and is greater than p, then it is moved backward one frame, and this is repeated until it is less than p. The corrected sub-timestamps with low confidence levels and the high-confidence word timestamps are used as the lyric timestamps of the current lyric text. That is, after obtaining the sub-timestamps with low confidence levels corrected by VAD, they and the high-confidence timestamps together form the final lyric timestamp result.
[0110] It can be seen that, in the embodiments of the present application, by calculating the confidence level and combining with the VAD strategy, a fallback strategy for lyric timestamps is formulated, so that the situation of the current lyric text deviation can be repaired, ensuring the reliability of the generated lyric timestamps, improving the performance of the automatic alignment model, and reducing the cost of manually annotating lyrics at the same time.
[0111] The embodiments of the present application also provide a device for generating lyric timestamps, as Figure 6 shown, including:
[0112] An acquisition unit 601, configured to acquire the lyric text corresponding to the target song and the target dry audio of the target song;
[0113] A first determination unit 602, configured to determine a phoneme state sequence corresponding to the lyric text according to the phonemes corresponding to each word in the lyric text; the phoneme state sequence includes a plurality of phoneme states corresponding to the audio frames of the target dry audio;
[0114] A second determination unit 603, configured to determine the state transition probability between adjacent two-frame phoneme states in the phoneme state sequence;
[0115] An input unit 604, configured to input the target dry audio and the phoneme state sequence into a pre-trained acoustic model to obtain a first assignment probability that each frame of audio corresponding to the lyric text is assigned to each phoneme state in the phoneme state sequence, and a second assignment probability that each frame of audio corresponding to the current lyric text is assigned to each phoneme state corresponding to the current lyric text; the current lyric text is a segment of text in the lyric text;
[0116] A third determination unit 605, configured to determine the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability in multiple frames of audio corresponding to the current lyric text according to the second assignment probability and the state transition probability;
[0117] An execution unit 606, configured to determine the target probability of phoneme state transition in multiple frames of audio corresponding to the current lyric text according to the first assignment probability corresponding to each frame of phoneme state in the state transition path and the state transition probability between adjacent two-frame phoneme states in the state transition path;
[0118] A fourth determination unit 607, configured to determine the confidence level of the word timestamps in the current lyric text according to the maximum probability and the target probability;
[0119] A generation unit 608, configured to generate lyric timestamps of the current lyric text according to the confidence level of the word timestamps.
[0120] The embodiments of the present application also provide a device 700 for generating lyrics timestamps, as Figure 7 shown, including:
[0121] A central processing unit 701, a memory 702, and an input / output interface 703;
[0122] The memory 702 is a transient storage memory or a persistent storage memory;
[0123] The central processing unit 701 is configured to communicate with the memory 702 and execute the instruction operations in the memory 702 to perform the above-mentioned generation method.
[0124] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0125] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.
[0126] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0128] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
Claims
1. A method for generating lyrics timestamps, characterized in that, it includes: Obtain the lyrics text corresponding to the target song and the target dry audio of the target song; According to the phonemes corresponding to each character in the lyrics text, determine the phoneme state sequence corresponding to the lyrics text; the phoneme state sequence includes a plurality of phoneme states corresponding to the audio frames of the target dry audio; Determine the state transition probability between adjacent two-frame phoneme states in the phoneme state sequence; Input the target dry audio and the phoneme state sequence into a pre-trained acoustic model to obtain the first assignment probability that each frame of audio corresponding to the lyrics text is assigned to each phoneme state in the phoneme state sequence, and the second assignment probability that each frame of audio corresponding to the current lyrics text is assigned to each phoneme state corresponding to the current lyrics text; the current lyrics text is a segment of text in the lyrics text; According to the second assignment probability and the state transition probability, determine the maximum probability of phoneme state transition in multiple frames of audio corresponding to the current lyrics text and the state transition path corresponding to the maximum probability; According to the first assignment probability corresponding to each frame of phoneme state in the state transition path and the corresponding state transition probability, determine the target probability of each frame of phoneme state corresponding to the current lyrics text; Determine the confidence of the word timestamps in the current lyrics text according to the state transition path and the target probability, and generate the lyrics timestamps of the current lyrics text according to the confidence of the word timestamps.
2. The generation method according to claim 1, characterized in that, the obtaining of the lyrics text corresponding to the target song includes: Obtain the initial lyrics text corresponding to the target song; Perform non-lyrics information filtering processing on the initial lyrics text to obtain the lyrics text.
3. The generation method according to claim 1, characterized in that, the determining of the phoneme state sequence corresponding to the lyrics text according to the phonemes corresponding to each character in the lyrics text includes: Obtain the phonemes corresponding to the audio frames in the target dry audio to obtain the phoneme sequence corresponding to the target dry audio; Match the phonemes corresponding to each character in the lyrics text with the phoneme sequence input into a pre-set language model to obtain the phonemes corresponding to each character in the lyrics text in the phoneme sequence; Convert the phonemes corresponding to each character in the lyrics text in the phoneme sequence into phoneme states to obtain the phoneme state sequence corresponding to the lyrics text.
4. The generation method according to claim 1, characterized in that, the determining of the state transition probability between adjacent two-frame phoneme states in the phoneme state sequence includes: Take multiple phoneme states of adjacent preset frames as a multi-phoneme state group, and determine the total number of phoneme states in the phoneme state sequence and the number of occurrences of the multi-phoneme state group in the phoneme state sequence; Divide the number of occurrences by the total number of phoneme states to obtain the occurrence probability of the multi-phoneme state group in the phoneme state sequence; Use the occurrence probability as the state transition probability between adjacent two-frame phoneme states in the multi-phoneme state group.
5. The generation method according to claim 1, wherein, the step of inputting the target dry audio and the phoneme state sequence into a pre-trained acoustic model to obtain the first assignment probability and the second assignment probability includes: extracting the audio features of the target dry audio, and inputting the audio features of the target dry audio and the phoneme states in the phoneme state sequence into the pre-trained acoustic model to obtain the first assignment probability; extracting the audio features of the dry audio corresponding to the current lyric text, and inputting the audio features of the dry audio and the phoneme states corresponding to the current lyric text into the pre-trained acoustic model to obtain the second assignment probability.
6. The generation method according to claim 1, wherein, the step of determining the maximum probability of phoneme state transition and the state transition path corresponding to the maximum probability in the multi-frame audio corresponding to the current lyric text according to the second assignment probability and the state transition probability includes: in the multi-frame audio corresponding to the current lyric text, calculating the sum of the products of the second assignment probability of the phoneme state of the current frame and the state transition probability of the phoneme state of the current frame transitioning to the phoneme state of the next frame; taking the maximum value in the sum of the products as the maximum probability of the phoneme state transition; determining the target phoneme state of each frame of audio corresponding to the current lyric text in the maximum probability, and taking the multiple target phoneme states corresponding to the multi-frame audio as the state transition path corresponding to the maximum probability.
7. The generation method according to claim 1, wherein, the step of determining the target probability of each frame of phoneme state corresponding to the current lyric text according to the first assignment probability and the corresponding state transition probability of each frame of phoneme state in the state transition path includes: calculating the product of the first assignment probability of each frame of phoneme state in the state transition path and the corresponding state transition probability to obtain the target probability of each frame of phoneme state corresponding to the current lyric text.
8. The generation method according to claim 1, wherein, the step of determining the confidence of the word timestamp in the current lyric text according to the state transition path and the target probability includes: dividing the target probability of each frame of phoneme state in the current lyric text by the probability of the corresponding phoneme state in the state transition path to obtain a quotient, taking the logarithm of the quotient and taking the negative value to obtain the pronunciation quality score of each audio frame in the current lyric text; taking the average value of the pronunciation quality scores of the audio frames corresponding to the words in the current lyric text to obtain the pronunciation quality score of the word; if the pronunciation quality score of the target word is greater than the preset threshold, determining that the timestamp of the target word is a high-confidence word timestamp; if the pronunciation quality score of the target word is less than the preset threshold, determining that the timestamp of the target word is a low-confidence word timestamp.
9. The generation method according to claim 8, wherein, Generating the lyric timestamp of the current lyric text according to the confidence of the word timestamp includes: Determining the start time and end time of the sub-timestamp with low confidence; If the first signal energy of a preset frame before the start time is less than a preset energy threshold, shifting the start time backward by one frame to correct the sub-timestamp with low confidence until the first signal energy is greater than the preset energy threshold; If the second signal energy of a preset frame after the end time is greater than the preset energy threshold, shifting the end time forward by one frame to correct the sub-timestamp with low confidence until the second signal energy is less than the preset energy threshold; Taking the corrected sub-timestamp with low confidence and the word timestamp with high confidence as the lyric timestamp of the current lyric text.
10. A device for generating lyric timestamps, characterized in that, it includes: A central processing unit, a memory, and an input / output interface; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that, it includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Humming type rhythm identification method based on hidden Markov model
CN101504834A
Methods and Systems for Performing Synchronization of Audio with Corresponding Textual Transcriptions and Determining Confidence Values of the Synchronization
US20110288862A1