Language model training method, audio recognition method and computer device

By constructing a hybrid pronunciation dictionary and training a language model, the problem of low accuracy when identifying lyric information of multiple languages ​​in the prior art is solved, and efficient recognition of lyric information of multiple languages ​​and audio styles is achieved.

CN114613359BActive Publication Date: 2025-05-16TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210331883.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-05-16
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

When the prior art recognizes lyric information of multiple languages ​​in songs, it is easy to lead to recognition errors and reduces recognition accuracy.

Method used

By obtaining sample audio and lyrics in multiple languages ​​and audio styles, adding language and audio style logos, building a hybrid pronunciation dictionary, and training the language model based on this to identify the lyric pronunciation sequence of input audio.

Benefits of technology

Improves the accuracy of audio recognition and can effectively identify lyric information containing multiple languages ​​and different audio styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114613359B_ABST
    Figure CN114613359B_ABST
Patent Text Reader

Abstract

The present application relates to a language model training method, an audio recognition method, an apparatus, a computer device, a storage medium and a computer program product. The audio features in the audio to be recognized and the pronunciation status of each frame of audio are input into a target language model, wherein the target language model is trained based on a mixed pronunciation dictionary, and the mixed pronunciation dictionary contains multiple sample phonemes with language and style identifiers, and the target language model is used to identify the association between each frame of audio, and the lyrics pronunciation sequence of the audio to be recognized is determined, thereby identifying the lyrics text corresponding to the audio to be recognized according to the lyrics pronunciation sequence. Compared with the traditional method of identifying the audio language and then identifying the audio through the corresponding language model, this solution uses a language model trained based on a mixed pronunciation dictionary, and jointly identifies the lyrics information in the audio based on the language and genre, which can improve the accuracy of audio recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to a language model training method, an audio recognition method, an apparatus, a computer device, a storage medium, and a computer program product. Background Art

[0002] With the development of computer technology, users can now play music and sing songs through mobile terminals such as mobile phones. However, for songs without words or audio input by users when singing songs, lyrics recognition is required to identify the lyrics information. Since the same song can contain lyrics in multiple languages, it is necessary to identify the lyrics information in each language of the song separately. The current method of identifying the lyrics information in a song is usually to identify the language of the song and then use the recognition model of the corresponding language to recognize the song audio. However, the method of identifying the audio language and then recognizing the audio through the corresponding language model is prone to errors in identifying the lyrics information and reducing the recognition accuracy.

[0003] Therefore, the current audio recognition method for songs has the defect of low recognition accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide a language model training method, audio recognition method, device, computer equipment, computer-readable storage medium and computer program product that can improve recognition accuracy in response to the above technical problems.

[0005] In a first aspect, the present application provides a language model training method, the method comprising:

[0006] Obtaining a plurality of sample audios and sample lyrics for each of the sample audios, wherein the plurality of sample audios correspond to a plurality of audio styles and the sample lyrics corresponding to the sample audios include a plurality of languages;

[0007] Adding a language identifier and an audio style identifier to the sample phonemes included in the sample lyrics according to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics;

[0008] Constructing a mixed pronunciation dictionary based on the sample phonemes with added labels corresponding to the sample lyrics of the plurality of sample audios;

[0009] The language model to be trained is trained according to the mixed pronunciation dictionary to obtain a target language model, and the target language model is used to identify the lyrics pronunciation sequence of the input audio based on the language and audio style of the input audio.

[0010] In one embodiment, obtaining sample lyrics for each sample audio includes:

[0011] The audio lyrics corresponding to the sample audio are obtained, and the audio lyrics are deduplicated to obtain the sample lyrics.

[0012] In one embodiment, adding a language identifier and an audio style identifier to the sample phonemes included in the sample lyrics according to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics includes:

[0013] Obtaining sample phonemes contained in the sample lyrics according to a mapping relationship between the sample lyrics and standard phonetic symbols;

[0014] A language identifier and an audio style identifier are added to the sample phonemes according to the language corresponding to the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics.

[0015] In a second aspect, the present application provides an audio recognition method, the method comprising:

[0016] Obtain audio features of each frame of audio to be recognized;

[0017] Acquiring the pronunciation state of each frame of audio according to the audio features of each frame of audio;

[0018] Inputting the audio features and pronunciation status of each frame of audio into a target language model, identifying the association between each frame of audio through the target language model and determining the lyrics pronunciation sequence of the audio to be identified based on the association; the target language model is trained according to the above method;

[0019] According to the lyrics pronunciation sequence, the lyrics text corresponding to the audio to be recognized is identified.

[0020] In one embodiment, the pronunciation state includes a pronunciation start state, a pronunciation middle state, and a pronunciation end state;

[0021] The obtaining, based on the audio features of each frame of audio, the pronunciation state of each frame of audio, includes:

[0022] For each frame of audio in the audio to be recognized, inputting audio features corresponding to the frame of audio into a preset state recognition model, and obtaining a first probability that the frame of audio is in a pronunciation start state, a second probability that the frame of audio is in a pronunciation middle state, and a third probability that the frame of audio is in a pronunciation end state, output by the preset state recognition model;

[0023] The pronunciation state corresponding to the maximum value among the first probability, the second probability and the third probability is used as the pronunciation state to which the frame of audio belongs.

[0024] In one embodiment, inputting the audio features and pronunciation status of each frame of audio into a target language model, identifying the association between each frame of audio through the target language model, and determining the lyrics pronunciation sequence of the audio to be identified based on the association, includes:

[0025] For each frame of audio in the audio to be recognized, inputting the audio features of the frame of audio and its corresponding pronunciation state into the target language model;

[0026] Identifying the language of the lyrics corresponding to the audio feature and the audio style of the audio to be identified corresponding to the audio feature through the target language model, and determining multiple associated pronunciation states of the frame of audio and a transition probability of each associated pronunciation state based on the identified language, audio style, and pronunciation state of the frame of audio; the associated pronunciation state represents the pronunciation state of the next frame of audio corresponding to the frame of audio; and the transition probability represents the probability of transitioning from the pronunciation state of the frame of audio to the associated pronunciation state;

[0027] Constructing multiple state sequences to be identified based on the pronunciation state corresponding to each frame of audio and the multiple associated pronunciation states corresponding to each frame of audio; each state sequence to be identified includes the pronunciation states corresponding to multiple frames of audio, and the pronunciation state of two adjacent frames of audio is the associated pronunciation state of the former;

[0028] According to the transition probability between the pronunciation states of two adjacent frames of audio in the state sequence to be recognized, the lyrics pronunciation sequence of the audio to be recognized is determined from a plurality of the state sequences to be recognized.

[0029] In one embodiment, determining the lyrics pronunciation sequence of the audio to be recognized from a plurality of state sequences to be recognized based on the transition probabilities between the pronunciation states of adjacent frames of audio in the state sequence to be recognized includes:

[0030] Determining a target state sequence from a plurality of state sequences to be identified based on a transition probability between two adjacent frames of pronunciation states in the state sequence to be identified;

[0031] The target state sequence is converted into a phoneme sequence as the lyrics pronunciation sequence of the audio to be recognized.

[0032] In one embodiment, determining a target state sequence from a plurality of state sequences to be identified based on the transition probabilities between pronunciation states of adjacent frames of audio in the state sequence to be identified includes:

[0033] Obtaining the product of transition probabilities between pronunciation states of two adjacent frames of audio in the state sequence to be identified;

[0034] A state sequence to be identified corresponding to the maximum value of the product is determined from a plurality of state sequences to be identified as a target state sequence.

[0035] In one embodiment, converting the target state sequence into a phoneme sequence as the lyrics pronunciation sequence of the audio to be recognized includes:

[0036] Inputting the target state sequence into the target language model, the target language model identifying associated phonemes corresponding to each known language and known audio style for each pronunciation state combination in the target state sequence and probabilities of the associated phonemes, thereby obtaining a plurality of associated phonemes and their probabilities for each pronunciation state combination; the pronunciation state combination includes a pronunciation start stage, a pronunciation middle stage, and a pronunciation end stage;

[0037] Determining multiple associated phoneme sequences according to the arrangement results between the associated phonemes corresponding to the multiple pronunciation state combinations in the target state sequence;

[0038] Obtaining the product of the probabilities of two adjacent associated phonemes in the associated phoneme sequence;

[0039] The lyrics pronunciation sequence of the audio to be recognized is obtained according to the associated phoneme sequence corresponding to the maximum value of the product of the probabilities.

[0040] In a third aspect, the present application provides a language model training device, the device comprising:

[0041] An audio acquisition module, configured to acquire a plurality of sample audios and sample lyrics for each of the sample audios, wherein the plurality of sample audios correspond to a plurality of audio styles and the sample lyrics corresponding to the sample audios include a plurality of languages;

[0042] An adding module, configured to add a language identifier and an audio style identifier to the sample phonemes contained in the sample lyrics according to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics;

[0043] The construction module is used to construct a mixed pronunciation dictionary based on the sample phonemes with added marks corresponding to the sample lyrics of the multiple sample audios.

[0044] The training module is used to train the language model to be trained according to the mixed pronunciation dictionary to obtain a target language model, and the target language model is used to recognize the lyrics pronunciation sequence of the input audio based on the language and audio style of the input audio.

[0045] In a fourth aspect, the present application provides an audio recognition device, comprising:

[0046] A state acquisition module is used to obtain audio features of each frame of audio to be recognized; and obtain the pronunciation state of each frame of audio according to the audio features of each frame of audio;

[0047] a determination module, configured to input the audio features and pronunciation status of each frame of audio into a target language model, identify the association between each frame of audio through the target language model, and determine the pronunciation sequence of lyrics of the audio to be identified based on the association; the target language model is trained according to the language model training method described above;

[0048] The recognition module is used to identify the lyrics text corresponding to the audio to be recognized based on the lyrics pronunciation sequence.

[0049] In a fifth aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0050] In a sixth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0051] In a seventh aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.

[0052] The above-mentioned language model training method, audio recognition method, apparatus, computer equipment, storage medium and computer program product obtain sample lyrics of sample audio containing lyrics in multiple languages, and add corresponding labels to sample phonemes corresponding to the sample lyrics according to the language and audio style of the sample lyrics. Based on the sample phonemes corresponding to the sample lyrics of multiple sample audios with added labels, a mixed pronunciation dictionary is constructed, thereby training the language model to be trained based on the mixed pronunciation dictionary to obtain a target language model; and by inputting the audio features in the audio to be recognized and the pronunciation status of each frame of audio into the target language model, the target language model identifies the association between each frame of audio, determines the pronunciation sequence of the lyrics of the audio to be recognized, and thereby recognizes the lyrics text corresponding to the audio to be recognized based on the lyrics pronunciation sequence. Compared with the traditional method of identifying the audio language and then recognizing the audio through the corresponding language model, this solution utilizes a language model trained based on a mixed pronunciation dictionary and jointly recognizes the lyrics information in the audio based on the language and genre, which can improve the accuracy of audio recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A diagram illustrating an application environment of a language model training method in one embodiment;

[0054] Figure 2 1 is a flow chart of a language model training method according to an embodiment;

[0055] Figure 3 1 is a flow chart of an audio recognition method according to an embodiment;

[0056] Figure 4 Schematic diagram of a flow chart of a state identification step in one embodiment;

[0057] Figure 5 Schematic diagram of a process of state transfer steps in one embodiment;

[0058] Figure 6 is a flowchart of an audio recognition method according to another embodiment;

[0059] Figure 7 is a structural block diagram of a language model training device in one embodiment;

[0060] Figure 8 is a structural block diagram of an audio recognition device in one embodiment;

[0061] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0063] The language model training method and audio recognition method provided in the embodiments of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can obtain the audio to be recognized input by the user, and send the audio to be recognized to the server 104. The server 104 can input the audio to be recognized into the target language model, and recognize the lyrics pronunciation sequence in the audio to be recognized through the target language model trained based on the mixed pronunciation dictionary, wherein the mixed pronunciation dictionary is constructed based on phonemes that identify language and audio style information. The server 104 can recognize the text information therein based on the lyrics pronunciation sequence to realize text recognition of the audio. Among them, the terminal 102 can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.

[0064] In one embodiment, Figure 2 As shown, a language model training method is provided, which is applied to Figure 1 The following steps are used as an example to illustrate the server in the example:

[0065] In step S202 , a plurality of sample audios and sample lyrics corresponding to each sample audio are obtained, wherein the plurality of sample audios correspond to a plurality of audio styles and the sample lyrics corresponding to the sample audios include a plurality of languages.

[0066] The sample audio may be an audio with known lyrics information, and the lyrics information in the sample audio may include multiple languages. For example, the sample audio may be a song audio containing Chinese and English lyrics. Server 104 may obtain multiple sample audios. The above-mentioned multiple sample audios may include lyrics in multiple languages. For example, taking the sample audio as a music song as an example, a sample audio may contain lyrics in multiple languages, such as Chinese-English mixed lyrics containing both Chinese lyrics and English lyrics, etc., wherein the multiple languages ​​may be two or more than two; and the above-mentioned multiple audios also correspond to multiple audio styles. For example, taking the sample audio as a music song as an example, each sample audio may have its corresponding audio style, such as pop music, hip-hop music, heavy metal music or other music. It should be noted that the audio style may also be a style other than the above-mentioned audio style.

[0067] Step S204 : adding a language identifier and an audio style identifier to the sample phonemes included in the sample lyrics according to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics.

[0068] The sample lyrics may be the lyrics of each sample audio obtained above, and there may be multiple sample audios, and each sample audio may correspond to a set of sample lyrics. The server 104 may obtain the sample lyrics corresponding to the sample audio after processing the lyrics corresponding to the sample audio. For example, in one embodiment, obtaining the sample lyrics of each sample audio includes: obtaining the audio lyrics corresponding to the sample audio, deduplicating the audio lyrics, and obtaining the sample lyrics. In this embodiment, the server 104 may obtain the audio lyrics corresponding to the sample audio, and deduplicate the audio lyrics to obtain the sample lyrics. For example, taking the sample audio as music audio, the sample audio may include mixed Chinese and English lyrics. The server 104 may separate the Chinese lyrics and English lyrics therefrom, and deduplicate the lyrics to obtain a word set of the sample lyrics.

[0069] After the server 104 obtains the sample lyrics of the above-mentioned sample audio, since the sample lyrics are in different languages ​​and the corresponding sample audios have different audio styles, for example, for a, its pronunciation in Chinese lyrics and English lyrics will be different, and its pronunciation in different audio styles will also be different. Therefore, the server 104 can also add corresponding identifiers to the sample phonemes corresponding to the sample lyrics according to the language of the sample lyrics and the above-mentioned audio style. Among them, the phoneme is the smallest speech unit divided according to the natural properties of the speech. The server 104 can first convert the sample lyrics into corresponding sample phonemes, and based on the language of the sample lyrics corresponding to the sample phonemes and the above-mentioned audio style, add corresponding identifiers to the sample phonemes, so as to distinguish between sample phonemes in different languages ​​and different audio styles.

[0070] Step S206 : constructing a mixed pronunciation dictionary based on the sample phonemes with added labels corresponding to the sample lyrics of the multiple sample audios.

[0071] There may be multiple sample audios, and thus multiple sample lyrics corresponding to the sample audios, and thus multiple sample phonemes. After server 104 identifies the multiple sample phonemes, it can construct a mixed pronunciation dictionary based on the identified sample phonemes. Specifically, the mixed pronunciation dictionary contains multiple sample phonemes, each with its corresponding language identifier and audio style identifier. This allows server 104 to train a language model based on the mixed pronunciation dictionary.

[0072] Step S208 , training the language model to be trained according to the mixed pronunciation dictionary to obtain a target language model, which is used to recognize the lyrics pronunciation sequence of the input audio based on the language and audio style of the input audio.

[0073] The mixed pronunciation dictionary may include multiple sample phonemes with language identifiers and audio style identifiers. Server 104 can train a language model to be trained based on the mixed pronunciation dictionary to obtain a target language model. The language model to be trained can be an N-Gram language model, which is an algorithm based on a statistical language model. Its basic concept is to perform a sliding window operation of size N on the content of the text according to bytes, forming a sequence of byte segments of length N. Since the above-mentioned sample phonemes have their corresponding language identifiers and audio style identifiers, server 104 can train the language model to be trained using sample phonemes with language identifiers and audio style identifiers to obtain a target language model for outputting a lyrics pronunciation sequence corresponding to the language and audio style of the input audio. The input audio can be audio input by the user, and the lyrics pronunciation sequence can be the pronunciation sequence of the lyrics corresponding to the audio input by the user. The pronunciation sequence can be composed of multiple phonemes. Specifically, taking the above-mentioned sample audio as a music song as an example, music songs can be divided into different genres, that is, different audio styles, and music songs can contain mixed Chinese and English lyrics. After the server 104 constructs a mixed pronunciation dictionary based on the Chinese and English sample lyrics based on the music genre, it can use the large amount of lyrics text in the mixed pronunciation dictionary to train the above-mentioned language model to be trained, thereby obtaining the target language model.

[0074] In the above-mentioned language model training method, sample lyrics of sample audio containing lyrics in multiple languages ​​are obtained, and language identifiers and audio style identifiers are added to sample phonemes corresponding to the sample lyrics according to the language and audio style of the sample lyrics. Based on the sample phonemes corresponding to the sample lyrics of multiple sample audios with added identifiers, a mixed pronunciation dictionary is constructed, and the language model to be trained is trained based on the mixed pronunciation dictionary to obtain a target language model; thus, the terminal can use the target language model to recognize the pronunciation sequence of lyrics in audio containing different audio styles in multiple languages. Compared with the traditional method of identifying the audio language and then identifying the audio through the corresponding language model, the language model trained based on the mixed pronunciation dictionary in this solution can jointly identify the lyrics information in the audio based on the language and genre, which can improve the accuracy of audio recognition.

[0075] In one embodiment, language identification and audio style identification are added to sample phonemes contained in the sample lyrics according to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics, including: obtaining the sample phonemes contained in the sample lyrics according to the mapping relationship between the sample lyrics and the standard phonetic symbols; adding language identification and audio style identification to the sample phonemes according to the language corresponding to the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics.

[0076] In this embodiment, the sample lyrics are lyrics corresponding to the sample audio. The sample lyrics have corresponding language information, and the corresponding sample audio also has a corresponding audio style. Server 104 can classify and distinguish lyrics in different languages ​​from the sample audio of different styles. Server 104 can obtain sample phonemes corresponding to the sample lyrics based on the mapping relationship between the sample lyrics and standard phonetic symbols. After obtaining the sample phonemes, server 104 can add corresponding identifiers to the sample phonemes based on the language corresponding to the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics, thereby adding corresponding language identifiers and audio style identifiers to the sample phonemes. The standard phonetic symbols can be the International Phonetic Alphabet. For example, if the sample audio is a song, server 104 can map the sample lyrics to the corresponding phonemes based on the mapping relationship between the International Phonetic Alphabet to obtain the sample phonemes corresponding to each sample lyric. Server 104 can then construct an original mixed dictionary based on the multiple sample phonemes. Server 104 can add corresponding identifiers to each sample phoneme in the original mixed dictionary based on the differences in language and audio style between the sample phonemes, thereby distinguishing phonemes of different languages ​​and audio styles.

[0077] Specifically, the server 104 can add different identifiers to the sample phonemes corresponding to the sample lyrics based on the language and audio style of the sample lyrics. For example, in one embodiment, the sample phonemes are added with corresponding identifiers based on the language corresponding to the sample lyrics and the audio style corresponding to the sample audio, including: adding a language identifier to the sample phoneme based on the language of the sample lyrics to obtain a first identified sample phoneme; adding an audio style identifier to the first identified sample phoneme based on the audio style of the sample audio corresponding to the first identified sample phoneme to obtain a second identified sample phoneme as the identified phoneme. In this embodiment, the server 104 can obtain the language corresponding to the sample lyrics, wherein the sample lyrics may include multiple languages, such as Chinese and English lyrics. The server 104 can then add a language identifier to the corresponding sample phoneme based on the language of each sample lyric to obtain the first identified sample phoneme. Specifically, for the same lyric a, its pronunciation will be different in different languages. For Chinese lyrics, the server 104 can add a "_chn" suffix to the sample phonemes corresponding to the sample lyrics; for English lyrics, the server 104 can add a "_eng" suffix to the sample phonemes corresponding to the sample lyrics, so that the server 104 can distinguish the differences between the same phonemes in Chinese and English languages, ensuring that the same phonemes can correspond to two different language decoding paths in the recognition stage.

[0078] After obtaining the aforementioned first-labeled sample phonemes with language identifiers, server 104 can obtain the audio style of the sample audio corresponding to each first-labeled sample phoneme. Based on the audio style of each sample audio, it can add the corresponding audio style identifier to the corresponding first-labeled sample audio, thereby distinguishing the audio styles of the sample phonemes, thereby obtaining second-labeled sample phonemes. Specifically, the second-labeled sample phonemes can include both the language identifier and the audio style identifier. Server 104 can use the second-labeled sample phonemes as the labeled phonemes for constructing the mixed pronunciation dictionary. Specifically, taking the aforementioned sample audio as a song, the aforementioned audio style can be a musical genre, including pop, hip-hop, heavy metal, and other music. Server 104 can add a corresponding audio style identifier to each of the aforementioned sample phonemes with language identifiers, such as adding a "_pop" suffix, a "_hiphop" suffix, a "_metal" suffix, or a "_other" suffix, to indicate the pronunciation differences of the same phoneme in different genres. Other music can be music that does not fall into the aforementioned three audio styles. Server 104 further distinguishes lyrics in different languages ​​based on genre. During the step of determining the lyrics pronunciation sequence, the decoding path can be further subdivided based on the language-based decoding path search. That is, the same phoneme can be subdivided from a single path into multiple paths after distinguishing between language and music genre. This allows server 104 to construct a mixed pronunciation dictionary based on sample phonemes in different languages ​​and audio styles, and train a target language model to achieve text recognition based on the target language model for audio containing lyrics in multiple languages ​​and different audio styles.

[0079] Through the above embodiments, the server 104 can add language identifiers to the phonemes corresponding to the sample lyrics in different languages, and add audio style identifiers to the phonemes of audios with different audio styles, so that the server 104 can build a mixed pronunciation dictionary based on the phonemes containing language identifiers and audio style identifiers, and obtain the target language model based on the mixed pronunciation dictionary training, so as to realize the recognition of the pronunciation sequence of lyrics of audios with mixed languages ​​and different audio styles, thereby improving the accuracy of audio recognition.

[0080] In one embodiment, Figure 3 As shown, an audio recognition method is provided, which is applied to Figure 1 The following steps are used as an example to illustrate the server in the example:

[0081] Step S302: obtaining audio features of each frame of audio in the audio to be recognized; and obtaining the pronunciation state of each frame of audio according to the audio features of each frame of audio.

[0082] The audio to be identified may be audio input by the user, and the user may record the audio to be identified through the terminal 102, so that the terminal 102 may send the audio to be identified to the server 104. The server 104 may obtain the user's audio to be identified and extract the audio features of the audio to be identified therefrom. The audio to be identified may be audio containing lyrics in multiple languages. For example, the audio to be identified may be a song containing mixed Chinese and English lyrics, and the server 104 may extract audio features from the mixed Chinese and English song. Specifically, the audio features may be MFCC (Mel Frequency Cepstrum Coefficient), that is, the server 104 may extract the Mel frequency cepstrum coefficients from the audio to be identified. The Mel frequency is proposed based on the auditory characteristics of the human ear, and it has a nonlinear correspondence with the Hz frequency. The Mel frequency cepstrum coefficient (MFCC) is the Hz spectrum feature calculated by utilizing this relationship between them.

[0083] After the server 104 extracts the audio features in the audio to be identified, it can obtain the pronunciation state of each frame of audio in the audio to be identified based on the audio features. The pronunciation state can represent the pronunciation of different parts of the pronunciation process of a phoneme in the audio. For example, for the phoneme "w", during its pronunciation process, it can correspond to three pronunciation states, including the pronunciation start state, the pronunciation middle state and the pronunciation end state, that is, the pronunciation state is the smallest pronunciation unit, and the pronunciation state can include the pronunciation start state, the pronunciation middle state and the pronunciation end state. The above-mentioned pronunciation state can be identified based on a preset state recognition model, wherein the above-mentioned state recognition model can be a neural network model.

[0084] In step S304, the audio features and pronunciation status of each frame of audio are input into the target language model, the association relationship between each frame of audio is identified through the target language model, and the pronunciation sequence of the lyrics of the audio to be identified is determined based on the association relationship; the target language model is trained according to the language model training method as described above.

[0085] Among them, the above audio features can be the MFCC coefficients of the audio. The server 104 can segment the to-be-recognized audio into frames to obtain multiple frames of audio. Moreover, each frame of audio corresponds to a pronunciation state. The server 104 can input the audio features of each frame of audio and its corresponding pronunciation state into the target language model, and through the target language model, identify the association relationship between each frame of audio for the input audio features and their pronunciation states. For example, the server 104 can identify the degree of association between the audio features of each frame of input audio and the degree of association between the pronunciation states of each frame of audio through the sample phonemes with language identifiers and audio style identifiers in the target language model. Thus, the target language model can combine the degree of association between audio features and the degree of association between pronunciation states to determine the association relationship between each frame of audio, and based on this association relationship, determine the lyric pronunciation sequence of the to-be-recognized audio. Among them, the above target language model can be a model trained by the above language model training method. The lyric pronunciation sequence can be a sequence composed of phonemes corresponding to multiple frames of audio. After the server 104 identifies the optimal state sequence including the pronunciation states of multiple frames through the target language model, it can then convert the state sequence into a lyric pronunciation sequence.

[0086] Step S306, according to the lyric pronunciation sequence, identify the lyric text corresponding to the to-be-recognized audio.

[0087] Among them, the lyric pronunciation sequence can be a phoneme sequence composed of phonemes including multiple frames. After the server 104 obtains the above lyric pronunciation sequence, it can convert the lyric pronunciation sequence into text, so as to obtain the final lyric recognition result. Among them, since one phoneme can correspond to multiple characters, for example, the phoneme "wan" can correspond to characters such as "玩 (play)", "万 (ten thousand)", or "晚 (evening)", etc., thus the server 104 can identify the language to which the phoneme in the phoneme sequence belongs and the audio style of its corresponding audio through the above target language model, and thus determine the character with the highest probability based on the language and audio style. The server 104 can identify multiple candidate characters for each phoneme through the target language model. These candidate characters can include characters of different languages and with different audio styles of the corresponding audio. There can be corresponding probabilities between the phoneme and each of its corresponding candidate characters. Then the server 104 can form multiple candidate character sequences based on the candidate characters associated with each phoneme in the phoneme sequence. Thus, the server 104 can calculate the product of the probabilities corresponding to each candidate character in these candidate character sequences, and obtain the maximum value of the product of the probabilities, and take the candidate character sequence corresponding to this maximum value as the final lyric recognition result, so as to obtain the lyric text of the to-be-recognized audio.

[0088] In the above-mentioned audio recognition method, the audio features of the audio to be recognized and the pronunciation status of each frame of audio are input into the target language model. The target language model, trained based on a mixed pronunciation dictionary, identifies the association between each frame of audio, determines the pronunciation sequence of the lyrics of the audio to be recognized, and then identifies the lyrics text corresponding to the audio to be recognized based on the lyrics pronunciation sequence. Compared to the traditional method of identifying the audio language and then using the corresponding language model to recognize the audio, this solution utilizes a language model trained based on a mixed pronunciation dictionary. The language model can jointly identify the lyrics information in the audio based on language and genre, which can improve the accuracy of audio recognition.

[0089] In one embodiment, based on the audio features of each frame of audio, the pronunciation state of each frame of audio is obtained, including: for each frame of audio in the audio to be identified, the audio features corresponding to the frame of audio are input into a preset state recognition model, and the first probability that the frame of audio is in the pronunciation start state, the second probability that the frame of audio is in the pronunciation middle state, and the third probability that the frame of audio is in the pronunciation end state output by the preset state recognition model are obtained; and the pronunciation state corresponding to the maximum value among the first probability, the second probability, and the third probability is used as the pronunciation state to which the frame of audio belongs.

[0090] In this embodiment, the pronunciation state may include a pronunciation start state, a pronunciation middle state, and a pronunciation end state. The server 104 may extract the audio features of each frame of audio in the audio to be identified, and identify the pronunciation state of each frame of audio based on the audio features of each frame of audio. The above-mentioned audio to be identified may include multiple frames of audio. For each frame of audio in the audio to be identified, the server 104 may input the audio features corresponding to the frame of audio into a preset state recognition model, and obtain the probability of the frame of audio corresponding to different pronunciation states output by the preset state recognition model, including a first probability that the frame of audio is in a pronunciation start state, a second probability that the frame of audio is in a pronunciation middle state, and a third probability that the frame of audio is in a pronunciation end state, so that the server 104 may determine the pronunciation state to which the frame of audio belongs based on these probabilities. For example, the server 104 may use the maximum value of the above-mentioned first probability, second probability, and third probability as the pronunciation state to which the frame of audio belongs, so that the server 104 may obtain the pronunciation state corresponding to each frame of audio in the audio to be identified by performing the above-mentioned processing on multiple frames of audio. Wherein, as Figure 4 As shown, Figure 4It is a flowchart of the state recognition step in one embodiment. The above-mentioned preset state recognition model can be a neural network model DNN (Deep Neural Networks, deep neural network), and the server 104 can input the above-mentioned audio features into the deep neural network. Each output node in the neural network represents an HMM (HiddenMarkov Model, hidden Markov model) state, that is, the above-mentioned pronunciation state can be an HMM state. For example, for a phoneme a_chn_pop with a language identifier and an audio style identifier, it can correspond to three HMM states, which respectively represent the pronunciation start stage, pronunciation middle stage, and pronunciation end stage of the phoneme, where a represents the phoneme name, chn represents the language identifier, and pop represents the audio style identifier; and each frame of audio can correspond to a state. The above-mentioned DNN model can be used as a state classifier to output the probability that each frame of audio in the audio to be identified belongs to each state. The probability can be the emission probability of the HMM model, such as Figure 4 In the above example, s1, ···, sk represent different HMM states, each of which can correspond to a transition probability and an observation probability. Layer N represents different layers in the above deep neural network model, and Speech features represents audio features. The server 104 can then use the above neural network model to obtain the probability of each pronunciation state corresponding to each frame in the audio to be recognized, and can also identify the optimal state path based on these probabilities. This identification can be achieved by identifying the transition probabilities of adjacent frame states using the above target language model.

[0091] Through this embodiment, the server 104 can identify the pronunciation state corresponding to each frame of audio based on the set state recognition model, so that the server 104 can determine the correlation between each frame of audio based on the pronunciation probability of each frame of audio, and thus perform text recognition on the audio to be recognized based on the correlation, thereby improving the accuracy of audio recognition.

[0092] In one embodiment, the audio features and pronunciation states of each frame of audio are input into a target language model, the association relationship between each frame of audio is identified through the target language model, and the lyrics pronunciation sequence of the audio to be identified is determined based on the association relationship, including: for each frame of audio in the audio to be identified, the audio features and the corresponding pronunciation state of the frame of audio are input into the target language model; the target language model is used to identify the language of the lyrics corresponding to the audio features and the audio style of the audio to be identified corresponding to the audio features, and based on the identified language, audio style and pronunciation state of the frame of audio, a plurality of associated pronunciation states of the frame of audio are determined, and each associated pronunciation state is determined. The transition probability of the sound state; the associated pronunciation state represents the pronunciation state of the next frame of audio corresponding to the current frame of audio; the transition probability represents the probability of converting from the pronunciation state of the current frame of audio to the associated pronunciation state; according to the pronunciation state corresponding to each frame of audio and the multiple associated pronunciation states corresponding to each frame of audio, multiple state sequences to be identified are constructed; each state sequence to be identified contains the pronunciation states corresponding to multiple frames of audio, and the latter of the pronunciation states of two adjacent frames of audio is the associated pronunciation state of the former; according to the transition probability between the pronunciation states of two adjacent frames of audio in the state sequence to be identified, the lyrics pronunciation sequence of the audio to be identified is determined from the multiple state sequences to be identified.

[0093] In this embodiment, the server 104 can input the audio features and pronunciation states of each frame of audio into the target language model, and identify the association between each frame of audio through the target language model. The audio to be identified can include multiple frames of audio. For each frame of audio in the audio to be identified, the server 104 can input the audio features of the frame of audio and its corresponding pronunciation state into the target language model, and identify the language of the lyrics corresponding to the audio features and the audio style of the audio to be identified corresponding to the audio features through the target language model. Thus, the server 104 can output the associated pronunciation state corresponding to the frame of audio and its transition probability based on the language of the lyrics corresponding to the audio features, the audio style of the audio to be identified corresponding to the audio features, and the pronunciation state of the frame of audio. The associated pronunciation state can represent the pronunciation state of the next frame of audio corresponding to the frame of audio, and the associated pronunciation state can be determined from the pronunciation states of other frames of audio other than the frame of audio; the transition probability can represent the probability of the pronunciation state of the frame of audio being converted to the associated pronunciation state. Since there can be multiple associated pronunciation states, there can also be multiple transition probabilities, that is, each pronunciation state and its associated pronunciation state can have a corresponding transition probability. Transition probability is an important concept in Markov chains. If a Markov chain is divided into m states, historical data is converted into a sequence consisting of these m states. Starting from any state, after any transition, one of the states 1, 2, ..., m will inevitably appear. The transition between these states is called the transition probability. Specifically, if Figure 5 As shown, Figure 5The server 104 can determine the pronunciation state of each frame of audio through a neural network, and identify the transition probability between the states of different frames of audio through the target language model, such as Figure 5 As shown, sa, sb and sc represent different pronunciation states. The three states can transition to each other or to any other state, and the probability of transition can be determined by the above target language model.

[0094] After the server 104 obtains the pronunciation state corresponding to each frame of audio and the associated pronunciation state corresponding to each frame of audio, it can construct a state sequence to be identified based on the pronunciation state corresponding to each frame of audio and the associated pronunciation state corresponding to each frame of audio, wherein each state sequence to be identified includes the pronunciation state corresponding to multiple frames of audio and the associated pronunciation state corresponding to each frame of audio, and there can be multiple associated pronunciation states corresponding to each frame of audio. The server 103 can determine a sequence from multiple state sequences to be identified based on the transition probability between the pronunciation states of adjacent frames of audio in the state sequence to be identified, as the lyrics pronunciation sequence of the audio to be identified. Specifically, the server 104 can calculate the emission probability (i.e., output probability) of the pronunciation state to which each frame of audio belongs through the above-mentioned DNN, and calculate the transition probability between the pronunciation state of each frame of audio and its associated pronunciation state through the above-mentioned target language model, so that the server 104 can obtain multiple decoding paths based on these probabilities, and the server 104 can select a sequence from the multiple state sequences to be identified formed by these decoding paths as the final recognition result. Among them, since the phonemes corresponding to each frame of audio can correspond to multiple languages ​​and multiple different audio styles, that is, one phoneme can be expanded into multiple phonemes, if all paths are listed to compare the probability values, the amount of calculation is large. Therefore, the server 104 can use the Viterbi decoding algorithm to obtain a final sequence from multiple state sequences to be identified as the lyrics pronunciation sequence of the audio to be identified.

[0095] Through this embodiment, the server 104 can determine an optimal state sequence by calculating the transition probability of adjacent frames in each state sequence in multiple state sequences, so that the server 104 can perform audio recognition based on the optimal state sequence, thereby improving the accuracy of audio recognition.

[0096] In one embodiment, a lyrics pronunciation sequence of audio to be recognized is determined from multiple state sequences to be recognized based on the transition probability between the pronunciation states of adjacent frames of audio in the state sequence to be recognized, including: determining a target state sequence from multiple state sequences to be recognized based on the transition probability between the pronunciation states of two adjacent frames in the state sequence to be recognized; and converting the target state sequence into a phoneme sequence as the lyrics pronunciation sequence of audio to be recognized.

[0097] In this embodiment, the state sequence to be recognized may include pronunciation states corresponding to multiple frames of audio, as well as associated pronunciation states corresponding to each frame of audio. Thus, the server 104 can determine a target state sequence from the multiple state sequences to be recognized based on the transition probabilities between the pronunciation states of adjacent frames in the state sequence to be recognized. The server 104 can determine the target state sequence based on the product of multiple transition probabilities in each state sequence to be recognized. For example, in one embodiment, determining the target state sequence from the multiple state sequences to be recognized based on the transition probabilities between the pronunciation states of adjacent frames of audio in the state sequence to be recognized includes: obtaining the product of the transition probabilities between the pronunciation states of adjacent frames of audio in the state sequence to be recognized; and determining the state sequence to be recognized corresponding to the maximum value of the products from the multiple state sequences to be recognized as the target state sequence. In this embodiment, the server 104 can obtain the product of the transition probabilities between the pronunciation states of adjacent frames of audio in the state sequence to be recognized. There may be multiple state sequences to be recognized, and the server 104 may also calculate multiple products. The server 104 may determine the state sequence to be recognized corresponding to the maximum value of the products of the transition probabilities as the target state sequence.

[0098] After the server 104 calculates the target state sequence, it can convert the target state sequence into a phoneme sequence, thereby using the phoneme sequence as the lyrics pronunciation sequence of the audio to be recognized. The server 104 can convert the target state sequence into a phoneme sequence based on the target language model. For example, in one embodiment, converting the target state sequence into a phoneme sequence as the lyrics pronunciation sequence of the audio to be recognized includes: inputting the target state sequence into the target language model, and using the target language model to identify the associated phonemes and the probabilities of the associated phonemes corresponding to each pronunciation state combination in the target state sequence for each known language and known audio style, thereby obtaining multiple associated phonemes and their probabilities for each pronunciation state combination; the pronunciation state combination includes a pronunciation start stage, a pronunciation middle stage, and a pronunciation end stage; determining multiple associated phoneme sequences based on the associated phonemes corresponding to each pronunciation state combination in the target state sequence; obtaining the product of the probabilities between two adjacent associated phonemes in the associated phoneme sequence; and obtaining the lyrics pronunciation sequence of the audio to be recognized based on the associated phoneme sequence corresponding to the maximum value of the product of the probabilities.

[0099] In this embodiment, the server 104 can obtain a corresponding pronunciation state combination based on the pronunciation states of multiple frames of audio in the target state sequence and their associated pronunciation states. The pronunciation state combination can include a pronunciation start stage, a pronunciation middle stage, and a pronunciation end stage. The server 104 can input the target state sequence obtained above into the target language model and identify the probability of the pronunciation state combination in the target state sequence corresponding to the phonemes of each language and each audio style of the target language model based on the target language model. In order to facilitate the distinction from other phonemes, the identified phonemes can be called associated phonemes. In other words, multiple associated phonemes and the corresponding probabilities of each associated phoneme can be identified based on the target language model. That is, the above pronunciation state combination can form a phoneme. For example, the server 104 forms the phoneme "w" based on the pronunciation state combination of three stages. Then, this phoneme corresponds to multiple recognition results of the target language model, such as the w_chn phoneme of Chinese lyrics, the w_eng phoneme of English lyrics, or the w phoneme of different audio styles. That is, each of the aforementioned pronunciation state combinations may correspond to multiple associated phonemes. The server 104 may, based on the multiple associated phonemes corresponding to each pronunciation state combination in the target state sequence, permutate and combine the associated phonemes corresponding to the multiple pronunciation state combinations to determine multiple associated phoneme sequences. Furthermore, the server 104 may determine the probability of mutual conversion between the associated phonemes in the associated phoneme sequence and obtain the probability product between adjacent associated phonemes in each associated phoneme sequence. Thus, the server 104 may obtain the lyric pronunciation sequence of the audio to be recognized based on the associated phoneme sequence corresponding to the maximum value of the probability product.

[0100] Through the above embodiment, the server 104 can obtain multiple state sequences and multiple phoneme sequences based on the target language model recognition, and determine the optimal sequence based on the product of each transition probability in these sequences, so that the server 104 obtains the lyrics pronunciation sequence of the audio to be recognized based on the optimal sequence, thereby improving the accuracy of audio recognition.

[0101] In one embodiment, Figure 6 As shown, Figure 6: is a flow chart of an audio recognition method in another embodiment. The following steps are included: taking the case where the audio to be recognized is a song containing mixed Chinese and English lyrics as an example, the audio style can be a music genre, including pop music, hip-hop music, heavy metal music, and other music that does not belong to these three categories is classified as other music. The server 104 can first construct a Chinese-English mixed pronunciation dictionary based on genre. Specifically, the server 104 can separate Chinese lyrics and English lyrics from a large number of Chinese-English mixed lyrics, and after deduplication, obtain a word set; the server 104 obtains the phonemes corresponding to the above Chinese lyrics and English lyrics based on the mapping relationship of the international phonetic symbols, and constructs the original mixed dictionary based on these phonemes. The server 104 can add language identifiers based on the language of the lyrics of the phonemes in the dictionary, such as the "_chn" suffix and the "_eng" suffix, to distinguish the differences between the same phonemes in Chinese and English. The server 104 can also add audio style identifiers based on the music genre information of the audio corresponding to the phonemes in the dictionary, such as the "_pop" suffix, the "_hiphop" suffix, the "_metal" suffix, and the "_other" suffix, to indicate the pronunciation differences of the same phoneme in different genres. Thus, the server 104 can construct a genre-based Chinese-English mixed pronunciation dictionary based on the phonemes with the language identifiers and audio style identifiers added. The server 104 can use a large amount of lyrics text based on the Chinese-English mixed pronunciation dictionary to train an N-Gram target language model (i.e., a language model).

[0102] In addition, the server 104 can also extract audio features from the Chinese and English mixed songs input by the user, and the extraction method can be as shown in the above method. The server 104 can input the above audio features into the above DNN model and HMM model to obtain the pronunciation state corresponding to each frame of audio features, the associated pronunciation state corresponding to each frame of audio features and its transition probability. The server 104 can combine the above target language model and determine the transition probability of state transition between adjacent frames in the state sequence combination by considering the matching degree between the pronunciation state of phonemes of different languages ​​and different audio styles and the state of each frame in the state sequence combination, and perform the optimal decoding path search based on these transition probabilities to obtain the optimal target state sequence. The server 104 can convert the optimal target state sequence into a phoneme sequence through the above state sequence-phoneme sequence conversion method, and then convert the phoneme sequence into text through the above phoneme sequence-text sequence conversion method, thereby obtaining the final lyrics recognition result.

[0103] Through this embodiment, the server 104 can use a language model trained based on a mixed pronunciation dictionary and jointly identify lyrics information in the audio based on language and genre, thereby improving the accuracy of audio recognition.

[0104] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0105] Based on the same inventive concept, the embodiments of the present application also provide a language model training device and an audio recognition device for implementing the above-mentioned language model training method and the audio recognition method. The implementation solution provided by the device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations in the one or more audio recognition device embodiments provided below can be referred to the limitations of the language model training method and the audio recognition method above, and will not be repeated here.

[0106] In one embodiment, Figure 7 As shown, a language model training device is provided, including: an audio acquisition module 500, an adding module 502, a construction module 504 and a training module 506, wherein:

[0107] The audio acquisition module 500 is used to acquire multiple sample audios and sample lyrics of each sample audio, wherein the multiple sample audios correspond to multiple audio styles and the sample lyrics corresponding to the sample audios include multiple languages.

[0108] The adding module 502 is configured to add a language identifier and an audio style identifier to the sample phonemes included in the sample lyrics according to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics.

[0109] The construction module 504 is configured to construct a mixed pronunciation dictionary based on the sample phonemes with added labels corresponding to the sample lyrics of the multiple sample audios.

[0110] The training module 506 is used to train the language model to be trained according to the mixed pronunciation dictionary to obtain a target language model. The target language model is used to recognize the lyrics pronunciation sequence of the input audio based on the language and audio style of the input audio.

[0111] In one embodiment, the audio acquisition module 500 is specifically used to obtain audio lyrics corresponding to the sample audio, deduplicate the audio lyrics, and obtain sample lyrics.

[0112] In one embodiment, the adding module 502 is specifically used to obtain the sample phonemes contained in the sample lyrics according to the mapping relationship between the sample lyrics and the standard phonetic symbols; and add a language identifier and an audio style identifier to the sample phonemes according to the language corresponding to the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics.

[0113] In one embodiment, the above-mentioned adding module 502 is specifically used to add a language identifier to the sample phoneme according to the language of the sample lyrics to obtain a first identified sample phoneme; and add an audio style identifier to the first identified sample phoneme according to the audio style of the sample audio corresponding to the first identified sample phoneme to obtain a second identified sample phoneme as the phoneme after adding the identifier.

[0114] In one embodiment, Figure 8 As shown, an audio recognition device is provided, including: a state acquisition module 600, a determination module 602 and a recognition module 604, wherein:

[0115] The state acquisition module 600 is used to acquire the audio features of each frame of audio to be recognized; and acquire the pronunciation state of each frame of audio according to the audio features of each frame of audio.

[0116] Determination module 602 is used to input the audio features and pronunciation status of each frame of audio into the target language model, identify the association between each frame of audio through the target language model and determine the lyrics pronunciation sequence of the audio to be identified based on the association; the target language model is trained according to the language model training method as described above.

[0117] The recognition module 604 is used to recognize the lyrics text corresponding to the audio to be recognized based on the lyrics pronunciation sequence.

[0118] In one embodiment, the above-mentioned state acquisition module 600 is specifically used to input the audio features corresponding to each frame of audio in the audio to be identified into a preset state recognition model, obtain the first probability that the frame of audio is in the pronunciation start state, the second probability that the frame of audio is in the pronunciation middle state, and the third probability that the frame of audio is in the pronunciation end state output by the preset state recognition model; and take the pronunciation state corresponding to the maximum value among the first probability, the second probability and the third probability as the pronunciation state to which the frame of audio belongs.

[0119] In one embodiment, the above-mentioned determination module 602 is specifically used to input the audio features and corresponding pronunciation states of each frame of audio in the audio to be identified into the target language model; identify the language of the lyrics corresponding to the audio features and the audio style of the audio to be identified corresponding to the audio features through the target language model, and determine multiple associated pronunciation states of the frame of audio and the transition probability of each associated pronunciation state based on the identified language, audio style and pronunciation state of the frame of audio; the associated pronunciation state represents the pronunciation state of the next frame of audio corresponding to the frame of audio; the transition probability represents the probability of converting from the pronunciation state of the frame of audio to the associated pronunciation state; construct multiple state sequences to be identified based on the pronunciation state corresponding to each frame of audio and the multiple associated pronunciation states corresponding to each frame of audio; each state sequence to be identified contains pronunciation states corresponding to multiple frames of audio, and the latter of the pronunciation states of two adjacent frames of audio is the associated pronunciation state of the former; based on the transition probability between the pronunciation states of two adjacent frames of audio in the state sequence to be identified, determine the lyrics pronunciation sequence of the audio to be identified from the multiple state sequences to be identified.

[0120] In one embodiment, the above-mentioned determination module 602 is specifically used to determine a target state sequence from multiple state sequences to be identified based on the transition probability between two adjacent frames of pronunciation states in the state sequence to be identified; and convert the target state sequence into a phoneme sequence as the lyrics pronunciation sequence of the audio to be identified.

[0121] In one embodiment, the above-mentioned determination module 602 is specifically used to obtain the product of the transition probabilities between the pronunciation states of adjacent frames of audio in the state sequence to be identified; and determine the state sequence to be identified corresponding to the maximum value of the product from multiple state sequences to be identified as the target state sequence.

[0122] In one embodiment, the above-mentioned determination module 602 is specifically used to input the target state sequence into the target language model, and the target language model identifies the associated phonemes and the probabilities of the associated phonemes corresponding to each known language and known audio style for each pronunciation state combination in the target state sequence, and obtains multiple associated phonemes and their probabilities for each pronunciation state combination; the pronunciation state combination includes the pronunciation starting stage, the pronunciation middle stage and the pronunciation ending stage; based on the arrangement results between the associated phonemes corresponding to the multiple pronunciation state combinations in the target state sequence, multiple associated phoneme sequences are determined; the product of the probabilities between two adjacent associated phonemes in the associated phoneme sequence is obtained; and based on the associated phoneme sequence corresponding to the maximum value of the product of the probabilities, the lyrics pronunciation sequence of the audio to be recognized is obtained.

[0123] Each module in the above-mentioned audio recognition device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0124] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a language model training method and an audio recognition method are implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0125] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0126] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above-mentioned language model training method and audio recognition method when executing the computer program.

[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the language model training method and audio recognition method described above are implemented.

[0128] In one embodiment, a computer program product is provided, including a computer program, which implements the above-mentioned language model training method and audio recognition method when executed by a processor.

[0129] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0130] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0131] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0132] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A language model training method, characterized in that: The method comprises: Acquire multiple sample audios and sample lyrics of each of the sample audios, wherein the multiple sample audios correspond to multiple audio styles and the sample lyrics corresponding to the sample audios include multiple languages; According to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics, adding a language identifier and an audio style identifier to the sample phonemes included in the sample lyrics; Constructing a mixed pronunciation dictionary based on the sample phonemes with added marks corresponding to the sample lyrics of the multiple sample audios; The language model to be trained is trained according to the mixed pronunciation dictionary to obtain a target language model, and the target language model is used to recognize the lyrics pronunciation sequence of the input audio based on the language and audio style of the input audio.

2. The method according to claim 1, characterized in that Obtaining sample lyrics of each sample audio, including: The audio lyrics corresponding to the sample audio are obtained, and the audio lyrics are deduplicated to obtain the sample lyrics.

3. The method according to claim 1, characterized in that The adding a language identifier and an audio style identifier to the sample phonemes included in the sample lyrics according to the language of the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics includes: According to the mapping relationship between the sample lyrics and the standard phonetic symbols, obtaining the sample phonemes contained in the sample lyrics; According to the language corresponding to the sample lyrics and the audio style of the sample audio corresponding to the sample lyrics, a language identifier and an audio style identifier are added to the sample phonemes.

4. An audio recognition method, characterized in that: The method comprises: Obtaining audio features of each frame of audio to be recognized; According to the audio features of each frame of audio, obtaining the pronunciation state of each frame of audio; Inputting the audio features and pronunciation states of each frame of audio into a target language model, identifying the association between each frame of audio through the target language model and determining the lyrics pronunciation sequence of the audio to be identified according to the association; the target language model is trained according to the method described in any one of claims 1 to 3; According to the lyrics pronunciation sequence, the lyrics text corresponding to the audio to be recognized is identified.

5. The method according to claim 4, characterized in that The pronunciation state includes a pronunciation start state, a pronunciation middle state and a pronunciation end state; The obtaining, according to the audio features of each frame of audio, the pronunciation state of each frame of audio, comprises: For each frame of audio in the audio to be recognized, input the audio features corresponding to the frame of audio into a preset state recognition model, and obtain a first probability that the frame of audio is in a pronunciation start state, a second probability that the frame of audio is in a pronunciation middle state, and a third probability that the frame of audio is in a pronunciation end state output by the preset state recognition model; The pronunciation state corresponding to the maximum value among the first probability, the second probability and the third probability is used as the pronunciation state to which the frame of audio belongs.

6. The method according to claim 4, characterized in that The step of inputting the audio features and pronunciation status of each frame of audio into a target language model, identifying the association between each frame of audio through the target language model, and determining the lyrics pronunciation sequence of the audio to be identified according to the association, comprises: For each frame of audio in the audio to be recognized, inputting the audio features of the frame of audio and its corresponding pronunciation state into the target language model; The target language model is used to identify the language of the lyrics corresponding to the audio feature and the audio style of the audio to be identified corresponding to the audio feature, and based on the identified language, audio style and pronunciation state of the frame of audio, multiple associated pronunciation states of the frame of audio and a transition probability of each associated pronunciation state are determined; the associated pronunciation state represents the pronunciation state of the next frame of audio corresponding to the frame of audio; the transition probability represents the probability of transitioning from the pronunciation state of the frame of audio to the associated pronunciation state; According to the pronunciation state corresponding to each frame of audio and the multiple associated pronunciation states corresponding to each frame of audio, multiple state sequences to be identified are constructed; each state sequence to be identified includes the pronunciation states corresponding to multiple frames of audio, and the latter of the pronunciation states of two adjacent frames of audio is the associated pronunciation state of the former; According to the transition probability between the pronunciation states of two adjacent frames of audio in the state sequence to be recognized, the lyrics pronunciation sequence of the audio to be recognized is determined from a plurality of the state sequences to be recognized.

7. The method according to claim 6, characterized in that The method of determining the lyrics pronunciation sequence of the audio to be identified from a plurality of state sequences to be identified according to the transition probabilities between the pronunciation states of adjacent frames of audio in the state sequence to be identified comprises: Determining a target state sequence from a plurality of state sequences to be identified according to a transition probability between two adjacent frames of pronunciation states in the state sequence to be identified; The target state sequence is converted into a phoneme sequence as a lyrics pronunciation sequence of the audio to be recognized.

8. The method according to claim 7, characterized in that The step of determining a target state sequence from a plurality of state sequences to be identified according to the transition probabilities between the pronunciation states of adjacent frames of audio in the state sequence to be identified comprises: Obtaining the product of transition probabilities between pronunciation states of adjacent frames of audio in the state sequence to be identified; A state sequence to be identified corresponding to the maximum value of the product is determined from a plurality of state sequences to be identified as a target state sequence.

9. The method according to claim 7, characterized in that: The step of converting the target state sequence into a phoneme sequence as a lyrics pronunciation sequence of the audio to be recognized includes: The target state sequence is input into the target language model, and the target language model identifies the associated phonemes corresponding to each known language and known audio style for each pronunciation state combination in the target state sequence and the probability of the associated phonemes, thereby obtaining a plurality of associated phonemes and their probabilities for each pronunciation state combination; the pronunciation state combination includes a pronunciation start stage, a pronunciation middle stage, and a pronunciation end stage; Determining multiple associated phoneme sequences according to the arrangement results between the associated phonemes corresponding to the multiple pronunciation state combinations in the target state sequence; Obtaining the product of the probabilities of two adjacent associated phonemes in the associated phoneme sequence; According to the associated phoneme sequence corresponding to the maximum value of the product of the probabilities, the lyrics pronunciation sequence of the audio to be recognized is obtained.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Weighted finite-state converter construction method and device, and speech recognition method and device

    CN112017648A

  • Speech synthesis method, device and equipment and storage medium

    CN112735373A