Audio data processing method, device, electronic device, medium and program product
By identifying phoneme fragments in singing audio, especially the expansion of drag phonemes, the problems of low efficiency and poor accuracy of music score generation in the prior art are solved, and more efficient and accurate music score generation is achieved.
Patent Information
- Application Number
- CN202210953220.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-09
AI Technical Summary
In the prior art, the song scores are inefficient and have poor accuracy, and they usually rely on manual score recording methods.
By obtaining singing audio and associated lyric text, identifying phonemes for each character, including the extended phonemes of the original phoneme and the drag phoneme, predicting the corresponding segments of the phonemes in the audio, and generating an accurate score.
It improves the efficiency and accuracy of music score generation, realizes phoneme-level alignment between lyrics text and singing audio, and generates a more complete and accurate music score.
Smart Images

Figure CN115331654B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud technology, and in particular, to an audio data processing method, apparatus, electronic device, medium, and program product. Background Art
[0002] Currently, based on some actual needs, users usually need to obtain the sheet music of a song. For example, in a singing software, in order to improve the singing experience, the sheet music can be generated for users to view. Among them, the sheet music of a song can be used to record the notes included in the song and the relationship between the notes and the lyrics.
[0003] In existing applications, usually the method of manual score extraction is adopted to generate the sheet music of a song, and using this method to generate the sheet music of a song will bring problems such as low efficiency and poor accuracy of the generated sheet music. Summary of the Invention
[0004] Embodiments of this application provide an audio data processing method, apparatus, electronic device, medium, and program product, which can improve the generation efficiency and accuracy of the sheet music of singing audio.
[0005] On the one hand, embodiments of this application provide an audio data processing method, and the method includes:
[0006] Obtain singing audio, and obtain the lyric text associated with the singing audio;
[0007] Obtain the phonemes of the pronunciation of each character in the lyric text respectively; the lyric text includes a target character, and the phonemes of the pronunciation of the target character include at least one original phoneme of the target character and an extended phoneme of a trailing phoneme, and the trailing phoneme refers to the last original phoneme in the original phonemes of the target character, and the extended phoneme of the trailing phoneme is used to represent the same phoneme after the pronunciation of the trailing phoneme changes;
[0008] Predict the audio segments corresponding to each phoneme of the pronunciation of the characters in the lyric text in the singing audio respectively; any one original phoneme corresponds to an audio segment, and any one extended phoneme corresponds to an audio segment or does not correspond to an audio segment;
[0009] Determine the target audio segments corresponding to each original phoneme of the characters in the lyric text according to the audio segments corresponding to each phoneme respectively, and generate the sheet music of the singing audio according to the target audio segments corresponding to each original phoneme and the lyric text.
[0010] On the one hand, embodiments of this application provide an audio data processing apparatus, and the apparatus includes:
[0011] An obtaining module, configured to obtain singing audio and obtain the lyric text associated with the singing audio;
[0012] An acquisition module is further configured to acquire the phonemes of the pronunciation of each character in the lyric text respectively; the lyric text includes target characters, and the phonemes of the pronunciation of the target characters include at least one original phoneme of the target characters and extended phonemes of the sustained phonemes. The sustained phoneme refers to the last original phoneme in the original phonemes of the target characters, and the extended phoneme of the sustained phoneme is used to represent the same phoneme after the pronunciation of the sustained phoneme changes.
[0013] A processing module is configured to predict the audio segments corresponding to the phonemes of the pronunciation of each character in the lyric text in the singing audio respectively; any one original phoneme corresponds to an audio segment, and any one extended phoneme corresponds to an audio segment or does not correspond to an audio segment.
[0014] The processing module is further configured to determine the target audio segments corresponding to the original phonemes of each character in the lyric text according to the audio segments corresponding to the phonemes respectively, and generate the musical score of the singing audio according to the target audio segments corresponding to the original phonemes and the lyric text.
[0015] On the one hand, an embodiment of the present application provides an electronic device, which includes a processor and a memory. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute some or all of the steps in the above method.
[0016] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute some or all of the steps in the above method.
[0017] Correspondingly, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes program instructions, and the program instructions are stored in a computer-readable storage medium. The processor of the computer device reads the program instructions from the computer-readable storage medium, and the processor executes the program instructions, so that the computer device executes some or all of the steps in the above method.
[0018] In the embodiments of the present application, a singing audio can be obtained, a lyric text associated with the singing audio can be obtained, and phonemes of the pronunciation of each character in the lyric text can be obtained respectively. The phonemes of the pronunciation of the character can include sustained phonemes and extended phonemes of the same phoneme after the pronunciation change of the sustained phoneme. Audio segments corresponding to each phoneme of the pronunciation of the characters in the lyric text can be predicted respectively, target audio segments corresponding to each original phoneme of the characters in the lyric text can be determined according to the audio segments corresponding to each phoneme respectively, and a musical score of the singing audio can be generated according to the target audio segments corresponding to each original phoneme and the lyric text. Through the above method, when using a sustained phoneme to represent the pronunciation process of a phoneme, if the pronunciation of the phoneme changes, the extended phoneme can be used to further represent the pronunciation process of the same phoneme after the pronunciation change, so that a more complete pronunciation process of a phoneme can be represented by the sustained phoneme and the extended phoneme, and the target audio segments corresponding to each original phoneme of the pronunciation of each character can be determined more accurately, so as to achieve an accurate alignment at the phoneme level between the lyric text and the singing audio, and generate a musical score based on the alignment result, which helps to improve the efficiency and accuracy of musical score generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 FIG. [ID] is a schematic diagram of an application architecture provided by an embodiment of the present application;
[0021] Figure 2 FIG. [ID] is a schematic flowchart of a method for processing audio data provided by an embodiment of the present application;
[0022] Figure 3 FIG. [ID] is a schematic diagram of a representation scenario of a phoneme provided by an embodiment of the present application;
[0023] Figure 4 FIG. [ID] is a schematic diagram of a scenario for determining an audio segment corresponding to a phoneme provided by an embodiment of the present application;
[0024] Figure 5 FIG. [ID] is a schematic diagram of a musical score provided by an embodiment of the present application;
[0025] Figure 6 FIG. [ID] is a schematic flowchart of a method for processing audio data provided by an embodiment of the present application;
[0026] Figure 7a FIG. [ID] is a schematic diagram of a scenario for determining an audio segment corresponding to a phoneme provided by an embodiment of the present application;
[0027] Figure 7b A schematic diagram of a scenario for determining an audio segment corresponding to a phoneme provided by an embodiment of the present application;
[0028] Figure 8 A schematic diagram of a scenario for determining a phoneme associated with a phoneme provided by an embodiment of the present application;
[0029] Figure 9 A schematic framework diagram of audio data processing provided by an embodiment of the present application;
[0030] Figure 10a A schematic diagram of an audio processing result provided by an embodiment of the present application;
[0031] Figure 10b A schematic diagram of a vibrato interval provided by an embodiment of the present application;
[0032] Figure 11 A schematic structural diagram of an audio data processing device provided by an embodiment of the present application;
[0033] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0035] The audio data processing method proposed in the embodiments of the present application is implemented on an electronic device, which can be a server or a terminal. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto.
[0036] Next, relevant introductions will be made to the technical terms involved in the technical fields to which the solutions of the embodiments of the present application may be applied:
[0037] I. Artificial Intelligence (AI):
[0038] The embodiments of the present application relate to the field of artificial intelligence technology. Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a manner similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0039] The embodiments of the present application may specifically relate to the field of machine learning (ML) in artificial intelligence. Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. For example, machine learning technology can be used in the technical solution of the present application to predict the audio segment corresponding to each phoneme.
[0040] II. Cloud Technology:
[0041] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The back-end services of technical network systems require a large amount of computing and storage resources, such as video websites, picture-based websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0042] In some embodiments, the audio data processing method provided in the embodiments of the present application can be combined with cloud technology. For example, predicting the audio clips corresponding to each audio, or generating the music score of the singing audio, etc., can be implemented through the artificial intelligence cloud service in the field of cloud technology to realize the intelligent generation of the music score. Among them, the so-called artificial intelligence cloud service is generally also referred to as AIaaS (AI as a Service, Chinese for "AI as a Service"). This is a mainstream service mode of an artificial intelligence platform. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI theme mall: all developers can access one or more artificial intelligence services provided by the platform through the API interface. Some senior developers can also use the AI framework and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services.
[0043] 3. Phone:
[0044] Phonemes are the smallest units of speech divided according to the natural properties of speech. According to the analysis of pronunciation actions, one action constitutes a phoneme. The pronunciation of a character can contain multiple phonemes. For characters in different languages, the expression and pronunciation of phonemes may be different. For example, taking Chinese as an example, the phonemes of Chinese pronunciation can be split based on the pinyin of a character pronunciation. A pinyin can generally be split into multiple phonemes, such as the pinyin of the character "吧" is "ba", and the phonemes of the character pronunciation can be split into consonants "b" and vowels "a". The specific composition of phonemes is not limited and can be set by relevant business personnel according to actual scenarios. The embodiment of the present application can pre-construct a dictionary to record the phonemes of different character pronunciations. For example, a Chinese phoneme dictionary can be constructed to record the phonemes of each Chinese character pronunciation, and the Chinese phoneme dictionary can specifically record the mapping relationship between different Chinese characters or Chinese characters' pinyin to phonemes, or an English phoneme dictionary can be constructed to record the phonemes of each word pronunciation, and the English phoneme dictionary can specifically record the mapping relationship between different words or word phonetic symbols to phonemes.
[0045] 4. Notes:
[0046] Notes are symbols used to record the progression of different types of sounds. Notes can be used to describe the melody of a song, that is, a musical score can be composed of different note arrangements. A note includes the pronunciation duration (i.e., the length of a sound) and the pitch (i.e., the height of a sound). Different pronunciation durations and different pitches correspond to different forms of note representations. For example, different pitches can correspond to symbols of different shapes (such as do, re, mi... in numbered musical notation), and different pronunciation durations can correspond to the lengths of the intervals between symbols; or different pronunciation durations correspond to symbols of different shapes, and different pitches correspond to different positions of the symbols (such as on which pitch height line the symbol is, etc.) and / or different pitch identifiers (such as C4 key, etc.). This form of representation can be different according to different actual application scenarios and is not specifically limited. It can be set by relevant business personnel according to the actual scenario.
[0047] In some embodiments, please refer to Figure 1 , Figure 1 , which is a schematic diagram of an application architecture provided by an embodiment of the present application. The audio data processing method proposed in the present application can be executed through this application architecture. As Figure 1 shown, the electronic device can obtain the singing audio and its associated lyric text, and respectively obtain the phonemes of each character pronunciation in the lyric text. Each character includes a target character with original phonemes and extended phonemes and other characters with only original phonemes. Predict the phoneme segments corresponding to each phoneme of each character pronunciation in the singing audio to achieve the alignment of the lyric text and the singing audio. Each original phoneme of the pronunciation of the target character and other characters can correspond to an audio segment. The extended phonemes of the pronunciation of the target character may or may not correspond to an audio segment. According to the audio segments corresponding to each phoneme respectively, the target audio segment corresponding to each original phoneme can be determined, and a musical score of the singing audio can be generated based on the target audio segment corresponding to each original phoneme and the lyric text. For example, the musical score can be jointly generated according to the relationship between each original phoneme and the characters in the lyric text and the corresponding target audio segment.
[0048] It can be understood that Figure 1 this is only an exemplary representation of the possible application architecture of the technical solution of the present application and does not limit the specific architecture of the technical solution of the present application. That is, the technical solution of the present application can also provide other forms of application architectures.
[0049] Optionally, in some embodiments, the electronic device may execute the audio data processing method according to actual service requirements to improve the efficiency and accuracy of score generation. The technical solution of the present application can be applied to any score generation scenario of singing audio associated with lyric texts, that is, the electronic device needs to obtain the lyric text of the singing audio, and determine the target audio segment of each original phoneme in the singing audio in combination with the phonemes of each character pronunciation in the lyric text, so as to align the singing audio with the lyric text, and can perform score transcription operations according to the target audio segment of each original phoneme and each character in the lyric text to generate the corresponding score. The score not only includes the notes that can describe the melody of the singing audio, but also includes the alignment relationship between the notes and the characters.
[0050] Optionally, the data involved in the present application, such as the phonemes of each character pronunciation, the score of the singing audio, etc., can be stored in a database, or can be stored in a blockchain, such as stored through a blockchain distributed system, which is not limited in the present application.
[0051] It should be noted that in the specific implementation manner of the present application, if obtaining the singing audio or lyric text involves relevant data for collecting user information, such as collecting the singing audio sung by the user or the lyric text made, then when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0052] It can be understood that the above scenarios are only examples and do not constitute a limitation on the application scenarios of the technical solutions provided by the embodiments of the present application. The technical solutions of the present application can also be applied to other scenarios. For example, as is known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0053] Based on the above description, an audio data processing method is proposed in the embodiments of the present application, and this method can be executed by the above-mentioned electronic device. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an audio data processing method provided by an embodiment of the present application. As Figure 2 shown, the process of the audio data processing method in the embodiments of the present application may include the following:
[0054] S201. Obtain the singing audio and obtain the lyric text associated with the singing audio.
[0055] In some embodiments, the obtained singing audio is an audio with associated lyric text, and each character in the lyric text is sung based on a certain melody in the singing audio. The singing audio is a dry audio, that is, an audio that only contains the sound track for the lyric text. The singing audio can be the sound track audio of the target object, such as a vocal track audio.
[0056] In some embodiments, the electronic device can directly obtain the singing audio, or can perform sound source separation on the initial singing audio mixed with human voice and accompaniment to obtain the vocal track audio as the singing audio. The singing audio or the initial singing audio can be obtained from an audio database, downloaded from a website, or uploaded by the target object (such as a user). There is no limitation here. In addition, the initial singing audio can be a song published online, or a song included in a music listening software, or can also be a song sung in real time, such as a song recorded by the target object in a singing software. There is no limitation on the type of the initial singing audio here.
[0057] In some embodiments, the lyric text associated with the singing audio can be directly obtained, such as obtained from an audio database, downloaded from a website, or uploaded by the target object. The lyric text can also be the text obtained by performing speech recognition on the singing audio, such as recognized by using a singing-based automatic speech recognition (ASR) technology. ASR can recognize the text in a certain piece of speech through a computer. The lyric text can only contain multiple lyric characters without other marking information, such as time stamps (such as the start and end times of a lyric sentence in the singing audio) or punctuation marks, etc. The lyric text is a text composed of any language, such as a text composed of Chinese language, or a text composed of English language, or a text composed of mixed languages (such as a text containing both Chinese and English). There is no limitation here. For singing audio and lyric text in different languages, the process and principle of generating a musical score are the same. For the convenience of description, the subsequent description will take the musical score generation scenario applied to Chinese singing audio as an example, that is, unless otherwise specifically mentioned later, the lyric text mentioned is a text composed of Chinese.
[0058] S202. Respectively obtain the phonemes of the pronunciation of each character in the lyric text.
[0059] In some embodiments, the electronic device may obtain the pinyin of the pronunciation of each character in the lyric text, and obtain the phonemes of the pronunciation of each character based on the pinyin. For example, a phoneme dictionary (such as a Chinese phoneme dictionary) that matches the language type of the character may be obtained, and the phonemes of the pronunciation of each character may be obtained from the Chinese phoneme dictionary according to the pinyin of the pronunciation of each character. Different phoneme dictionaries may be constructed according to the phoneme composition of the pronunciations of characters in different languages. Among them, a character is a language unit that constitutes a sentence. For example, a character may represent a Chinese character, or a character may represent an English word.
[0060] In some embodiments, the Chinese phoneme dictionary records the original phonemes of the pinyin pronunciation of each character. The original phonemes can be obtained by disassembling based on the pinyin. For example, if the pinyin of a character is "ba", the original phonemes determined based on the pronunciation action of the pinyin may include "b" and "a". Further, due to the particularity of the singing scenario, the phonemes of a specified pinyin pronunciation may also be extended to determine the trailing phoneme in the original phonemes of the corresponding character pronunciation, and the extended phonemes of the trailing phoneme.
[0061] In some embodiments, the last original phoneme in the original phonemes of a character pronunciation is the last pronunciation action of the character pronunciation, that is, the last original phoneme can represent the last sound in the character pronunciation process. Therefore, the trailing phoneme can be determined as the last original phoneme, and the extended phonemes of the trailing phoneme also represent the last sound in the character pronunciation process, and the extended phonemes can be expressed as the same phoneme after the pronunciation of the trailing phoneme changes. It can be understood that in the singing scenario, there may be a situation where the pronunciation of a character is dragged very long. For example, the last word in a lyric has a pronunciation duration of more than 3 seconds, and the change in the trailing duration is mainly reflected in the vowel, that is, the vowel duration in a character is relatively long, while the consonant duration is similar to that in the normal speaking scenario. That is, at this time, it can be understood that the duration of the last sound of a character pronunciation is dragged very long, and there may be a change in melody during the duration of the last sound. For example, the last sound suddenly rises from C major to G major during the duration. At this time, the moment of pitch increase of this sound can be regarded as a change in pronunciation.
[0062] Therefore, during the pronunciation process of a character, different pronunciation actions can be represented by primitive phonemes. When the pronunciation action represented by the last primitive phoneme is prolonged, if the pronunciation changes during this prolongation process, the continuous pronunciation process after the change of this pronunciation action can be represented by an extended phoneme, that is, the prolonged phoneme and the extended phoneme are the same phoneme. The Chinese phoneme dictionary containing extended phonemes can be regarded as a dictionary extended from a dictionary containing only primitive phonemes. Since the extended phoneme is the last primitive phoneme in the primitive phonemes, and the last primitive phoneme is usually a vowel phoneme, the Chinese phoneme dictionary can also be called a vowel extension dictionary.
[0063] In some embodiments, the phoneme composition of the pinyin pronunciation in the Chinese phoneme dictionary can be set by relevant business personnel according to empirical values. For example, there may be extended phonemes only in the phonemes of some pinyin pronunciations. In addition, the extended phoneme of a prolonged phoneme can be one or more, and each extended phoneme represents the same phoneme after the pronunciation of the previous phoneme changes. For example, as Figure 3 shown Figure 3 is a schematic diagram of the representation scenario of a phoneme provided by an embodiment of the present application; among them, the phonemes of the character "ba" have extended phonemes, and the prolonged phoneme is represented as "a1", and the extended phonemes are successively represented as "a2" and "a3". Both the prolonged phoneme and the extended phoneme represent the same phoneme "a". During the pronunciation process of "ba", for the pronunciation process of the pronunciation action of "a", it is first represented by "a1". If there is a change in melody when the pronunciation is prolonged, such as a sudden rise in melody, "a2" can represent the process after the pronunciation of "a1" changes, and if there is another change in melody when the pronunciation is prolonged, such as a sudden rise in melody again, "a3" can represent the process after the pronunciation of "a2" changes. The number of extended phonemes in the phonemes of a character's pronunciation can be specifically set by relevant business personnel, and the number of extended phonemes of different characters' pronunciations can be different.
[0064] Therefore, the embodiments of the present application can represent different pronunciation processes of characters through such extended phonemes, and are more suitable for the phoneme alignment task in the singing scenario. The phonemes representing a pronunciation process must include all the primitive phonemes of a character's pronunciation and may include the extended phonemes (i.e., spare phonemes) of a character's pronunciation, so as to achieve the flexibility of phoneme representation.
[0065] In some embodiments, the lyric text includes target characters with extended phonemes and other characters without extended phonemes. The phoneme composition for the pronunciation of a target character can be expressed as: original phoneme 1, original phoneme 2,....., the last original phoneme, extended phoneme 1, extended phoneme 2,...., the last extended phoneme. Different pronunciation processes can be characterized according to different combinations of the foregoing original phonemes and extended phonemes. The phoneme composition for the pronunciation of other characters can be expressed as: original phoneme 1, original phoneme 2,....., the last original phoneme. The pronunciation process can be characterized according to the foregoing original phonemes.
[0066] For example, if the character "ba" is a target character, the phonemes for its pronunciation may include: "b", "a1", "a2", "a3", where "a1" is the last original phoneme, and "a2" and "a3" are extended phonemes, both representing the same phoneme "a". In addition, for different pronunciation processes of "ba", there are multiple phoneme representation entries that can be obtained, such as "b / a1", "b / a1 / a2", "b / a1 / a2 / a3". Another example is that if the character "ba" is an other character, the original phonemes for its pronunciation may include: "b", "a1". In addition, for the pronunciation process of "ba", the phoneme representation entry that can be obtained is, for example, "b / a1".
[0067] It should be noted that for the characters in the lyric text, there may be the same characters. However, based on the order between the characters, they can be regarded as different characters in the lyric text, but the phoneme compositions for the pronunciation of the same characters are the same. For the phonemes for the pronunciation of any two characters, there may be the same phonemes. However, since the characters to which they belong are regarded as different characters, they can be regarded as different phonemes, but the pronunciation actions they represent are the same. For example, for the phoneme "a", in a lyric text, the u-th character is "ba" and the phonemes for its pronunciation include the phoneme "a", and the v-th character is also "ba" and the phonemes for its pronunciation also include the phoneme "a", where u ≠ v. Then the u-th character and the v-th character are different characters, and the phoneme "a" included in the u-th character and the phoneme "a" included in the v-th character are different phonemes.
[0068] S203. Predict the audio segments corresponding to each phoneme of the pronunciation of the characters in the lyric text in the singing audio.
[0069] In some embodiments, the electronic device may predict the audio segments corresponding to each phoneme of the pronunciation of each character in the singing audio. Since the representation of each character pronunciation process includes the original phonemes of the pronunciation, any one of the original phonemes in the phonemes of each character pronunciation in the lyric text corresponds to an audio segment, and any one extended phoneme may or may not correspond to an audio segment. This audio segment may characterize the pronunciation process for the phoneme in the singing audio, thereby enabling phoneme-level alignment. Alignment, also known as forced alignment, in the field of speech recognition, the forced alignment technology can output the start and end times of each word (or phoneme) when inputting a segment of speech and the corresponding text.
[0070] It can be understood that in a singing audio, for two identical characters (denoted as character A and character B), they have the same extended phonemes, but the pronunciation processes corresponding to character A and character B in the singing audio are different. Therefore, there may be a situation where the extended phoneme of character A's pronunciation corresponds to an audio segment, while the extended phoneme of character B's pronunciation does not correspond to an audio segment.
[0071] In some embodiments, the electronic device may divide the singing audio into multiple audio frames, represent each audio frame as a corresponding audio sequence respectively, and predict the audio segments corresponding to each phoneme in the singing audio based on this audio sequence. Specifically, it may be to determine the phoneme associated with each audio frame based on the audio sequence of each audio frame respectively, and determine the audio segment corresponding to each phoneme according to the phoneme associated with each audio frame.
[0072] In some embodiments, an audio sequence may be the spectral feature corresponding to an audio frame. For example, the audio sequence may be the MFCC (Mel-scale Frequency Cepstral Coefficients) of the audio frame, or it may also be an acoustic feature. For example, the audio sequence may be the FBK (filter bank) acoustic feature of the audio frame, or it may also be a combined feature. For example, it may be a concatenated sequence of MFCC and FBK, or it may be a concatenated sequence of MFCC and the pitch corresponding to the audio frame, etc. The specific composition of the audio sequence is not limited here. The length of the audio frame is usually less than the shortest pronunciation duration of a phoneme, and the length of the audio frame can be set by relevant business personnel according to the actual scenario.
[0073] In some embodiments, determining the phoneme associated with each audio frame based on the audio sequence may be obtained by predicting based on the audio sequence through a phoneme model. The phoneme model may be a model with any structure. For example, it may include an audio processing layer and a phoneme prediction layer or only include a phoneme prediction layer. The division of audio frames and the extraction of the audio sequence may be performed by the audio processing layer or may also be performed by an electronic device. The phoneme prediction layer may be constructed based on a DNN (Dynamic Neural Network)-HMM (Hidden Markov Model) model, a GMM (Gaussian Mixture Model)-HMM model, etc. For the specific method of predicting the phoneme associated with each audio frame through the phoneme model, reference may be made to the relevant descriptions in the following embodiments.
[0074] In some embodiments, determining the audio segment corresponding to each phoneme according to the phoneme associated with each audio frame may specifically be that if the phonemes associated with M consecutive audio frames among multiple audio frames are all target phonemes, then the M consecutive audio frames are spliced, and the spliced audio frames are determined as the audio segment corresponding to the target phoneme, that is, the audio frames corresponding to the same phoneme are spliced to obtain the audio segment corresponding to the phoneme. Among them, an audio frame may be associated with one phoneme among the phonemes of the pronunciation of the characters in the lyric text or be associated with a default phoneme. The default phoneme may be a non-phoneme, that is, a phoneme indicating no pronunciation action. The audio frame associated with the default phoneme can be understood as an audio frame without pronunciation. The consecutive audio frames associated with the default phoneme may be determined as the audio segment corresponding to the default phoneme, that is, the pause segment without pronunciation in the singing audio.
[0075] In some embodiments, when predicting the phoneme associated with each audio frame, it may occur that an audio frame is associated with a default phoneme, but the audio frames on the left and right of this audio frame are both associated with the same phoneme, that is, the same phoneme corresponding to the pronunciation of the same character. For example, due to prediction errors, since the phonemes associated with the audio frames corresponding to the complete pronunciation process of a target phoneme will all be the target phoneme, each audio frame associated with a phoneme can be detected. If there is an audio frame associated with a default phoneme and the audio frames on the left and right of this audio frame are associated with the same phoneme, then this audio frame can be corrected so that this audio frame is associated with the same phoneme as the audio frames on the left and right, thereby consecutive audio frames associated with one phoneme can be determined, and the consecutive audio frames are spliced to obtain the corresponding audio segment, and the audio segment corresponding to the default phoneme is determined based on the remaining audio frames associated with the default phoneme.
[0076] For example, as Figure 4 shown, Figure 4A schematic diagram of a scenario for determining an audio segment corresponding to a phoneme provided by an embodiment of the present application; among them, the phonemes of the character pronunciation in the lyric text include phoneme 1, phoneme 2, and phoneme 3. It is predicted that audio frames 1-4 are associated with phoneme 1, audio frames 5-7 are associated with phoneme 2, audio frames 9-10 are associated with phoneme 3, and audio frames 12-14 are associated with phoneme 3. Therefore, the audio frames 1-4 are spliced to obtain audio segment 1 corresponding to phoneme 1, the audio frames 5-7 are spliced to obtain audio segment 2 corresponding to phoneme 2, audio frame 8 is the audio segment 3 corresponding to the default phoneme, and since both audio frames 9-10 and audio frames 12-14 are associated with phoneme 3, the phoneme associated with audio frame 11 is modified from the default phoneme to phoneme 3, and the audio frames 9-14 are spliced to obtain audio segment 4 corresponding to phoneme 3.
[0077] S204. Determine a target audio segment corresponding to each original phoneme of the characters in the lyric text according to the audio segments corresponding to each phoneme respectively, and generate a musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text.
[0078] In some embodiments, the target audio segment corresponding to each original phoneme of the character pronunciation in the lyric text can be determined according to the audio segments corresponding to each phoneme respectively, and the processes and principles of the target audio segments corresponding to each original phoneme are the same. Any one of each original phoneme is represented as a target original phoneme. To determine the target audio segment corresponding to the target original phoneme, specifically, if the target original phoneme does not have an extended phoneme, the audio segment corresponding to the target original phoneme is determined as the target audio segment corresponding to the target original phoneme; if the target original element has an extended phoneme and the extended phoneme of the target original phoneme corresponds to an audio segment, the audio segment corresponding to the target original phoneme and the audio segment corresponding to the extended phoneme of the target original phoneme are spliced, and the spliced audio segment is determined as the target audio segment corresponding to the target original phoneme. The audio duration indicated by the target audio segment is the pronunciation duration of the corresponding original phoneme, that is, the time stamp of the original phoneme in the singing audio, and this time stamp includes the start time and the end time of the original phoneme in the singing audio.
[0079] In some embodiments, the electronic device can obtain the pitch curve of the target audio segment corresponding to each original phoneme, and generate a corresponding musical score based on this pitch curve. It can be to generate a complete pitch curve corresponding to the singing audio, and intercept the curve segment corresponding to the start and end points from the complete pitch curve based on the start and end points indicated by the target audio segment in the singing audio as the pitch curve corresponding to each original phoneme, or it can be to generate a pitch curve for the target audio segment corresponding to each original phoneme respectively.
[0080] In some embodiments, a musical score is composed of musical notes. Therefore, generating a corresponding musical score based on a pitch curve can be to determine the musical note associated with each original phoneme according to the pitch curve corresponding to each original phoneme, and generate a musical score according to the musical note associated with each original phoneme and the lyric text. Among them, the pronunciation duration and pitch included in the musical note associated with the original phoneme can be determined according to the pitch curve corresponding to the original phoneme. For example, the duration indicated by the pitch curve corresponding to the original phoneme can be determined as the pronunciation duration of the associated musical note, which is equivalent to determining the pronunciation duration of the original phoneme as the pronunciation duration of the associated musical note. For another example, the average pitch indicated by the pitch curve can be determined as the pitch of the associated musical note. At this time, one phoneme is associated with one musical note. Or, at least one musical note associated with the original phoneme can also be determined in combination with the change in the curve shape of the pitch curve corresponding to the original phoneme. The specific method can refer to the relevant description in the following embodiments.
[0081] In some embodiments, generating a musical score according to the musical note associated with each original phoneme and the lyric text can specifically be to determine the character to which each original phoneme belongs in the lyric text, and align and display the musical note (or multiple associated musical notes) associated with each original phoneme with the character to which each original phoneme belongs, so as to obtain the musical score corresponding to the singing audio. The musical note associated with the original phoneme of the character pronunciation describes the singing method of this character. The musical score at this time is a vocal score, which represents the association relationship between the singing audio and the lyric text, specifically represents the corresponding relationship between different positions in the singing audio and the characters in the lyric text, and represents the singing melody for the lyric text in the singing audio.
[0082] In some embodiments, a musical score generation rule can be obtained to determine the display method of phonemes and characters. For example, the manifestation form of the musical note (such as a symbol, symbol position, or pitch identifier, etc.), the arrangement rule of the musical notes, the alignment rule between the phoneme and the character, etc. can be determined from the musical score generation rule according to the pronunciation duration and pitch. The specific musical score generation rule can be set by relevant business personnel, and the musical score generation rules in different scenarios can be different, the manifestation forms of the musical notes are different, and the specific types of the musical scores are different.
[0083] For example, as Figure 5 shown, Figure 5 is a schematic diagram of a musical score provided by an embodiment of the present application; if the musical score is a numbered musical notation, taking the character "ba" as an example, it includes two original phonemes "b" and "a". The musical note associated with the original phoneme "b" determines the corresponding symbol as 1 (i.e., pronounced "si") based on the pitch mapping relationship, and determines the interval length from the next musical note according to the pronunciation duration mapping relationship. The musical note associated with the original phoneme "a" determines the corresponding symbol as 1 (i.e., pronounced "la") based on the pitch mapping relationship, and determines the interval length from the next musical note according to the pronunciation duration mapping relationship. Therefore, aligning and displaying the musical note associated with the original phoneme with this character, the generated musical score can be asFigure 5 as shown in (1) of Figure 5 ; or, if the music score is a singing score for guiding a singer in a singing software, the symbol of each note is a horizontal line. The length of the horizontal line represents the corresponding pronunciation duration, and the height of the horizontal line represents the corresponding pitch. Taking the character "ba" as an example, the lengths and heights of the horizontal lines corresponding to the notes respectively associated with the original phonemes "b" and "a" are determined (such as on which pitch height line), and the foregoing notes are displayed based on the length and height of the horizontal lines, and the character "ba" is aligned with the displayed notes, and what can be generated is as shown in Figure 5 (2) of
[0084] The music score shown is only an example. According to the different rules for generating music scores in different scenarios, there can be various types of music scores, which are not limited herein.
[0085] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of an audio data processing method provided by an embodiment of the present application, and this method can be executed by the electronic device mentioned above. As shown in Figure 6 , the process of the audio data processing method in the embodiment of the present application can include the following:
[0086] S601. Obtain a singing audio and obtain the lyric text associated with the singing audio. Among them, the specific implementation manner of step S601 can refer to the relevant steps of the foregoing embodiment, and will not be elaborated herein.
[0087] S602. Obtain the phonemes of the pronunciation of each character in the lyric text respectively.
[0088] In some embodiments, the electronic device can obtain a Chinese phoneme dictionary and obtain the pinyin of each character, and obtain the phonemes for pronouncing each character in the lyrics text from the Chinese phoneme dictionary according to the pinyin of each character. The lyrics text includes a target character and other characters, and the phonemes for pronouncing the target character include at least one original phoneme of the target character and an extended phoneme of a dragged phoneme, wherein the dragged phoneme refers to the last original phoneme among the original phonemes of the target character, and the extended phoneme of the dragged phoneme is used to characterize the same phoneme after the pronunciation of the dragged phoneme changes. The phonemes for pronouncing other characters include at least one original phoneme of other characters. For a detailed description of the Chinese phoneme dictionary, original phonemes, dragged phonemes, and extended phonemes, please refer to the relevant description of the above embodiments.
[0089] In some embodiments, the Chinese phoneme dictionary may set different numbers of extended phonemes for all characters, or may set different numbers of extended phonemes for some characters. For example, relevant business personnel may count the frequency of dragging sounds of each character in different singing audios, and set a corresponding number of extended phonemes for the pinyin of high-frequency characters.
[0090] In addition, since characters may be polyphonic, the pinyin of each character can be determined in combination with the contextual meaning of the character in the lyrics text. For example, the pinyin of the character "单" can be "dan" or "shan". If the context of the character in the lyrics text is "孤单", the pinyin of the character is determined to be "dan" based on the contextual meaning. For example, the lyrics text can be segmented and the target feature of each character can be obtained. The target feature can include the encoding feature of the character, the part of speech of the segmented word to which the character belongs, and the position of the segmented word in the lyrics text. The upper and lower features of each character are obtained through a bidirectional long short-term memory model, and the target feature and the upper and lower features of each character are spliced and input into the pinyin prediction model to obtain the pinyin of each character. The pinyin prediction model can be trained by the sample features and sample pinyin labels of each sample character in the sample sentence.
[0091] In some embodiments, the tonal characters involved in the embodiments of the present application can all be represented as toneless characters, that is, each Chinese character usually has a tone, and the tone information can be removed here. That is, the pinyin of each character can be a toneless pinyin. In the singing scene, since each character does not have a fixed tone, when obtaining the pinyin of each character, the pinyin after removing the tone information can be obtained, and the Chinese phoneme dictionary can also record the correspondence between the toneless pinyin and phonemes of different characters.
[0092] For example, for the character "小", the pronunciation is the third tone, so the pinyin is usually expressed as "xiǎo", and the phoneme entry recorded for this pinyin in a conventional phoneme dictionary with tones may be "xiǎo:x / iao3", but in an embodiment of the present application, the pinyin of the character "小" can be removed with tones and expressed as "xiao", and the phoneme entry recorded for this character in the Chinese phoneme dictionary may be xiao:x / iao..... and so on. In addition, the recording form of the phoneme dictionary can be recorded in the form of phoneme entries, such as xiao:x / iao1, x / iao1 / iao2, x / iao1 / iao2 / iao3, indicating that there are original phoneme entries "x / iao1" and extended phoneme entries "x / iao1 / iao2", "x / iao1 / iao2 / iao3", and "iao1" is the original phoneme, "iao2" and "iao3" are extended phonemes; or it can be recorded in the form of phoneme arrangement, such as xiao:x, iao1, iao2, iao3, indicating that "iao1" is the original phoneme, "iao2" and "iao3" are extended phonemes. There can be many specific recording forms in the phoneme dictionary, which are not limited here.
[0093] S603, dividing the singing audio into a plurality of audio frames, and determining the phonemes associated with each audio frame based on the character order between each character in the lyrics text and the pronunciation order corresponding to the phonemes of each character pronunciation.
[0094] The specific method of dividing the singing audio into multiple audio frames can refer to the relevant description of the above embodiment.
[0095] In some embodiments, each audio frame can be represented as a corresponding audio sequence, and the specific description of the audio sequence can refer to the relevant description of the above embodiment. The phonemes associated with each audio frame are predicted in combination with the audio sequence, character order and pronunciation order of the phonemes of each audio frame. The determination process and principle of the phonemes associated with each audio frame are the same, and the determination of the phonemes associated with each audio frame needs to be combined with the associated phoneme results of the previous audio frame. Take an audio frame as an example for illustration. Suppose that each audio frame includes the i-th audio frame, i is an integer and i is greater than 1, and the character to which the phoneme associated with the i-th audio frame belongs is represented as a reference character, which can be specifically determined based on the character order, the pronunciation order of the phonemes, and the phonemes associated with the i-1-th audio frame. The candidate phoneme set associated with the i-th audio frame includes at least one candidate phoneme, and the association probability between each candidate phoneme in the candidate phoneme set and the i-th audio frame is predicted based on the audio sequence of the i-th audio frame, so as to determine the phoneme associated with the i-th audio frame based on the association probability.
[0096] In some embodiments, the next character of the reference character in the lyric text may be determined according to the character order, and a candidate phoneme set may be determined according to the next character, the pronunciation order of phonemes, and the phoneme associated with the (i - 1)-th audio frame. Specifically, it may be to obtain the first phoneme among the phonemes of the pronunciation of the next character of the reference character in the lyric text, and determine the candidate phoneme set associated with the i-th audio frame based on the first phoneme, the phoneme associated with the (i - 1)-th audio frame, and the pronunciation order corresponding to the phoneme of the reference character's pronunciation.
[0097] It can be understood that multiple audio frames are sequentially associated with each phoneme of the pronunciation of each character. For each original phoneme in each phoneme, there must be an associated audio frame. For the extended phonemes in each phoneme, there may or may not be an associated audio frame. The phoneme associated with the current audio frame may be the phoneme associated with the previous audio frame (i.e., the pronunciation of the phoneme has not ended or has not changed) or may change to the next phoneme (i.e., the pronunciation of the phoneme has ended or has changed). When an audio frame is associated with the last original phoneme, the next audio frame may still be in the pronunciation process of the character, so the next audio frame may still be the last original phoneme, or at this time, the pronunciation of the last original phoneme may change, so the next audio frame may be an extended phoneme, or at this time, the pronunciation process of the character may end and the pronunciation of the next character may start, so the next audio frame may be the first original phoneme of the next character.
[0098] In some embodiments, the candidate phoneme set associated with the i-th audio frame may specifically include one or more of the following situations: If the phoneme associated with the (i - 1)-th audio frame is any original phoneme other than the last original phoneme among the original phonemes of the reference character's pronunciation, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the next phoneme of the phoneme associated with the (i - 1)-th audio frame in the phonemes of the reference character's pronunciation, or the default phoneme; the default phoneme represents no phoneme.
[0099] If the phoneme associated with the (i - 1)-th audio frame is the last original phoneme of the reference character's pronunciation and the last original phoneme does not have an extended phoneme, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the above-mentioned first phoneme, or the default phoneme.
[0100] If the phoneme associated with the (i - 1)-th audio frame is the sustained phoneme of the reference character's pronunciation or any extended phoneme other than the last extended phoneme among the extended phonemes of the sustained phoneme, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the next phoneme of the phoneme associated with the (i - 1)-th audio frame in the phonemes of the reference character's pronunciation, the first phoneme, or the default phoneme.
[0101] If the phoneme associated with the (i - 1)-th audio frame is the last extended phoneme of the pronunciation of the reference character, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the first phoneme, or the default phoneme.
[0102] In some embodiments, when i = 1, the candidate phoneme set associated with the first audio frame includes the first original phoneme of the first character in the lyric text and the default phoneme. If the default phoneme is associated with the first audio frame, the candidate phoneme set associated with the next audio frame is the same as the aforementioned candidate phoneme set until there is an audio frame associated with the first original phoneme of the first character.
[0103] In some embodiments, if the phoneme associated with the i-th audio frame is determined to be the default phoneme, the candidate phoneme set associated with the (i + 1)-th audio frame is determined according to the above determination process and based on the phoneme associated with the (i - 1)-th audio frame. That is, when the phoneme associated with the previous audio frame of the current audio frame is the default phoneme, the candidate phoneme set associated with the current audio frame can be determined based on the nearest audio frame associated with a non-default phoneme.
[0104] It can be understood that the extended phoneme can be used as a backup phoneme for the prolonged phoneme to jointly represent the pronunciation process. The probability that the audio frame corresponding to the pronunciation change of a phoneme is predicted to be associated with other phonemes is relatively high. Therefore, during the pronunciation process of a character, the audio frame corresponding to the pronunciation change of the last original phoneme may be predicted as other phonemes (such as the first phoneme or the default phoneme of the pronunciation of the next character). When the pronunciation change at this time represents the pronunciation variation of the same phoneme, there may be a situation where the phoneme predicted for the audio frame is incorrect, resulting in a certain error in the alignment between the lyric text and the singing audio. That is to say, if a phoneme model is trained using a phoneme dictionary without extension, the predicted pronunciation duration of a phoneme by this model may be inaccurate.
[0105] Therefore, by training the phoneme model with a phoneme dictionary with phoneme extension, when the phoneme model predicts the phoneme associated with an audio frame, it will also consider the extended phoneme as a backup phoneme when it is highly likely to be predicted as other phonemes, so that the audio frame in this case can have a certain probability of being predicted as the extended phoneme. Since this extended phoneme and the prolonged phoneme represent the same phoneme, a relatively complete audio segment corresponding to the same phoneme can be obtained according to the audio frames corresponding to these two phonemes, thereby improving the alignment accuracy and reliability, as well as the flexibility of phoneme prediction, and avoiding the problem that characters with long pronunciations in the lyric text cannot be aligned and thus the correct notes cannot be determined. By training the model with an extended phoneme dictionary, the model can learn the pronunciation change characteristics during different pronunciation processes of a phoneme and the characteristic information of more phonemes (extended phonemes).
[0106] For example, the actual duration of the audio segment corresponding to the phoneme "a" is 3 seconds. However, due to a pitch increase at the 2nd second during singing, when the phoneme model predicts the audio frame corresponding to this pitch increase point, it may be predicted as the first phoneme of the pronunciation of the next character, resulting in an incorrect prediction of this audio frame and affecting the prediction results of subsequent audio frames. Also, the predicted duration of the audio segment for the phoneme "a" may be 2 seconds. For the extended phoneme dictionary, the phonemes obtained for the phoneme "a" can include the original phoneme "a1" and the extended phoneme "a2". It can be achieved that the phoneme associated with the audio frame during the continuous pronunciation process after a 2-second pronunciation change is "a2", thereby making the predicted duration of the audio segment for the phoneme "a" more accurate.
[0107] For example, as Figure 7a - Figure 7b shown shown, Figure 7a and Figure 7b FIG. is a schematic diagram of a scenario for determining an audio segment corresponding to a phoneme provided by an embodiment of the present application. Among them, the singing audio includes multiple audio frames, the lyric text includes character 1, character 2, and character 3, and the phonemes obtained for the pronunciation of character 1 through the extended phoneme dictionary include phoneme A and phoneme B1, the phonemes obtained for the pronunciation of character 2 include phoneme C, phoneme D1, and phoneme D2, and the phonemes obtained for the pronunciation of character 3 include phoneme E1. A possible associated phoneme path for a singing audio is as Figure 7a shown. Therefore, the candidate phoneme set associated with the 1st audio frame includes phoneme A and the default phoneme. If the phoneme associated with the current audio frame is phoneme B1, then the candidate phoneme set associated with the next audio frame includes phoneme B1, phoneme C, and the default phoneme. If the phoneme associated with the current audio frame is phoneme C, then the candidate phoneme set associated with the next audio frame includes phoneme C, phoneme D1, and the default phoneme. If the phoneme associated with the current audio frame is phoneme D1, then the candidate phoneme set associated with the next audio frame includes phoneme D1, phoneme D2, phoneme E1, and the default phoneme. If the phoneme associated with the current audio frame is phoneme D2, then the candidate phoneme set associated with the next audio frame includes phoneme D2, phoneme E1, and the default phoneme.
[0108] As Figure 7b, which can represent the alignment relationship between singing audio and phonemes. Taking the character 2 as an example, each phoneme in the character 2 corresponds to an audio segment, which is obtained by splicing the audio frames associated with the corresponding phonemes. The audio segments corresponding to the phonemes C, D1, and D2 are spliced to obtain the audio segment of the character 2 in the singing audio. Moreover, the phonemes D1 and D2 represent the same phoneme D, and the audio segments corresponding to the phonemes D1 and D2 are spliced to obtain the audio segment of the phoneme D (i.e., the target audio segment of the original phoneme D1). The audio segment of each character can be obtained based on the audio segments of each phoneme. In addition, compared with the unexpanded phoneme dictionary, the phonemes obtained for the pronunciation of the character 2 are the phonemes C and D1. For the pronunciation process of the character 2 in this singing audio, when aligning only through the phonemes C and D1, it may cause this pronunciation process to be unable to be represented completely, and there will be a negative impact on the subsequent alignment of audio frames, that is, there may be many alignment errors and inaccurate alignment durations. For example, the pronunciation process of the character 2 represented by the phonemes C, D1, and D2 can be 3 seconds, while the pronunciation process of the character 2 represented by the phonemes C and D1 is only within 0-1 second, and the audio frames corresponding to the remaining pronunciation process may be recognized as the phonemes associated with the pronunciation of the next character. For example, the pronunciation process within 2-3 seconds is also recognized as the pronunciation process of the phoneme E1, resulting in an error in the alignment result of the character 3.
[0109] In some embodiments, it may be to use a phoneme model to predict the association probability between each candidate phoneme in the candidate phoneme set and the i-th audio frame based on the audio sequence of the i-th audio frame. Specifically, the phoneme model can be called to predict the target audio sequence of the i-th audio frame to obtain the initial association probability between each candidate phoneme and the i-th audio frame. Based on the phoneme model, the phoneme transition matrix corresponding to the phoneme associated with the (i-1)-th audio frame is obtained, and the association probability between each candidate phoneme and the i-th audio frame is determined according to the initial association probability and the phoneme transition matrix. The target audio sequence can be the audio sequence of the i-th audio frame, or an ordered splicing sequence of the i-th audio frame and adjacent audio frames (the (i-1)-th and (i+1)-th). That is, a sequence that combines the upper and lower audio information of the i-th audio frame. For example, the target audio sequence of the 4th audio frame is the spliced audio sequence of the audio sequences of the 3rd audio frame, the 4th audio frame, and the 5th audio frame respectively. When an audio frame has no adjacent audio frames, a default audio sequence can be preset to replace the audio sequence of the adjacent audio frame.
[0110] In some embodiments, the phoneme model can be, for example, a GMM-HMM model. The GMM model contains Gaussian distribution models corresponding to each phoneme. The Gaussian distribution model is used to characterize the distribution of the audio sequence of the audio frames corresponding to the phoneme. Determining the initial association probability can be to call the Gaussian distribution models corresponding to each candidate phoneme to predict the target audio sequence (such as the audio sequence) of the i-th audio frame, and determine the probability that the audio sequence of the i-th audio frame conforms to the distribution indicated by each Gaussian distribution model. This probability can characterize the probability that the phoneme associated with the i-th audio frame is the candidate phoneme, and this probability is determined as the association probability between each candidate phoneme and the i-th audio frame respectively. The HMM model contains a state transition model for each phoneme. The state transition model is used to characterize the phoneme transition matrix of the phoneme. The phoneme transition matrix indicates the probability that a phoneme transitions to each phoneme (including the default phoneme) respectively.
[0111] In some embodiments, determining the association probability between each candidate phoneme and the i-th audio frame according to the initial association probability and the phoneme transition matrix can specifically be to determine the transition probability that the phoneme associated with the (i - 1)-th audio frame transitions to each candidate phoneme from the phoneme transition matrix, and determine the product of the initial association probability corresponding to each candidate phoneme and the corresponding transition probability as the association probability corresponding to each candidate phoneme.
[0112] For example, the candidate phonemes of the i-th audio frame include phoneme 1 and phoneme 2. By predicting the audio sequence of the i-th audio frame through the Gaussian distribution model corresponding to phoneme 1, the initial association probability between phoneme 1 and the i-th audio frame is 0.5. By predicting the audio sequence of the i-th audio frame through the Gaussian distribution model corresponding to phoneme 2, the initial association probability between phoneme 2 and the i-th audio frame is 0.7. Obtain the transition matrix of the phoneme associated with the (i - 1)-th audio frame (such as phoneme 1), and get the probability that phoneme 1 transitions to phoneme 1 is 0.3, and the probability that phoneme 1 transitions to phoneme 2 is 0.5. Therefore, the association probability between phoneme 1 and the i-th audio frame is 0.5 * 0.3 = 0.15, and the association probability between phoneme 2 and the i-th audio frame is 0.5 * 0.7 = 0.35.
[0113] In some embodiments, when the above phoneme model is a GMM-HMM model, the training process may be as follows: obtain a sample audio, divide the sample audio into N audio frames, and obtain the sample phoneme sequence corresponding to the sample audio. The sample phoneme sequence is the phonemes sequentially associated with the N audio frames, which are set by relevant business personnel according to the specific sample audio. Distribute each phoneme in the sample phoneme sequence equally to the N audio frames in turn to represent the associated phoneme states of the first round of training. Construct the Gaussian distribution model corresponding to the phoneme based on the audio sequences of the audio frames corresponding to the same phoneme, and use the phonemes associated with each audio frame to determine the transition probability of each phoneme under the sample phoneme sequence to obtain the corresponding phoneme transition matrix. Use the Gaussian distribution model and the phoneme state probability matrix at this time to correct the associated phoneme states of the first round to obtain the associated phoneme states of the second round, and correct the previously constructed Gaussian distribution model according to the above process. The corrected Gaussian distribution model represents the audio sequence distribution of the new audio frames corresponding to the phoneme, and re-determine the phoneme transition matrix. Continuously iterate the above process to obtain the finally trained Gaussian distribution model and phoneme transition matrix as the model parameters of the phoneme model.
[0114] For example, sample audio 1 includes 10 audio frames, and sample audio sequence 1 is phoneme 1, phoneme 2, and phoneme 3 in sequence. Therefore, set the phonemes associated with audio frames 1-3 as phoneme 1, the phonemes associated with audio frames 4-7 as phoneme 2, and the phonemes associated with audio frames 8-10 as phoneme 3. Construct Gaussian distribution model 1 corresponding to phoneme 1 based on the audio sequence of audio frames 1-3, construct Gaussian distribution model 1 corresponding to phoneme 2 based on the audio sequence of audio frames 4-7, construct Gaussian distribution model 1 corresponding to phoneme 3 based on the audio sequence of audio frames 8-10, determine phoneme transition matrix 1 based on the above associated phoneme states, and re-determine the phonemes associated with each audio frame in turn according to the Gaussian distribution model 1 and phoneme transition matrix 1 as the associated phoneme states of the second round of training. Correct the parameters in Gaussian distribution model 1 and phoneme transition matrix 1 according to the new associated phoneme states, and continuously iterate the above process to obtain the final Gaussian distribution model 1 and phoneme transition matrix 1.
[0115] Accordingly, the sample audio 2 includes 10 audio frames, and the sample audio sequence 2 is phoneme 1, phoneme 2, phoneme 3 in sequence. Therefore, the phonemes associated with audio frames 1 - 3 are set as phoneme 1, the phonemes associated with audio frames 4 - 7 are set as phoneme 2, and the phonemes associated with audio frames 8 - 10 are set as phoneme 3. The phonemes associated with audio frames 1 - 3 are set as phoneme 1, the phonemes associated with audio frames 4 - 7 are set as phoneme 2, and the phonemes associated with audio frames 8 - 10 are set as phoneme 3. The Gaussian distribution model 1 corresponding to the above phoneme 1 is corrected according to the audio sequence of audio frames 1 - 3, the Gaussian distribution model 1 corresponding to the above phoneme 2 is corrected according to the audio sequence of audio frames 4 - 7, the Gaussian distribution model 1 corresponding to phoneme 3 is corrected according to the audio sequence of audio frames 8 - 10, and the phoneme transition matrix 2 is determined based on the above associated phoneme states. Based on the corrected Gaussian distribution model and phoneme transition matrix, the phoneme associated with each audio frame is determined in sequence, and the above process is continuously iterated to obtain the finally corrected Gaussian distribution model 1 (Gaussian distribution model 2). The phoneme transition matrix 1 is corrected by using the obtained phoneme transition matrix 2. For example, in the two matrices, for the transition probability of the same phoneme to the same phoneme, the average is obtained as the final phoneme transition probability of this phoneme, and the corrected matrix 2 is determined as the final phoneme transition matrix 2. Thus, the finally trained phoneme model can be determined according to a large number of sample audios and sample audio sequences according to the above process.
[0116] In some embodiments, the phoneme model can be, for example, a DNN - HMM model. The DNN model can predict the probability of each phoneme associated with an audio frame. Determining the initial association probability can be to call the DNN model to predict the target audio sequence (such as an ordered splicing sequence) of the i - th audio frame, and based on the prediction result, obtain the probability of the audio sequence of the i - th audio frame associated with each candidate phoneme as the corresponding association probability. The phoneme transition matrix of each phoneme is included in the HMM model, and this phoneme transition matrix indicates the probability of a phoneme transitioning to each phoneme respectively.
[0117] In some embodiments, the manner of determining the association probability between each candidate phoneme and the i - th audio frame according to the initial association probability and the phoneme transition matrix can be the same as the manner of determining the association probability for the phoneme model of the GMM - HMM model described above.
[0118] In some embodiments, when the above phoneme model is a DNN-HMM model, the training process may be as follows: use the pre-trained GMM-HMM model to label the phoneme tags for each audio frame in the sample audio, and train the DNN model through the sample audio sequence of each audio frame and the corresponding phoneme tags to obtain the trained DNN model. For example, the DNN model can be used to predict the sample audio sequence to obtain the phoneme prediction result, determine the prediction deviation based on the phoneme prediction result and the phoneme tag, and use the prediction deviation to correct the model parameters of the DNN model until the model converges. The model parameters in this HMM model can be the relevant parameters in the pre-trained GMM-HMM model. For example, the phoneme transition matrix of each phoneme in the DNN-HMM model can be obtained from the GMM-HMM model during training.
[0119] In some embodiments, determining the phoneme associated with the i-th audio frame based on the association probability between each candidate phoneme and the i-th audio frame respectively may specifically be to determine the candidate phoneme corresponding to the maximum association probability as the phoneme associated with the i-th audio frame.
[0120] S604. Determine the audio segment corresponding to each phoneme of the pronunciation of the characters in the lyrics text in the singing audio respectively based on the phoneme associated with each audio frame.
[0121] S605. Determine the target audio segment corresponding to each original phoneme in the lyrics text according to the audio segment corresponding to each phoneme respectively. The specific implementation manners of steps S604 - S605 can refer to the relevant steps in the above embodiments and will not be elaborated here.
[0122] S606. Generate the musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyrics text.
[0123] In some embodiments, the electronic device may obtain the pitch curve of the target audio segment corresponding to each original phoneme, and generate the musical score of the singing audio based on the pitch curve and the lyrics text. Among them, it may be to obtain the complete pitch curve of the singing audio and extract the pitch curve corresponding to each target audio segment from the complete pitch curve; or it may also be to obtain the pitch curve corresponding to each target audio segment separately. The method of determining the pitch curve may be to divide the singing audio (or the target audio segment) into multiple audio sub-segments, and extract the fundamental frequency of each audio sub-segment (when the sounding body makes a sound due to vibration, the sound can generally be decomposed into many simple sine waves, that is, all natural sounds are basically composed of many sine waves with different frequencies, and the sine wave with the lowest frequency is the fundamental tone, and the frequency of the fundamental tone is the fundamental frequency). According to the corresponding relationship between the fundamental frequency and the pitch, determine the pitch corresponding to each audio sub-segment (or target audio segment), and generate the pitch curve of the singing audio (or target audio segment) according to the pitch corresponding to each audio sub-segment. The generation method of the fundamental frequency or the pitch curve may be generated by relevant audio tools or executed by the electronic device. For example, it may be the MFA (Montreal Forced Alignment) tool, and the MFA tool can be used to output the waveform diagram, fundamental frequency diagram, phoneme alignment information, etc. of the singing audio. The corresponding relationship between the fundamental frequency and the pitch can be specifically set by relevant business personnel.
[0124] In some embodiments, generating the musical score of the singing audio based on the pitch curve and the lyrics text may be to respectively determine at least one note associated with each original phoneme according to the pitch curve corresponding to each original phoneme, and generate the musical score according to at least one note associated with each original phoneme and the lyrics text. The musical score can be understood as the result of audio transcription of the singing audio based on the lyrics text, that is, the process of converting the singing audio and the lyrics text into a representation form of a note sequence (such as including pronunciation duration and pitch sequence) (for example, MIDI (Musical Instrument Digital Interface)) through association processing. Among them, generating the musical score according to the associated notes and the lyrics text can be generated corresponding to the musical score generation rules, and the specific method can refer to the method of generating the musical score in the above embodiments.
[0125] In some embodiments, any one of the original phonemes of the pronunciation of the characters in the lyric text is represented as a target original phoneme. Determining at least one note associated with the original phoneme can be determined according to the change in the curve shape of the pitch curve. In this way, the associated notes and the pitch corresponding to the notes can be determined more flexibly in combination with the specific changes in the musical score, and the accuracy of the pitch corresponding to the notes can be improved. Specifically, determining the note associated with the original phoneme can be, if the pitch curve corresponding to the target original phoneme belongs to a vibrato curve, then determining the note associated with the target original phoneme based on the average pitch indicated by the pitch curve corresponding to the target original phoneme. Specifically, the average pitch can be determined as the pitch of the note, and the pronunciation duration of the target original phoneme can be determined as the pronunciation duration of the note. Among them, if the pitch curve segment used to determine the pitch of the note belongs to a vibrato curve, a vibrato identifier can be generated for the note, and the vibrato identifier can be displayed at the corresponding position on the musical score to indicate that the singing method of the corresponding phoneme belongs to the vibrato singing method, etc.
[0126] The above-mentioned vibrato curve is a curve segment with stable fluctuations in the pitch curve. Vibrato can be understood as the fluctuation of the pitch curve with a stable frequency, that is, the periodic change of the pitch within a certain frequency range. A pitch curve can have one or more segments of vibrato curves. The method for determining the vibrato curve can be to detect the morphological change of the pitch curve, determine the fluctuation interval with stable fluctuations and a fluctuation process greater than a specified threshold (such as 1 second), and determine the pitch curve corresponding to the fluctuation interval as the vibrato curve. Or, it can also be to perform frequency analysis on the pitch curve. When analyzing, all the pitch curves corresponding to the original phonemes can be analyzed together or separately, or the complete pitch curve corresponding to the singing audio can be analyzed. For example, the frequency analysis can be implemented by the FDM (Frequency-division multiplexing) algorithm or the FFT (fast Fourier transform) algorithm. The FDM algorithm has higher resolution in the low-frequency band compared with the FFT algorithm and can more accurately find the frequency range of the vibrato vibration. If the frequency of a segment of the pitch curve is within the vibrato frequency range and the frequency changes periodically, then this segment of the curve can be determined as the vibrato curve. By identifying the vibrato interval in the singing audio, it is helpful for the division of notes and the determination of pitch.
[0127] The above-mentioned pitch curve belonging to the vibrato curve means that the part of the pitch curve that is the vibrato curve needs to be greater than or equal to a specified range. For example, if 80% of the curve part of the pitch curve is the vibrato curve, then it is determined that the pitch curve belongs to the vibrato curve. Otherwise, the pitch curve can be divided, and multiple notes associated with the target original phoneme can be determined according to the divided pitch curve segments. The determination method based on the pitch curve segment can be the same as the method for determining the associated note based on the pitch curve.
[0128] The above division method can be as follows: if the part that is the vibrato curve is the starting segment (or ending segment) of the pitch curve, then the pitch curve is divided into two pitch curve segments, one is the pitch curve segment corresponding to the part that is the vibrato curve, and the other is the remaining pitch curve; if the part that is the vibrato curve is the middle segment of the pitch curve, then the pitch curve can be divided into three pitch curve segments, one is the pitch curve segment corresponding to the part that is the vibrato curve, one is the segment corresponding to the starting segment, and one is the segment corresponding to the ending segment. Or, further, if the segment corresponding to the starting segment (or ending segment) is less than the specified length threshold (such as 0.03 seconds or 5% of the pitch curve, etc.), then the segment corresponding to the starting segment (or ending segment) and the pitch curve segment corresponding to the part that is the vibrato curve can also be spliced to be used as one pitch curve segment;
[0129] If the pitch curve includes two vibrato curves, then the pitch curve can be divided into multiple pitch curve segments, including the segment corresponding to the part that is the first vibrato curve, the segment corresponding to the part that is the second vibrato curve, and the segment corresponding between the two vibrato curves. Or, further, if the segment corresponding between the two vibrato curves is less than the specified length threshold, then the middle segment and the segments corresponding to the two vibrato curves can also be spliced to be used as one pitch curve segment.
[0130] In some embodiments, determining the note associated with the original phoneme specifically can be: if the pitch curve corresponding to the target original phoneme does not belong to the vibrato curve, then detect the pitch mutation points in the pitch curve corresponding to the target original phoneme, and determine at least one note associated with the target original phoneme based on the detection result. The pitch mutation point can be a point where the pitch curve suddenly rises, suddenly drops, or suddenly slows down, etc. Specifically, it can be set by relevant business personnel.
[0131] In some embodiments, determining the associated note based on the detection result specifically can be: if the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is no pitch mode, then determine the note associated with the target original phoneme according to the ending pitch of the pitch curve corresponding to the target original phoneme. Specifically, the ending pitch can be determined as the pitch of the note, and the pronunciation duration of the target original phoneme can be determined as the pronunciation duration of the note.
[0132] In some embodiments, determining the associated note based on the detection result specifically can be: if the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is a pitch mode, then determine the note associated with the target original phoneme according to the pitch mode. Specifically, the pitch mode can be determined as the pitch of the note, and the pronunciation duration of the target original phoneme can be determined as the pronunciation duration of the note.
[0133] In some embodiments, determining the associated notes based on the detection result may specifically be as follows: If the detection result includes a pitch mutation point detected in the pitch curve corresponding to the target original phoneme, then the pitch curve corresponding to the target original phoneme is divided based on the detected pitch mutation point to obtain multiple pitch curve segments of the pitch curve corresponding to the target original phoneme, and the notes associated with the target original phoneme are determined according to each pitch curve segment. At this time, the target original phoneme is associated with multiple notes. Among them, the specific method for determining the associated notes based on each pitch curve segment may be the same as the method for determining the associated notes based on the pitch curve described above. The pronunciation duration of the notes determined according to each pitch curve segment is the duration indicated by the pitch curve segment in the singing audio.
[0134] For example, as Figure 8 shown, Figure 8 is a schematic diagram of a scenario for determining the phonemes associated with a phoneme provided by an embodiment of the present application; among them, the pitch curve corresponding to phoneme 1 is from 0 to 0.8 seconds, the pitch curve corresponding to phoneme 2 is from 0.8 to 1.2 seconds, the pitch curve corresponding to phoneme 3 is from 1.2 to 1.4 seconds, and the pitch curve corresponding to phoneme 4 is from 1.4 to 2 seconds; frequency analysis of the pitch curve using the FDM algorithm and the FFT algorithm respectively obtains analysis results indicating the frequency information corresponding to pitch curves with different morphological changes, and the frequency range of vibrato vibration is approximately between 4 - 8 Hz, and the interval with periodic frequency changes within this frequency range is the vibrato interval, that is, the pitch curve corresponding to 0.02 seconds - 1 second is the vibrato curve;
[0135] Phoneme 1: The part of the pitch curve that is the vibrato curve is greater than the specified range, that is, it belongs to the vibrato curve, then the average pitch of the pitch curve is determined as the pitch of the note associated with phoneme 1, and the pronunciation duration of phoneme 1 is the pronunciation duration of the associated note;
[0136] Phoneme 2: The part of the pitch curve that is the vibrato curve is less than the specified range, then the pitch curve is divided into two pitch curve segments, which are 0.8 - 1 second and 1 - 1.12 seconds respectively. The average pitch of the pitch curve segment 1 (belonging to the vibrato curve) corresponding to 0.8 - 1 second is determined as the pitch of note 1 associated with phoneme 2, and the pronunciation duration of this note is the pronunciation duration indicated by the pitch curve segment; the pitch mode or the ending pitch of the pitch curve segment 2 (not a vibrato curve and without a pitch mutation point) corresponding to 1 - 1.12 seconds is determined as the pitch of note 2 associated with phoneme 2, and the pronunciation duration of this note is the pronunciation duration indicated by the pitch curve segment;
[0137] Phoneme 3: The pitch curve does not belong to the vibrato curve and there is no pitch mutation point. If the pitch curve has a pitch mode, the pitch mode is determined as the pitch of the note associated with phoneme 3; if there is no pitch mode, the ending pitch (i.e., the pitch corresponding to 1.4 seconds) is determined as the pitch of the note associated with phoneme 3, and the pronunciation duration of phoneme 3 is the pronunciation duration of the associated note.
[0138] Phoneme 4: The pitch curve does not belong to the vibrato curve and there are pitch mutation points (1.6 seconds and 1.8 seconds). Then the pitch curve is divided into pitch curve segment 1, pitch curve segment 2, and pitch curve segment 3, which are 1.4 - 1.6 seconds, 1.6 - 1.8 seconds, and 1.8 - 2 seconds respectively. And the notes (note 1, note 2, and note 3) associated with phoneme 4 are determined respectively based on each pitch curve segment according to the above process.
[0139] In some embodiments, different types of sheet music can be generated according to different scenarios. The sheet music can be applied in a singing software to guide the user's singing. Or it can be used to compare the sheet music of the user's actual sung audio with the sheet music of the standard version of the sung audio for scoring the user's singing. Or the sheet music can be used as an input for singing synthesis to generate an AI song with a specified timbre. Or it can also be based on the alignment relationship between the lyric text and the sung audio to provide accurate lyric timestamps for a music player software.
[0140] For example, as Figure 9 shown, Figure 9Schematic diagram of a framework for audio data processing provided by an embodiment of this application; among them, ① obtain singing audio and the lyric text associated with the singing audio; for example, it can be to obtain a mixed audio of human voice and accompaniment, and perform sound source separation on the mixed audio to obtain the vocal track audio as the singing audio; ② obtain the untoned pinyin of each character in the lyric text and the extended phoneme dictionary (the untoned audio of a character is represented by several phonemes, including the original phonemes, or may also include extended phonemes); for example, the lyric text is "He is tumbling only by strength", the untoned pinyin is "ta jinping li liang zai fan teng a", and obtain the corresponding pronunciation phonemes from the phoneme dictionary, such as he - ta - t / a1 / a2; ③ in the alignment module, perform phoneme alignment according to the phoneme dictionary, untoned pinyin, and singing audio to obtain phoneme timestamps, that is, the target audio segment corresponding to each original phoneme; for example, the MFA tool can be used in the alignment module to output phoneme timestamps. Specifically: obtain the phonemes of each character's pronunciation according to the untoned pinyin of each character, and generate the audio sequence of each audio frame in the singing audio. Based on the audio sequence, the trained phoneme model predicts the phoneme associated with each audio frame in the singing audio, thereby achieving alignment; the training process can be obtained by training hundreds of sample audios and the sample phoneme sequences corresponding to the sample audios. The sample phoneme sequence can be determined based on the untoned audio of each character in the sample text corresponding to the sample audio and the phoneme dictionary; ④ the MFA tool can also output the waveform diagram and fundamental frequency diagram of the singing audio. The output result of the MFA tool is as Figure 10a ; this fundamental frequency diagram (or the pitch curve used later) is an intermediate processing result of the MFA tool (or phoneme model); ⑤ determine the pitch curve corresponding to each original phoneme according to the fundamental frequency diagram; ⑥ in the vibrato recognition module, perform frequency analysis on the pitch curve to determine the vibrato curve and obtain the vibrato interval in the singing audio. The display of the vibrato interval on the pitch curve is as Figure 10b ; ⑦ generate a musical score according to the phoneme timestamps, pitch curve, and vibrato curve; specifically: determine the notes associated with the original phonemes according to the phoneme timestamps, pitch curve, and vibrato curve, and generate a musical score according to the notes associated with the original phonemes and the characters to which the original phonemes belong in the lyric text.
[0141] In an embodiment of the present application, a singing audio can be obtained, and a lyric text associated with the singing audio can be obtained. Phonemes of the pronunciation of each character in the lyric text are respectively obtained. The phonemes of the pronunciation of the character can include a sustaining phoneme and an extended phoneme of the same phoneme after the pronunciation change of the sustaining phoneme. The singing audio is divided into multiple audio frames, and the phoneme associated with each audio frame is determined based on the character order between each character in the lyric text and the pronunciation order corresponding to the phoneme of the pronunciation of each character. An audio segment corresponding to each phoneme of the pronunciation of the characters in the lyric text is determined based on the phoneme associated with each audio frame. A target audio segment corresponding to each original phoneme of the characters in the lyric text is determined according to the audio segment corresponding to each phoneme. A musical score of the singing audio is generated according to the target audio segment corresponding to each original phoneme and the lyric text. Through the above method, when using a sustaining phoneme to represent the pronunciation process of a phoneme, if the pronunciation of the phoneme changes, an extended phoneme can be used to further represent the pronunciation process of the same phoneme after the pronunciation change. Therefore, a more complete pronunciation process of a phoneme can be represented by the sustaining phoneme and the extended phoneme, and the target audio segment corresponding to each original phoneme of the pronunciation of each character can be determined more accurately, so as to achieve an accurate alignment at the phoneme level between the lyric text and the singing audio, and generate a musical score based on the alignment result, which helps to improve the efficiency and accuracy of musical score generation.
[0142] The method of the embodiment of the present application is elaborated in detail above. To facilitate better implementation of the above solution of the embodiment of the present application, correspondingly, a device of the embodiment of the present application is provided below.
[0143] Figure 11 The structural schematic diagram of an audio data processing device provided by an exemplary embodiment of the present application is shown; the audio data processing device can be used as a computer program (including program code) running in an electronic device. For example, the audio data processing device can be an application program in the electronic device; the audio data processing device can be used to execute Figure 2 and Figure 6 Some or all of the steps in the method embodiment shown. Please refer to Figure 11 ; the audio data processing device includes the following modules:
[0144] An obtaining module 1101, configured to obtain a singing audio and obtain a lyric text associated with the singing audio;
[0145] An acquisition module 1101 is further configured to respectively acquire phonemes of the pronunciation of each character in the lyric text; the lyric text includes a target character, and the phonemes of the pronunciation of the target character include at least one original phoneme of the target character and an extended phoneme of a sustained phoneme, where the sustained phoneme refers to the last original phoneme in the original phonemes of the target character, and the extended phoneme of the sustained phoneme is used to represent the same phoneme after the pronunciation of the sustained phoneme changes;
[0146] A processing module 1102 is configured to predict an audio segment corresponding to each phoneme of the pronunciation of a character in the lyric text in the singing audio respectively; any one original phoneme corresponds to an audio segment, and any one extended phoneme corresponds to an audio segment or does not correspond to an audio segment;
[0147] The processing module 1102 is further configured to determine a target audio segment corresponding to each original phoneme of the character in the lyric text according to the audio segment corresponding to each phoneme respectively, and generate a musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text.
[0148] In some embodiments, any one of the phonemes of the pronunciation of a character in the lyric text is represented as a target phoneme; when the processing module 1102 is used to predict an audio segment corresponding to each phoneme of the pronunciation of a character in the lyric text in the singing audio respectively, it is specifically configured to:
[0149] Divide the singing audio into multiple audio frames, and represent each audio frame as a corresponding audio sequence;
[0150] Based on the audio sequence of each audio frame, the character order between each character in the lyric text, and the pronunciation order corresponding to the phoneme of the pronunciation of each character, predict the phoneme associated with each audio frame respectively;
[0151] If the phonemes associated with M consecutive audio frames among multiple audio frames are all target phonemes, perform splicing processing on the M consecutive audio frames, and determine the spliced audio frame as the audio segment corresponding to the target phoneme.
[0152] In some embodiments, each audio frame includes the i-th audio frame, i is an integer and i>1, and the character to which the phoneme associated with the i-th audio frame belongs is represented as a reference character;
[0153] When the processing module 1102 is used to predict the phoneme associated with each audio frame respectively based on the audio sequence of each audio frame, the character order between each character in the lyric text, and the pronunciation order corresponding to the phoneme of the pronunciation of each character, it is specifically configured to:
[0154] Acquire the first phoneme in the phonemes of the pronunciation of the next character of the reference character in the lyric text; the next character of the reference character is determined based on the character order;
[0155] Determine a candidate phoneme set associated with the i-th audio frame based on the pronunciation order corresponding to the phoneme associated with the first phoneme, the (i-1)-th audio frame, and the phoneme of the reference character pronunciation; the candidate phoneme set includes at least one candidate phoneme;
[0156] Predict the association probability between each candidate phoneme in the candidate phoneme set and the i-th audio frame respectively based on the audio sequence of the i-th audio frame;
[0157] Select the phoneme associated with the i-th audio frame from the candidate phoneme set based on the association probability between each candidate phoneme and the i-th audio frame respectively.
[0158] In some embodiments, if the phoneme associated with the (i-1)-th audio frame is any one of the original phonemes in the reference character pronunciation except the last original phoneme, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i-1)-th audio frame, the next phoneme of the phoneme associated with the (i-1)-th audio frame in the phonemes of the reference character pronunciation, or, the default phoneme;
[0159] If the phoneme associated with the (i-1)-th audio frame is the last original phoneme of the reference character pronunciation and the last original phoneme does not have an extended phoneme, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i-1)-th audio frame, the first phoneme, or, the default phoneme;
[0160] If the phoneme associated with the (i-1)-th audio frame is a sustained phoneme of the reference character pronunciation or any one of the extended phonemes of the sustained phoneme except the last extended phoneme, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i-1)-th audio frame, the next phoneme of the phoneme associated with the (i-1)-th audio frame in the phonemes of the reference character pronunciation, the first phoneme, or, the default phoneme;
[0161] If the phoneme associated with the (i-1)-th audio frame is the last extended phoneme of the reference character pronunciation, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i-1)-th audio frame, the first phoneme, or, the default phoneme.
[0162] In some embodiments, any one of the original phonemes in the character pronunciation of the lyric text is denoted as the target original phoneme; when the processing module 1102 is used to determine the target audio segment corresponding to each original phoneme of the character in the lyric text according to the audio segment corresponding to each phoneme respectively, it is specifically used for:
[0163] If the target original phoneme does not have an extended phoneme, determine the audio segment corresponding to the target original phoneme as the target audio segment corresponding to the target original phoneme;
[0164] If the target original phoneme has an extended phoneme and there is an audio segment corresponding to the extended phoneme of the target original phoneme, the audio segment corresponding to the target original phoneme and the audio segment corresponding to the extended phoneme of the target original phoneme are spliced, and the spliced audio segment is determined as the target audio segment corresponding to the target original phoneme.
[0165] In some embodiments, when the processing module 1102 is used to generate the musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyrics text, it is specifically used for:
[0166] Obtain the pitch curve of the target audio segment corresponding to each original phoneme;
[0167] Determine at least one note associated with each original phoneme according to the pitch curve corresponding to each original phoneme;
[0168] Generate a musical score according to at least one note associated with each original phoneme and the lyrics text.
[0169] In some embodiments, any one of the original phonemes of the pronunciation of the characters in the lyrics text is represented as the target original phoneme; when the processing module 1102 is used to determine at least one note associated with each original phoneme according to the pitch curve corresponding to each original phoneme, it is specifically used for:
[0170] If the pitch curve corresponding to the target original phoneme belongs to a vibrato curve, determine the note associated with the target original phoneme based on the average pitch indicated by the pitch curve corresponding to the target original phoneme;
[0171] If the pitch curve corresponding to the target original phoneme does not belong to a vibrato curve, detect the pitch mutation point in the pitch curve corresponding to the target original phoneme, and determine at least one note associated with the target original phoneme based on the detection result.
[0172] In some embodiments, when the processing module 1102 is used to determine at least one note associated with the target original phoneme based on the detection result, it is specifically used for:
[0173] If the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is no pitch mode, determine the note associated with the target original phoneme according to the ending pitch of the pitch curve corresponding to the target original phoneme;
[0174] If the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is a pitch mode, determine the note associated with the target original phoneme according to the pitch mode;
[0175] If the detection result includes pitch mutation points detected in the pitch curve corresponding to the target original phoneme, the pitch curve corresponding to the target original phoneme is divided based on the detected pitch mutation points to obtain a plurality of pitch curve segments of the pitch curve corresponding to the target original phoneme, and the notes associated with the target original phoneme are determined according to each pitch curve segment respectively.
[0176] According to an embodiment of the present application, Figure 11 each module in the audio data processing device shown can be separately or all combined into one or several other modules to form, or some of the modules can be further split into a plurality of smaller modules in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module can also be implemented by multiple modules, or the functions of multiple modules can be implemented by one module. In other embodiments of the present application, the audio data processing device can also include other modules. In practical applications, these functions can also be assisted by other modules and can be achieved by the cooperation of multiple modules. According to another embodiment of the present application, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding methods shown in Figure 2 and Figure 6 on a general computing device such as a computer including processing elements and storage elements such as a central processing module (CPU), a random access storage medium (RAM), a read-only storage medium (ROM), etc., to construct the audio data processing device shown in Figure 11 and to implement the audio data processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0177] In an embodiment of the present application, an acquisition module acquires singing audio and acquires a lyric text associated with the singing audio; the acquisition module respectively acquires phonemes of the pronunciation of each character in the lyric text, and the phonemes of the pronunciation of the character may include a sustained phoneme and an extended phoneme which is the same phoneme after the pronunciation change of the sustained phoneme; a processing module predicts an audio segment corresponding to each phoneme of the pronunciation of a character in the lyric text in the singing audio; the processing module determines a target audio segment corresponding to each original phoneme of a character in the lyric text according to the audio segment corresponding to each phoneme respectively, and generates a musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text. Through the above device, when using a sustained phoneme to represent the pronunciation process of a phoneme, if the pronunciation of the phoneme changes, an extended phoneme can be used to further represent the pronunciation process of the same phoneme after the pronunciation change, so that a more complete pronunciation process of a phoneme can be represented by the sustained phoneme and the extended phoneme, and the target audio segment corresponding to each original phoneme of the pronunciation of each character can be determined more accurately, so as to achieve accurate alignment at the phoneme level between the lyric text and the singing audio, and generate a musical score based on the alignment result, which helps to improve the efficiency and accuracy of musical score generation.
[0178] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 12 shown, the electronic device 1200 includes: at least one processor 1201 and a memory 1202. Optionally, the electronic device may further include a network interface. Among them, data can be exchanged between the processor 1201, the memory 1202 and the network interface. The network interface is controlled by the processor 1201 to send and receive messages. The memory 1202 is used to store a computer program, and the computer program includes program instructions. The processor 1201 is used to execute the program instructions stored in the memory 1202. Among them, the processor 1201 is configured to call the program instructions to execute the above method.
[0179] Among them, the memory 1202 may include a volatile memory, such as a random-access memory (RAM); the memory 1202 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the memory 1202 may further include a combination of the above types of memories.
[0180] Among them, the processor 1201 can be a central processing unit (CPU). In one embodiment, the processor 1201 can also be a Graphics Processing Unit (GPU). The processor 1201 can also be a combination of a CPU and a GPU.
[0181] In a possible implementation, the memory 1202 is used to store program instructions, and the processor 1201 can call the program instructions to execute the following steps:
[0182] Obtain singing audio and obtain the lyric text associated with the singing audio;
[0183] Obtain the phonemes of the pronunciation of each character in the lyric text respectively; the lyric text includes target characters, and the phonemes of the pronunciation of the target characters include at least one original phoneme of the target character and the extended phoneme of the sustained phoneme. The sustained phoneme refers to the last original phoneme in the original phonemes of the target character, and the extended phoneme of the sustained phoneme is used to represent the same phoneme after the pronunciation of the sustained phoneme changes;
[0184] Predict the audio segments corresponding to each phoneme of the pronunciation of the characters in the lyric text in the singing audio respectively; any one original phoneme corresponds to one audio segment, and any one extended phoneme corresponds to one audio segment or does not correspond to an audio segment;
[0185] Determine the target audio segments corresponding to each original phoneme of the characters in the lyric text according to the audio segments corresponding to each phoneme respectively, and generate the musical score of the singing audio according to the target audio segments corresponding to each original phoneme and the lyric text.
[0186] In some embodiments, any one of the phonemes of the pronunciation of the characters in the lyric text is represented as a target phoneme; when the processor 1201 is used to predict the audio segments corresponding to each phoneme of the pronunciation of the characters in the lyric text in the singing audio respectively, it is specifically used for:
[0187] Divide the singing audio into multiple audio frames, and represent each audio frame as a corresponding audio sequence;
[0188] Based on the audio sequence of each audio frame, the character order between each character in the lyric text, and the pronunciation order corresponding to each phoneme of the pronunciation of each character, predict the phoneme associated with each audio frame respectively;
[0189] If the phonemes associated with M consecutive audio frames among multiple audio frames are all target phonemes, splice the M consecutive audio frames, and determine the spliced audio frame as the audio segment corresponding to the target phoneme.
[0190] In some embodiments, each audio frame includes the i-th audio frame, where i is an integer and i > 1, and the character to which the phoneme associated with the i-th audio frame belongs is represented as a reference character;
[0191] When the processor 1201 is used to predict the phoneme associated with each audio frame based on the audio sequence of each audio frame, the character order between each character in the lyric text, and the pronunciation order corresponding to the phoneme of each character pronunciation respectively, it is specifically used for:
[0192] Obtain the first phoneme in the phonemes of the pronunciation of the next character of the reference character in the lyric text; the next character of the reference character is determined based on the character order;
[0193] Based on the first phoneme, the phoneme associated with the (i - 1)-th audio frame, and the pronunciation order corresponding to the phoneme of the reference character pronunciation, determine the candidate phoneme set associated with the i-th audio frame; the candidate phoneme set includes at least one candidate phoneme;
[0194] Predict the association probability between each candidate phoneme in the candidate phoneme set and the i-th audio frame based on the audio sequence of the i-th audio frame;
[0195] Select the phoneme associated with the i-th audio frame from the candidate phoneme set based on the association probability between each candidate phoneme and the i-th audio frame respectively.
[0196] In some embodiments, if the phoneme associated with the (i - 1)-th audio frame is any one of the original phonemes in the original phonemes of the reference character pronunciation except the last original phoneme, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the next phoneme of the phoneme associated with the (i - 1)-th audio frame in the phonemes of the reference character pronunciation, or, the default phoneme;
[0197] If the phoneme associated with the (i - 1)-th audio frame is the last original phoneme of the reference character pronunciation and the last original phoneme does not have an extended phoneme, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the first phoneme, or, the default phoneme;
[0198] If the phoneme associated with the (i - 1)-th audio frame is the sustained phoneme of the reference character pronunciation or any one of the extended phonemes of the sustained phoneme except the last extended phoneme, the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the next phoneme of the phoneme associated with the (i - 1)-th audio frame in the phonemes of the reference character pronunciation, the first phoneme, or, the default phoneme;
[0199] If the phoneme associated with the (i - 1)-th audio frame is the last extended phoneme of the pronunciation of the reference character, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the first phoneme, or the default phoneme.
[0200] In some embodiments, any one of the original phonemes of the pronunciation of a character in the lyric text is represented as a target original phoneme; when the processor 1201 is used to determine the target audio segment corresponding to each original phoneme of the character in the lyric text according to the audio segments corresponding to each phoneme respectively, it is specifically used for:
[0201] If the target original phoneme does not have an extended phoneme, then determine the audio segment corresponding to the target original phoneme as the target audio segment corresponding to the target original phoneme;
[0202] If the target original phoneme has an extended phoneme and the extended phoneme of the target original phoneme corresponds to an audio segment, then splice the audio segment corresponding to the target original phoneme and the audio segment corresponding to the extended phoneme of the target original phoneme, and determine the spliced audio segment as the target audio segment corresponding to the target original phoneme.
[0203] In some embodiments, when the processor 1201 is used to generate the musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text, it is specifically used for:
[0204] Obtain the pitch curve of the target audio segment corresponding to each original phoneme;
[0205] Determine at least one note associated with each original phoneme respectively according to the pitch curve corresponding to each original phoneme;
[0206] Generate a musical score according to at least one note associated with each original phoneme and the lyric text.
[0207] In some embodiments, any one of the original phonemes of the pronunciation of a character in the lyric text is represented as a target original phoneme; when the processor 1201 is used to determine at least one note associated with each original phoneme respectively according to the pitch curve corresponding to each original phoneme, it is specifically used for:
[0208] If the pitch curve corresponding to the target original phoneme belongs to a vibrato curve, then determine the note associated with the target original phoneme based on the average pitch indicated by the pitch curve corresponding to the target original phoneme;
[0209] If the pitch curve corresponding to the target original phoneme does not belong to a vibrato curve, then detect the pitch mutation points in the pitch curve corresponding to the target original phoneme, and determine at least one note associated with the target original phoneme based on the detection result.
[0210] In some embodiments, when the processor 1201 is used to determine at least one note associated with the target original phoneme based on the detection result, it is specifically used for:
[0211] If the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is no pitch mode, then determine the note associated with the target original phoneme according to the ending pitch of the pitch curve corresponding to the target original phoneme;
[0212] If the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is a pitch mode, then determine the note associated with the target original phoneme according to the pitch mode;
[0213] If the detection result includes a pitch mutation point detected in the pitch curve corresponding to the target original phoneme, then divide the pitch curve corresponding to the target original phoneme based on the detected pitch mutation point to obtain multiple pitch curve segments of the pitch curve corresponding to the target original phoneme, and determine the note associated with the target original phoneme according to each pitch curve segment respectively.
[0214] In the embodiments of the present application, the processor can obtain the singing audio, obtain the lyric text associated with the singing audio, and respectively obtain the phonemes of each character pronunciation in the lyric text. The phonemes of the character pronunciation may include a sustained phoneme and an extended phoneme representing the same phoneme after the pronunciation change of the sustained phoneme. Predict the audio segment corresponding to each phoneme of the character pronunciation in the lyric text in the singing audio, determine the target audio segment corresponding to each original phoneme of the character in the lyric text according to the audio segment corresponding to each phoneme respectively, and generate the score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text. In the above solution, when using a sustained phoneme to represent the pronunciation process of a phoneme, if the pronunciation of the phoneme changes, the extended phoneme can be used to further represent the pronunciation process of the same phoneme after the pronunciation change, so that a more complete pronunciation process of a phoneme can be represented by the sustained phoneme and the extended phoneme, and the target audio segment corresponding to each original phoneme of each character pronunciation can be determined more accurately, so as to achieve accurate alignment at the phoneme level between the lyric text and the singing audio, and generate a score based on the alignment result, which helps to improve the efficiency and accuracy of score generation.
[0215] In specific implementation, the above-described device, processor, memory, etc. can execute the implementation manners described in the method embodiments above, and can also execute the implementation manners described in the embodiments of the present application, which will not be elaborated here.
[0216] In an embodiment of the present application, a computer (readable) storage medium is further provided. The computer storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor can execute some or all of the steps performed in the above method embodiment. Optionally, the computer storage medium can be volatile or non-volatile. The computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the blockchain node, etc.
[0217] An embodiment of the present application further provides a computer program product. The computer program product includes computer instructions (program instructions). When the computer instructions are executed by a processor, some or all of the steps in the above audio data processing method can be implemented. Optionally, the computer instructions can be stored in a computer-readable storage medium. A processor of a computer device, such as an electronic device, reads the program instructions from the computer-readable storage medium, and the processor executes the program instructions, so that the computer device executes the above-provided audio data processing method.
[0218] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0219] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the present application.
[0220] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.
[0221] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An audio data processing method, characterized in that, The method includes: Obtaining a singing audio and obtaining the lyric text associated with the singing audio; Respectively obtaining the phonemes of the pronunciation of each character in the lyric text; the lyric text includes a target character, and the phonemes of the pronunciation of the target character include at least one original phoneme of the target character and an extended phoneme of a trailing phoneme, where the trailing phoneme refers to the last original phoneme in the original phonemes of the target character, and the extended phoneme of the trailing phoneme is used to represent the same phoneme after the pronunciation of the trailing phoneme changes; Predicting the audio segment corresponding to each phoneme of the pronunciation of a character in the lyric text in the singing audio respectively; any one original phoneme corresponds to an audio segment, and any one extended phoneme corresponds to an audio segment or does not correspond to an audio segment; Determining the target audio segment corresponding to each original phoneme of the character in the lyric text according to the audio segment corresponding to each phoneme respectively, and generating a musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text; Wherein, the singing audio is divided into multiple audio frames, and each audio frame is represented as a corresponding audio sequence, and any one of the phonemes of the pronunciation of the character in the lyric text is a target phoneme, and the audio segment corresponding to the target phoneme is obtained through the phonemes associated with the multiple audio frames respectively; the multiple audio frames include the i-th audio frame, i is an integer and i>1, and the character to which the phoneme associated with the i-th audio frame belongs is a reference character; the process of predicting the phoneme associated with each audio frame respectively includes: Obtaining the first phoneme in the phonemes of the pronunciation of the next character of the reference character in the lyric text; the next character of the reference character is determined based on the character order; determining a candidate phoneme set associated with the i-th audio frame based on the first phoneme, the phoneme associated with the (i - 1)-th audio frame, and the pronunciation order corresponding to the phonemes of the pronunciation of the reference character; the candidate phoneme set includes at least one candidate phoneme; predicting the association probability between each candidate phoneme in the candidate phoneme set and the i-th audio frame based on the audio sequence of the i-th audio frame; selecting the phoneme associated with the i-th audio frame from the candidate phoneme set based on the association probability between each candidate phoneme and the i-th audio frame respectively.
2. The method according to claim 1, wherein The predicting the audio segment corresponding to each phoneme of the pronunciation of a character in the lyric text in the singing audio respectively includes: If the phonemes associated with M consecutive audio frames in the multiple audio frames are all the target phonemes, then performing splicing processing on the M consecutive audio frames, and determining the spliced audio frame as the audio segment corresponding to the target phoneme.
3. The method according to claim 1, characterized in that If the phoneme associated with the (i - 1)-th audio frame is any one of the original phonemes of the pronunciation of the reference character except the last original phoneme, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i - 1)-th audio frame, the next phoneme of the phoneme associated with the (i - 1)-th audio frame in the phonemes of the pronunciation of the reference character, or, a default phoneme; If the phoneme associated with the (i-1)-th audio frame is the last original phoneme in the pronunciation of the reference character, and the last original phoneme does not have an extended phoneme, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i-1)-th audio frame, the first phoneme, or the default phoneme; If the phoneme associated with the (i-1)-th audio frame is the sustained phoneme in the pronunciation of the reference character, or any extended phoneme other than the last extended phoneme in the extended phonemes of the sustained phoneme, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i-1)-th audio frame, the next phoneme of the phoneme associated with the (i-1)-th audio frame among the phonemes in the pronunciation of the reference character, the first phoneme, or the default phoneme; If the phoneme associated with the (i-1)-th audio frame is the last extended phoneme in the pronunciation of the reference character, then the candidate phonemes in the candidate phoneme set include at least one of the following: the phoneme associated with the (i-1)-th audio frame, the first phoneme, or the default phoneme.
4. The method according to claim 1, characterized in that Any one of the original phonemes in the pronunciation of the characters in the lyric text is denoted as the target original phoneme; determining the target audio segment corresponding to each original phoneme of the characters in the lyric text according to the audio segments corresponding to each phoneme respectively includes: If the target original phoneme does not have an extended phoneme, determining the audio segment corresponding to the target original phoneme as the target audio segment corresponding to the target original phoneme; If the target original phoneme has an extended phoneme and the extended phoneme of the target original phoneme corresponds to an audio segment, splicing the audio segment corresponding to the target original phoneme and the audio segment corresponding to the extended phoneme of the target original phoneme, and determining the spliced audio segment as the target audio segment corresponding to the target original phoneme.
5. The method according to claim 1, characterized in that, Generating the musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text includes: Obtaining the pitch curve of the target audio segment corresponding to each original phoneme; Determining at least one note associated with each original phoneme according to the pitch curve corresponding to each original phoneme respectively; Generating the musical score according to at least one note associated with each original phoneme and the lyric text.
6. The method according to claim 5, wherein Any one of the original phonemes in the pronunciation of the characters in the lyric text is denoted as the target original phoneme; determining at least one note associated with each original phoneme according to the pitch curve corresponding to each original phoneme respectively includes: If the pitch curve corresponding to the target original phoneme belongs to a vibrato curve, determining the note associated with the target original phoneme based on the average pitch indicated by the pitch curve corresponding to the target original phoneme; If the pitch curve corresponding to the target original phoneme does not belong to a vibrato curve, detecting the pitch mutation points in the pitch curve corresponding to the target original phoneme, and determining at least one note associated with the target original phoneme based on the detection result.
7. The method according to claim 6, wherein Determining at least one note associated with the target original phoneme based on the detection result includes: If the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is no pitch mode, then determine the note associated with the target original phoneme according to the ending pitch of the pitch curve corresponding to the target original phoneme; If the detection result is used to indicate that there is no pitch mutation point in the pitch curve corresponding to the target original phoneme and there is a pitch mode, then determine the note associated with the target original phoneme according to the pitch mode; If the detection result includes a pitch mutation point detected in the pitch curve corresponding to the target original phoneme, then divide the pitch curve corresponding to the target original phoneme based on the detected pitch mutation point to obtain multiple pitch curve segments of the pitch curve corresponding to the target original phoneme, and determine the notes associated with the target original phoneme according to each pitch curve segment respectively.
8. An audio data processing device, characterized in that, The device includes: An acquisition module, configured to acquire a singing audio and acquire the lyric text associated with the singing audio; The acquisition module is further configured to respectively acquire the phonemes of the pronunciation of each character in the lyric text; the lyric text includes a target character, and the phonemes of the pronunciation of the target character include at least one original phoneme of the target character and an extended phoneme of a trailing phoneme, where the trailing phoneme refers to the last original phoneme in the original phonemes of the target character, and the extended phoneme of the trailing phoneme is used to represent the same phoneme after the pronunciation of the trailing phoneme changes; A processing module, configured to predict the audio segment corresponding to each phoneme of the pronunciation of a character in the lyric text in the singing audio; any one original phoneme corresponds to one audio segment, and any one extended phoneme corresponds to one audio segment or does not correspond to an audio segment; The processing module is further configured to determine the target audio segment corresponding to each original phoneme of the character in the lyric text according to the audio segment corresponding to each phoneme respectively, and generate a musical score of the singing audio according to the target audio segment corresponding to each original phoneme and the lyric text; Wherein, the singing audio is divided into multiple audio frames, and each audio frame is represented as a corresponding audio sequence, and any one of the phonemes of the pronunciation of a character in the lyric text is a target phoneme, and the audio segment corresponding to the target phoneme is obtained through the phonemes respectively associated with the multiple audio frames; the multiple audio frames include the i-th audio frame, i is an integer and i>1, and the character to which the phoneme associated with the i-th audio frame belongs is a reference character; the process by which the processing module predicts the phoneme respectively associated with each audio frame includes: Obtain the first phoneme in the phonemes of the pronunciation of the next character of the reference character in the lyric text; the next character of the reference character is determined based on the character order; determine the candidate phoneme set associated with the i-th audio frame based on the first phoneme, the phonemes associated with the (i - 1)-th audio frame, and the pronunciation order corresponding to the phonemes of the pronunciation of the reference character; the candidate phoneme set includes at least one candidate phoneme; predict the association probability between each candidate phoneme in the candidate phoneme set and the i-th audio frame based on the audio sequence of the i-th audio frame; select the phoneme associated with the i-th audio frame from the candidate phoneme set based on the association probability between each candidate phoneme and the i-th audio frame respectively.
9. An electronic device, characterized in that, Comprising a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 - 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 - 7.
11. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 - 7 is implemented.
Citation Information
Patent Citations
Method for determining lyric timestamp information and training method of acoustic model
CN112735429A
Note pitch value determination method and device, equipment and storage medium
CN113140230A
Music score generation method, electronic equipment and readable storage medium
CN113763913A