Speech processing model training method, speech processing method and song processing method

By training a speech processing model and using character and note processing units to process speech, the problem of inaccurate correspondence between characters and notes in speech is solved, thus improving the accuracy of speech processing tasks.

CN121662028APending Publication Date: 2026-03-13ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies cannot accurately determine the correspondence between characters and notes in speech, resulting in inaccurate results for speech processing tasks.

Method used

By training a speech processing model, the sample speech is processed using character processing units and note processing units to predict characters and notes. The model is then trained using training labels to achieve a precise correspondence between characters and notes.

Benefits of technology

It achieves accurate correspondence between characters and notes in speech, improving the accuracy of speech processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121662028A_ABST
    Figure CN121662028A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice processing model training method, a voice processing method and a song processing method.The voice processing model training method comprises the steps that a to-be-trained voice processing model is determined, and a training sample and a training label associated with the to-be-trained voice processing model are determined, the speech processing model to be trained comprises a character processing unit and a note processing unit; according to a character processing unit, processing the sample voice to obtain a prediction character corresponding to the character audio, and a prediction character audio duration and a prediction note number corresponding to the prediction character; according to a note processing unit, processing the sample voice, the predicted character audio duration and the predicted note quantity to obtain predicted notes corresponding to the predicted characters; and training a to-be-trained voice processing model according to the training label, the predicted character, the predicted character audio duration, the predicted note number and the predicted notes to obtain a trained voice processing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a method for training a speech processing model. One or more embodiments of this specification also relate to a speech processing method, a song processing method, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the continuous development of computer technology and voice technology, voice has become an indispensable part of people's daily lives. People can perform voice processing tasks such as voice synthesis and artistic creation according to their actual needs.

[0003] Currently, in the process of performing speech processing tasks, the correspondence between characters and notes in speech cannot be accurately determined. As a result, the characters and their corresponding notes cannot be accurately processed, leading to inaccurate task results. Therefore, how to determine the notes corresponding to characters in speech has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a speech processing model training method. One or more embodiments of this specification also relate to a speech processing method, a song processing method, a speech processing model training device, a speech processing device, a song processing device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies in the prior art that cannot accurately determine the correspondence between characters and notes in speech.

[0005] According to a first aspect of the embodiments of this specification, a method for training a speech processing model is provided, comprising:

[0006] The speech processing model to be trained is determined, as well as the training samples and training labels associated with the speech processing model to be trained. The speech processing model to be trained includes a character processing unit and a note processing unit. The training samples are sample speech containing character audio. The training labels are the target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes.

[0007] The character processing unit processes the sample speech to obtain the predicted character corresponding to the character audio, as well as the predicted character audio duration and the number of predicted notes.

[0008] According to the note processing unit, the sample speech, the audio duration of the predicted character, and the number of predicted notes are processed to obtain the predicted note corresponding to the predicted character;

[0009] The speech processing model to be trained is trained based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes to obtain the trained speech processing model.

[0010] According to a second aspect of the embodiments of this specification, a speech processing model training apparatus is provided, comprising:

[0011] The data determination module is configured to determine the speech processing model to be trained, and to determine the training samples and training labels associated with the speech processing model to be trained. The speech processing model to be trained includes a character processing unit and a note processing unit. The training samples are sample speech containing character audio. The training labels are the target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes.

[0012] The character processing module is configured to process the sample speech according to the character processing unit to obtain the predicted character corresponding to the character audio, as well as the predicted character audio duration and the number of predicted notes corresponding to the predicted character.

[0013] The note processing module is configured to process the sample speech, the audio duration of the predicted character, and the number of predicted notes according to the note processing unit to obtain the predicted note corresponding to the predicted character;

[0014] The model training module is configured to train the speech processing model to be trained based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes, so as to obtain the trained speech processing model.

[0015] According to a third aspect of the embodiments of this specification, a speech processing method is provided, comprising:

[0016] The speech to be processed is determined, wherein the speech to be processed contains character audio;

[0017] The speech to be processed is processed using a speech processing model to obtain the character corresponding to the character audio and the note corresponding to the character. The speech processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the speech to be processed to obtain the character corresponding to the character audio, as well as the duration of the character audio and the number of notes. The note processing unit is used to process the speech to be processed, the duration of the character audio, and the number of notes to obtain the note corresponding to the character.

[0018] According to a fourth aspect of the embodiments of this specification, a voice processing apparatus is provided, comprising:

[0019] The voice determination module is configured to determine the voice to be processed, wherein the voice to be processed contains character audio.

[0020] The note determination module is configured to process the speech to be processed using a speech processing model to obtain the character corresponding to the character audio and the note corresponding to the character. The speech processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the speech to be processed to obtain the character corresponding to the character audio, as well as the duration of the character audio and the number of notes. The note processing unit is used to process the speech to be processed, the duration of the character audio, and the number of notes to obtain the note corresponding to the character.

[0021] According to a fifth aspect of the embodiments of this specification, a song processing method is provided, comprising:

[0022] Identify the song to be processed, wherein the song to be processed contains character audio;

[0023] The song to be processed is processed using a song processing model to obtain the lyrics characters corresponding to the character audio and the notes corresponding to the lyrics characters. The song processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the song to be processed to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and the number of notes corresponding to the lyrics characters. The note processing unit is used to process the song to be processed, the character audio duration, and the number of notes to obtain the notes corresponding to the lyrics characters.

[0024] According to a sixth aspect of the embodiments of this specification, a song processing apparatus is provided, comprising:

[0025] The voice determination module is configured to determine the song to be processed, wherein the song to be processed contains character audio.

[0026] The note determination module is configured to process the song to be processed using a song processing model to obtain the lyrics characters corresponding to the character audio and the notes corresponding to the lyrics characters. The song processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the song to be processed to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and the number of notes corresponding to the lyrics characters. The note processing unit is used to process the song to be processed, the character audio duration, and the number of notes to obtain the notes corresponding to the lyrics characters.

[0027] According to a seventh aspect of the embodiments of this specification, a computing device is provided, comprising:

[0028] Memory and processor;

[0029] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of any of the above methods.

[0030] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of any of the above-described methods.

[0031] According to a ninth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0032] The speech processing model training method in one or more embodiments of this specification provides a speech processing model including a character processing unit and a note processing unit. During model training, the character processing unit processes sample speech to predict the predicted character corresponding to the character audio in the sample speech, as well as the predicted character audio duration and the number of predicted notes. The note processing unit processes the sample speech, the predicted character audio duration, and the number of predicted notes to obtain the predicted notes corresponding to the predicted characters. Then, based on the target character, the target note corresponding to the target character, the target character audio duration, and the number of target notes, as well as the predicted character, the predicted character audio duration, the number of predicted notes, and the predicted notes, the speech processing model to be trained is trained. This enables the speech processing model to perform more refined processing operations on speech using the character processing unit and the note processing unit, analyzing the correspondence between characters and notes, thereby obtaining a speech processing model that can accurately identify the notes corresponding to characters in speech. This achieves accurate determination of the correspondence between characters and notes in speech, facilitating accurate processing of characters and their corresponding notes during speech processing tasks, and improving the accuracy of task results. Attached Figure Description

[0033] Figure 1 This is a schematic diagram illustrating the application of a speech processing method provided in one embodiment of this specification;

[0034] Figure 2 This is a flowchart illustrating a speech processing model training method provided in one embodiment of this specification;

[0035] Figure 3 This is a flowchart illustrating the processing procedure of a speech processing model training method provided in one embodiment of this specification.

[0036] Figure 4 This is a flowchart illustrating a speech processing method provided in one embodiment of this specification;

[0037] Figure 5 This is a flowchart illustrating a song processing method provided in one embodiment of this specification;

[0038] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0039] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0040] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0041] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0042] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0043] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0044] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0045] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0046] Whisper: refers to the Whisper model, an automatic speech recognition (ASR) system designed to efficiently convert speech into text.

[0047] Cross-entropy is a metric used to measure the difference between two probability distributions.

[0048] MFA (Montreal Forced Aligner) is a tool used in the audio field to align audio and text, predicting the duration of speech for each word.

[0049] Notes: In the field of speech synthesis, each word corresponds to multiple pronunciations; the pitch and duration of each pronunciation combined together form a note.

[0050] UVR (Ultimate Vocal Remover): A model that can separate the accompaniment and vocals of a song.

[0051] Softmax (normalized exponential function): The Softmax function is a function that transforms a vector of real numbers into a probability distribution. It converts each element into a positive number through exponential operations and normalizes the output so that the sum of all outputs is 1. It is often used in multi-class classification tasks.

[0052] ROSVOT (Robust Singing Voice Transcription Serves Synthesis): A model of musical note transcription.

[0053] With the continuous development of computer technology, speech has become an indispensable part of people's daily lives. People can perform speech processing tasks such as speech synthesis and artistic creation according to their actual needs. For example, music is an essential part of people's daily lives, and with the continuous development of the field of artificial intelligence, many AI-based music creation products have achieved great success. For such music synthesis and music creation tasks, a large amount of vocal data with lyrics and note annotations is needed for training. For example, given a song, during the training process, the model needs to know the lyrics of each line, how many pitches each word in the lyrics needs to be sung, and the corresponding time for each pitch. Therefore, a large amount of processed vocal data is crucial to the field of vocal synthesis. However, due to copyright issues or data silos, there is not much available open-source vocal data with annotations; and existing music annotation tools cannot map lyrics to notes and have limited performance, requiring very time-consuming data preprocessing.

[0054] Based on this, a speech processing model training method is provided in this specification. One or more embodiments of this specification also relate to a speech processing method, a song processing method, a speech processing model training device, a speech processing device, a song processing device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0055] See Figure 1 , Figure 1 This diagram illustrates an application illustration of a speech processing method according to an embodiment of this specification, based on... Figure 1 It can be seen that users can upload songs to server 104 via terminal 102. After obtaining the song, server 104 can process it using a song transcription model. Specifically, the song is input into the autoregressive model in the song transcription model for prediction to obtain the lyrics corresponding to the lyrics audio, the duration of the lyrics audio, and the number of notes. Then, the song, the lyrics audio duration, and the number of notes are input into a non-autoregressive model for prediction to obtain the note corresponding to each lyric. Then, server 104 sends each lyric and its corresponding note to terminal 102, thereby achieving accurate annotation of the characters in the song corresponding to the notes.

[0056] See Figure 2 , Figure 2 A flowchart of a speech processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0057] Step 202: Determine the speech processing model to be trained, and determine the training samples and training labels associated with the speech processing model to be trained. The speech processing model to be trained includes a character processing unit and a note processing unit. The training samples are sample speech containing character audio. The training labels are the target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes.

[0058] The speech processing model can be used to process speech. The speech processing model training method provided in this specification will vary depending on the application scenario. For example, when the speech processing model training method provided in this specification is applied to a song processing scenario, the speech processing model can be a song processing model used to process songs; for example, the song processing model can be a song transcription model. When the speech processing model training method provided in this specification is applied to a recitation audio processing scenario, the speech processing model can be a recitation audio processing model used to process recitation audio. When the speech processing model training method provided in this specification is applied to a traditional opera processing scenario, the speech processing model can be a traditional opera processing model used to process traditional opera audio. The speech processing model can also be a large-scale model.

[0059] The speech processing model to be trained can be a speech processing model that needs to be trained.

[0060] Sample speech can be understood as speech data used as samples. For example, the sample speech can be data such as songs, recitation audio, opera audio, etc.; that is to say, the sample speech can be sample songs, sample recitation audio, sample opera audio, etc.

[0061] The sample speech can be audio, which can be understood as audio data obtained by recording human voice or simulated human voice. For example, the audio can be the human voice audio in a song (such as human voice singing audio), the human voice recitation audio in a recitation audio, etc.; the audio can be obtained by performing an accompaniment separation operation on the accompaniment audio in the speech.

[0062] The character audio contained in this sample speech can be understood as audio that expresses the meaning of specific characters. For example, character audio can be the audio of lyrics sung by a human voice in a song, or the audio of words recited by a human voice in a recitation audio; that is to say, this sample speech can be human voice audio.

[0063] The character can be the character information corresponding to the audio of the character, such as lyrics or recitation words. The target character can be understood as the character used as a training label, and the predicted character is the character predicted by the speech processing model.

[0064] The duration of a character audio can be understood as the length of time corresponding to a character audio. The duration of a character audio can be represented by the number of audio frames. In other words, the duration of a character audio can be the number of audio frames, and the audio frames can be the audio frames that constitute the sample speech.

[0065] A character processing unit can be understood as a unit used to process character audio in speech; the character processing unit can be one or more network layers in a speech processing model, or the character processing unit can be a sub-model in a speech processing model; for example, the character processing unit can be an autoregressive model in a speech processing model, used to process character audio in speech.

[0066] A note processing unit can be understood as a unit used to predict the note corresponding to a character; the note processing unit can be one or more network layers in a speech processing model, or the note processing unit can be a sub-model in a speech processing model; for example, the note processing unit can be a non-autoregressive model in a speech processing model, used to predict notes in speech.

[0067] The number of notes can be understood as the number of notes corresponding to a character. For example, the character A can correspond to one, two, or four notes. The number of target notes is the number of notes corresponding to the target character used as a training label. The number of predicted notes is the number of notes corresponding to the predicted character predicted by the speech processing model.

[0068] In one or more embodiments provided in this specification, in order to improve the performance of training samples, a two-stage screening operation can be performed on multiple speech samples to be screened to obtain the speech samples associated with the speech processing model to be trained (i.e., training samples). The two-stage screening operation includes a screening operation based on speech source information and a screening operation based on the similarity between speech associated characters and speech recognition characters. The specific implementation method is as follows.

[0069] Determining the training samples associated with the speech processing model to be trained includes:

[0070] The speech processing model to be trained is associated with multiple speech samples to be selected, and the speech source information and speech associated characters corresponding to each speech sample are determined.

[0071] Based on the voice source information, the multiple voices to be processed are filtered in the first stage to obtain the first stage voice;

[0072] Speech recognition is performed on the audio contained in the first stage of speech to obtain the speech recognition characters corresponding to the audio.

[0073] Based on the similarity between the speech-associated characters and the speech-recognition characters, the first-stage speech is subjected to a second-stage filtering to obtain the second-stage speech, which is then used as the training sample.

[0074] This method can determine the training samples and the sample labels as training data associated with the speech processing model to be trained.

[0075] Here, the speech to be filtered can be understood as the speech that needs to undergo two-stage filtering processing; the speech source information can be understood as information that represents the source of the speech to be filtered, for example, the speech source information can be the data source identifier of the speech to be filtered (such as a website address, ID, etc.); in the case that the speech to be filtered is a song to be filtered, the speech source information can be the singer information or creator information corresponding to the song to be filtered; the speech associated characters can be understood as characters associated with the speech to be filtered, and the speech associated characters correspond to the character audio in the speech to be filtered, for example, the speech associated characters can be the lyrics of a song, the recitation words of a recitation audio, the opera lyrics of an opera audio, etc.

[0076] The first-stage speech can be understood as the speech obtained by filtering multiple speech samples from the first-stage speech samples through the first-stage filtering operation; the second-stage speech is the speech obtained by filtering multiple speech samples from the first-stage speech samples through the second-stage filtering operation.

[0077] Taking the application of the speech processing model training method provided in this manual in the context of singing transcription as an example, the speech processing model training method is explained. The speech processing model can be a singing transcription model that combines an autoregressive model and a non-autoregressive model. The character processing unit is an autoregressive model, the note processing unit is a non-autoregressive model, and the sample speech is a sample song.

[0078] Based on this, this method can obtain multiple songs from the network or from a database; then, the accompaniment of these songs can be separated to obtain the human voice audio (i.e., the voice to be selected). The accompaniment separation operation can be performed using a UVR tool; subsequently, the separated human voice audio is used for annotation to improve the quality of the annotated notes.

[0079] After separating the accompaniment, the lyrics need to be transcribed. However, considering the significant noise in lyrics data (i.e., speech-related characters) obtained from the internet—for example, the beginning of the lyrics might contain singer information, and a sentence might be missing—and the poor quality of songs uploaded by non-professional musicians, this method first performs a preliminary screening of the song data based on singer information (speech source information) (first-stage screening), retaining the vocal audio obtained from separating the accompaniment of songs by mainstream pop singers (i.e., first-stage speech).

[0080] Then, speech recognition is performed on the human voice audio to obtain the corresponding human voice lyrics (i.e., speech recognition characters), and the lyrics are transcribed. The speech recognition can be performed using the Whisper model.

[0081] By comparing the similarity between the identified vocal lyrics and the obtained lyrics, and removing songs whose similarity is less than a preset similarity threshold (second-stage filtering operation), low-quality vocal audio is filtered out to obtain high-quality vocal audio (i.e., second-stage speech); and this high-quality vocal audio is used as sample speech.

[0082] It should be noted that this method uses the filtered songs (i.e., sample speech) to train a singing transcription model (i.e., speech processing model).

[0083] In one or more embodiments provided in this specification, in order to improve the accuracy of training labels, accurate note annotation can be performed based on sample speech to obtain accurate training labels. The specific implementation method is as follows:

[0084] Determining the training labels associated with the speech processing model to be trained includes:

[0085] The silent segments in the sample speech are identified, and the silent segments are used to segment the sample speech to obtain multiple sample speech segments, wherein the sample speech segments do not contain the silent segments;

[0086] Speech recognition is performed on each sample speech segment to obtain the target character corresponding to the character audio, and the duration of each sample speech segment is predicted to obtain the duration of the character audio corresponding to the target character;

[0087] From the plurality of sample speech segments, determine the preceding sample speech segment connected to the silence segment, and determine the target character corresponding to the character audio in the preceding sample speech segment as the associated character of the silence segment;

[0088] The audio duration of the segment corresponding to the silent segment is merged with the audio duration of the character corresponding to the associated character to obtain the audio duration of the target character.

[0089] Based on the audio duration of the target character, a musical note is marked for the target character to obtain the target note and the number of target notes corresponding to the target character;

[0090] The target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes are used as training labels.

[0091] Among them, silent segments can be understood as audio segments in the sample speech that do not contain human voices.

[0092] Among them, the associated characters of the silent segment can be understood as the characters of the previous sample speech segment (i.e. the preceding sample speech segment).

[0093] Following the previous example, the specific method for determining sample labels in this approach is as follows:

[0094] 1. In order to align lyrics and notes, this method needs to predict the duration of each word in the lyrics (i.e., the number of audio frames).

[0095] The prediction of the duration of each word can be achieved using a voice character alignment model, which can be understood as a model that aligns voices with characters in speech. For example, the voice character alignment model can be an MFA, which is a tool for aligning voices with characters. However, for singing data, the voice character alignment model will incorrectly predict the ending sounds in the voice audio as silence, which will be mixed with the silence segments that already exist in the singing and cannot be recognized. Based on this, this method can use the following approach for time prediction:

[0096] (1) Identify the silent segments in the human voice audio;

[0097] (2) The human voice audio is re-segmented using the silent segment to obtain multiple human voice segments (i.e., sample speech segments) that do not contain silent segments;

[0098] (3) Use the lyrics recognition model (i.e. Whisper) to re-identify the lyrics of the human voice segments, thereby labeling the lyrics corresponding to each human voice segment;

[0099] (4) Use MFA to process the vocal segments and predict the duration of each character in the lyrics of the vocal segments;

[0100] (5) The human voice segment that was incorrectly predicted as silent by the MFA is merged with the lyrics (i.e., associated characters) preceding the human voice segment to obtain the duration (i.e. the number of audio frames) of each word in the lyrics of the complete human voice audio.

[0101] 2. Mark the musical notes.

[0102] (1) Mark the notes on the optimized lyrics to determine the note corresponding to each word in the lyrics;

[0103] (2) For segments that are incorrectly predicted as silent, the characters in the preceding and following segments are re-marked according to their pitches, so as to obtain the note corresponding to each character in the lyrics of the complete human voice audio.

[0104] In one or more embodiments provided in this specification, in order to improve the accuracy and efficiency of speech processing, the complete speech (sample speech) can be segmented to obtain multiple speech segments, and the model can be trained based on the speech segments, so that the model can process the speech segments accurately and quickly. The specific implementation method is as follows.

[0105] The step of determining the training samples and training labels associated with the speech processing model to be trained includes:

[0106] The sample speech containing character audio is used as the sample speech to be segmented, and the training labels are used as the training labels to be segmented.

[0107] Multiple character sentences are determined from the target characters, and based on each character sentence, the speech sample to be segmented is segmented to obtain multiple speech segments.

[0108] Based on each sample speech segment, the training label to be segmented is segmented to obtain sample label segments corresponding to each sample speech segment. The sample label segments include the target segment character corresponding to the character audio in the sample speech segment, the target segment note corresponding to the target segment character, the target segment character audio duration, and the number of target segment notes.

[0109] The plurality of sample speech segments are used as training samples, and the sample label segments corresponding to each sample speech segment are used as sample labels.

[0110] Among them, character sentences can be understood as multiple sentences in speech recognition characters, such as a line of lyrics in the lyrics, or a sentence in the recitation audio.

[0111] Following the previous example, this scheme can segment sample songs based on lyrics and timestamps to obtain multiple song fragments. It can also determine the corresponding lyric fragments (i.e., the target characters in the sample labels) from the lyrics (i.e., the target characters in the sample speech fragments) to obtain multiple sentence-level lyric and song pairs. Then, the notes, character audio duration, and number of notes corresponding to the lyric fragments are used as sample labels, and the song fragments are used as training samples.

[0112] Step 204: Process the sample speech according to the character processing unit to obtain the predicted character corresponding to the character audio, as well as the predicted character audio duration and the number of predicted notes corresponding to the predicted character.

[0113] Here, the predicted character audio duration can be understood as the character audio duration predicted by the model; the predicted number of notes can be the number of notes predicted by the model.

[0114] Specifically, this scheme uses a character processing unit to perform prediction processing on sample speech, identify one or more characters in the sample speech, and determine the predicted character audio duration and the number of predicted notes for each character.

[0115] In one or more embodiments provided in this specification, the character processing unit includes a duration prediction subunit and a quantity prediction subunit. By utilizing the duration prediction subunit and the quantity prediction subunit, the sample speech is processed to accurately predict the predicted character corresponding to the character audio, as well as the predicted character audio duration and the predicted note number corresponding to the predicted character. The specific implementation method is as follows.

[0116] The step of processing the sample speech according to the character processing unit to obtain the predicted character corresponding to the character audio, as well as the predicted character audio duration and the number of predicted notes, includes:

[0117] The sample speech is input into the character processing unit, wherein the character processing unit includes a duration prediction subunit and a quantity prediction subunit, and the sample speech is a sample song;

[0118] The duration prediction subunit is used to perform speech recognition on the sample speech, predict the predicted character corresponding to the character audio, and the duration of the predicted character audio corresponding to the predicted character;

[0119] The quantity prediction subunit is used to identify the number of notes in the sample speech, predict the predicted character corresponding to the character audio, and the number of predicted notes corresponding to the predicted character.

[0120] The duration prediction subunit can be understood as a subunit in the character processing unit used to predict the duration of character audio. The duration prediction subunit can be one or more network layers in the character processing unit, or it can be a sub-model in the character processing unit. For example, the duration prediction subunit can be a Whisper model in the character processing unit.

[0121] The quantity prediction subunit can be understood as a subunit in the character processing unit used to predict the quantity of notes. This quantity prediction subunit can be one or more network layers in the character processing unit, or it can be a sub-model in the character processing unit. For example, the quantity prediction subunit can be another Whisper model in the character processing unit.

[0122] Continuing with the above example, the voice processing model provided in this specification can be a singing transcription model that combines an autoregressive model and a non-autoregressive model. Among them, the autoregressive model part can achieve end-to-end prediction by fine-tuning two Whisper models, obtaining two data pairs of [lyrics + lyrics duration] and [lyrics + number of notes], so as to achieve the prediction of lyrics and the alignment of lyrics and note sequences, which is convenient for subsequent training of the model. The specific implementation method is as follows:

[0123] First, the song transcription model (songtrans) takes human voice audio as input; processes the human voice audio by fine-tuning the first Whisper model to predict the lyrics and the corresponding duration of the lyrics (Word duration prediction); thus obtaining the data pair of [lyrics + lyrics duration].

[0124] Among them, the lyrics duration can be understood as the audio frames corresponding to a character in the lyrics. For example, Figure 3 each character in the lyrics "Looking up" in corresponds to 21, 18, 44, and 22 audio frames respectively.

[0125] Secondly, also taking human voice audio as input, processes the human voice audio by fine-tuning the second Whisper model to predict the lyrics and the corresponding number of notes of the lyrics (note number prediction); thus obtaining the data pair of [lyrics + number of notes].

[0126] Among them, the number of notes can be understood as the number of notes corresponding to a character in the lyrics. For example, Figure 3 each character in the lyrics "Looking up" in corresponds to 1, 1, 2, and 1 notes respectively.

[0127] Step 206: According to the note processing unit, process the sample voice, the predicted character audio duration, and the predicted number of notes to obtain the predicted notes corresponding to the predicted character.

[0128] In one or more embodiments provided in this specification, the note processing unit can process the sample voice, the predicted character audio duration, and the predicted number of notes through the note duration prediction subunit and the note frequency prediction subunit included in itself, so as to accurately obtain the predicted notes corresponding to the predicted character. The specific implementation method is as follows.

[0129] The process of processing the sample voice, the predicted character audio duration, and the predicted number of notes according to the note processing unit to obtain the predicted notes corresponding to the predicted character includes steps one to four:

[0130] Step 1: Input the sample speech, the predicted character audio duration, and the predicted number of notes into the note processing unit, wherein the note processing unit includes a note duration prediction subunit and a note frequency prediction subunit.

[0131] The note duration prediction subunit can be understood as a subunit in the note processing unit used for predicting note duration. The note duration prediction subunit can be one or more network layers in the note processing unit, or the note duration prediction subunit can be a sub-model in the note processing unit; for example, the note duration prediction subunit can be a subunit composed of an audio encoder and a linear layer.

[0132] The note frequency prediction subunit can be understood as a subunit in the note processing unit used for note frequency prediction. The note frequency prediction subunit can be one or more network layers in the note processing unit, or the note frequency prediction subunit can be a sub-model in the note processing unit; for example, the note frequency prediction subunit can be a subunit composed of an average pooling layer and a linear layer.

[0133] Step 2: Using the note duration prediction subunit, determine the audio frame features of multiple audio frames in the sample speech, and based on the audio frame features, the predicted character audio duration, and the number of predicted notes, determine the note audio duration corresponding to the predicted character.

[0134] In this context, an audio frame can be understood as an audio frame that constitutes the sample speech, and the features of an audio frame can be the feature code or feature vector corresponding to the audio frame.

[0135] The duration of a musical note audio can be understood as the duration of each musical note audio, such as the number of audio frames for each musical note audio.

[0136] Specifically, the step of using the note duration prediction subunit to determine the audio frame features of multiple audio frames in the sample speech, and determining the note audio duration corresponding to the predicted character based on the audio frame features, the predicted character audio duration, and the number of predicted notes, includes:

[0137] The audio encoder in the note duration prediction subunit is used to encode the multiple audio frames in the sample speech to obtain the audio frame features corresponding to each audio frame.

[0138] Using the linear layer in the note duration prediction subunit, prediction is performed based on the audio frame features to obtain the prediction score for each audio frame;

[0139] Based on the predicted character audio duration, determine the associated audio frame related to the predicted character from the plurality of audio frames;

[0140] When the number of predicted notes is one, the duration of the note audio corresponding to the predicted character is determined based on the associated audio frame;

[0141] When there are multiple predicted notes, boundary audio frames are determined from the associated audio frames based on the prediction scores. The associated audio frames are then segmented using the boundary audio frames to obtain multiple associated audio frame segments. Based on each associated audio frame segment, the note audio duration of the multiple predicted notes corresponding to the predicted character is determined.

[0142] Among them, the associated audio frame can be understood as an audio frame that is associated with the predicted character among multiple audio frames. The associated audio frame is the audio frame that constitutes the character audio corresponding to the predicted character.

[0143] A boundary audio frame is an audio frame that defines the boundaries between notes. In multiple related audio frames, it is necessary to use this boundary audio frame to separate the audio frame corresponding to each note.

[0144] Following the previous example, the speech processing model provided in this manual can be a singing transcription model that combines autoregressive and non-autoregressive models. The non-autoregressive model is divided into two parts: determining note duration and determining pitch. The note duration and pitch are used to determine the notes corresponding to the lyrics. The specific execution method is as follows:

[0145] The specific execution method for determining the duration of a note is as follows:

[0146] 1. Obtain the frame features of each audio frame of human voice audio through the encoder part of a Whisper model;

[0147] 2. For each character (i.e., the predicted character) and its corresponding frame features, a linear layer and a Softmax function are used to predict the score (i.e., the predicted score) for each frame.

[0148] 3. Based on the predicted duration of each character (i.e., the predicted audio duration of the character), determine the corresponding audio frame (i.e., the associated audio frame) for each character;

[0149] 4. Based on the number of predicted notes, select the frame with the highest score from the associated audio frames as the boundary frame, and obtain the duration of each note (i.e., the duration of the note audio).

[0150] For example, if a word corresponds to a certain number of musical notes, then the duration of the musical notes corresponding to that word (i.e., the number of audio frames) can be determined based on the duration of that word (i.e., the number of audio frames); see reference for details. Figure 3 , Figure 3In NoteDur: 21, 18, 20, which are the note durations of the three characters "raise", "head", and "look up" in the lyrics "Look up".

[0151] If a character corresponds to two or more notes, then it is necessary to determine the boundary frame (i.e., the boundary audio frame) from multiple audio frames. This method can select the frame with the highest score as the boundary frame; through this boundary frame, the audio frames corresponding to each note can be divided from multiple audio frames, thereby determining the duration corresponding to each note; specifically, reference can be made to Figure 3 , Figure 3 In the lyrics "Look up", the character "raise" corresponds to two notes, and Figure 3 In NoteDur: 19, 25, which are the note durations of the two notes corresponding to "raise".

[0152] Step 3: Use the note frequency prediction subunit to determine the note audio features of multiple note audios in the sample speech according to the note audio duration, and determine the note frequency corresponding to the predicted character according to the note audio features.

[0153] Among them, the note audio feature can be understood as the feature corresponding to each note audio; this note frequency can be understood as the high or low degree of the note audio. For example, this note frequency can be the pitch.

[0154] Specifically, the use of the note frequency prediction subunit to determine the note audio features of multiple note audios in the sample speech and determine the note frequency corresponding to the predicted character according to the note audio features includes:

[0155] Use the average pooling layer in the note frequency prediction subunit to perform note feature extraction on the audio frame features according to the note audio duration, and obtain the note audio features of the multiple note audios in the sample speech;

[0156] Use the linear layer in the note frequency prediction subunit to perform note frequency recognition according to each note audio feature, and determine the note frequency corresponding to the predicted character.

[0157] Continuing with the above example, determine the frame features corresponding to the predicted note duration, and obtain the feature (note feature) corresponding to each note through an average pooling layer; predict the pitch corresponding to each note through a linear layer and a Softmax function.

[0158] Step 4: Based on the note frequency and the note audio duration, determine the predicted note audio corresponding to the character audio from the multiple note audios.

[0159] Continuing with the previous example, since a note is composed of pitch and time, the specific method for determining pitch is as follows:

[0160] 1. From the frame features corresponding to multiple audio frames, determine the frame features corresponding to each predicted note duration, and process them through an average pooling layer to obtain the note feature corresponding to each note.

[0161] 2. For each note feature, the pitch corresponding to each note is predicted through a linear layer and a Softmax function.

[0162] Specifically, a linear layer and a Softmax function are used to predict multiple predicted pitches and scores for each note. The predicted pitch with the highest score is taken as the pitch of the note. Then, based on the pitch and the duration of the note audio, the note corresponding to each word can be accurately determined, thus achieving the annotation of the notes corresponding to the lyrics.

[0163] Step 208: Train the speech processing model to be trained based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes to obtain the trained speech processing model.

[0164] Specifically, this method can calculate two loss functions based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes, and use the two loss functions to adjust the model parameters of the speech processing model to be trained, thereby obtaining the trained speech processing model.

[0165] In one or more embodiments provided in this specification, training the speech processing model to be trained based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes to obtain the trained speech processing model includes:

[0166] A first loss function is calculated based on the target character, the audio duration of the target character, the number of target notes, the predicted character, the audio duration of the predicted character, and the number of predicted notes;

[0167] Calculate the second loss function based on the target note and the predicted note;

[0168] The first loss function is used to adjust the model parameters of the character processing unit, and the second loss function is used to adjust the model parameters of the note processing unit to obtain the trained speech processing model.

[0169] The first loss function can be the cross-entropy loss function, and the second loss function can be the binary cross-entropy loss function.

[0170] Following the previous example, in the model training process of this scheme, the first Whisper model is trained under supervision using the real lyrics and their duration obtained during the previous data annotation process. Furthermore, the second Whisper model is trained under supervision using the real lyrics and their corresponding note counts obtained during the previous annotation process.

[0171] Specifically, based on the lyrics and their corresponding durations in the sample labels, as well as the lyrics and their corresponding durations predicted by the first Whisper model, the cross-entropy loss function is calculated, and the first Whisper model is fine-tuned using the cross-entropy loss function for each character in the lyrics.

[0172] Furthermore, based on the lyrics and the number of notes corresponding to the lyrics in the sample labels, as well as the lyrics and the number of notes corresponding to the lyrics predicted by the second Whisper model, the cross-entropy loss function is calculated, and the second Whisper model is fine-tuned using the cross-entropy loss function of each character in the lyrics.

[0173] In training the non-autoregressive model, the model is supervised by using the musical notes corresponding to the real lyrics obtained during the previous data annotation process.

[0174] Specifically, based on the musical notes in the lyrics of the sample labels and the musical notes predicted by the non-autoregressive model, a binary cross-entropy loss function is calculated, and the non-autoregressive model is fine-tuned using the binary cross-entropy loss function of each character in the lyrics.

[0175] The speech processing model training method in one or more embodiments of this specification provides a speech processing model including a character processing unit and a note processing unit. During model training, the character processing unit processes sample speech to predict the predicted character corresponding to the character audio in the sample speech, as well as the predicted character audio duration and the number of predicted notes. The note processing unit processes the sample speech, the predicted character audio duration, and the number of predicted notes to obtain the predicted notes corresponding to the predicted characters. Then, based on the target character, the target note corresponding to the target character, the target character audio duration, and the number of target notes, as well as the predicted character, the predicted character audio duration, the number of predicted notes, and the predicted notes, the speech processing model to be trained is trained. This enables the speech processing model to perform more refined processing operations on speech using the character processing unit and the note processing unit, analyzing the correspondence between characters and notes, thereby obtaining a speech processing model that can accurately identify the notes corresponding to characters in speech. This achieves accurate determination of the correspondence between characters and notes in speech, facilitating accurate processing of characters and their corresponding notes during speech processing tasks, and improving the accuracy of task results.

[0176] The following is in conjunction with the appendix Figure 3 Taking the application of the speech processing model training method provided in this specification in singing transcription as an example, the speech processing model training method will be further explained. Among other things, Figure 3 The flowchart of a speech processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0177] Step 302: Determine the training sample and the corresponding sample label.

[0178] The specific method for determining the training samples is as follows:

[0179] 1. Several songs were obtained from the Internet, and the accompaniment of these songs was separated using UVR tools to obtain the vocal audio.

[0180] It should be noted that in the subsequent labeling process for the sample labels, the separated human voice audio was used for labeling, thereby improving the quality of the labeling.

[0181] 2. The songs were initially screened based on the singer information, retaining the vocal audio of mainstream pop singers.

[0182] Specifically, considering the significant amount of noise present in lyrics data obtained from the internet; for example, the lyrics may begin with the singer's information, or a sentence may be missing in the middle; furthermore, some songs uploaded by non-professional musicians are of poor quality.

[0183] Based on this, this method first performs preliminary screening of song data according to singer information, retaining the vocal audio obtained by separating the accompaniment of songs by mainstream pop singers.

[0184] 3. Use the Whisper model to perform speech recognition on human voice audio and obtain the corresponding lyrics.

[0185] 4. Compare the similarity between the recognized lyrics output by the Whisper model and the acquired lyrics, filter out the vocal audio corresponding to the songs with lower quality, and retain the high-quality vocal audio corresponding to the songs with higher quality.

[0186] 5. Use high-quality human voice audio and the corresponding lyrics as training samples.

[0187] The specific method for determining sample labels is as follows:

[0188] 1. Predict the duration of lyrics.

[0189] Specifically, in order to align lyrics and notes, this method needs to predict the duration of each word in the lyrics (i.e., the number of audio frames).

[0190] The prediction of the duration of each word can be achieved through MFA; however, for singing data, MFA will incorrectly predict the ending notes in the human voice audio as silence, which will be mixed with the silence segments that already exist in the singing and cannot be identified.

[0191] Based on this, the method can be used for time prediction in the following way:

[0192] (1) Identify the silent segments in the human voice audio;

[0193] (2) The human voice audio is re-segmented using the silent segment to obtain multiple human voice segments (i.e., human voice audio segments) that do not contain the silent segment;

[0194] (3) Use the lyrics recognition model (i.e. Whisper) to re-identify the lyrics of the human voice segments, thereby labeling the lyrics corresponding to each human voice segment;

[0195] (4) Use MFA to process the vocal segments and predict the duration of each character in the lyrics of the vocal segments;

[0196] (5) The human voice segment that was incorrectly predicted as silent by the MFA is merged with the lyrics preceding the human voice segment to obtain the duration (i.e. the number of audio frames) of each word in the lyrics of the complete human voice audio.

[0197] 2. Mark the musical notes.

[0198] (1) Use the ROSVOT tool to annotate the notes for the song based on the optimized MFA results, so as to determine the notes corresponding to each word in the lyrics.

[0199] (2) For the segments mispredicted as silence, re-annotate them according to the pitch of the previous and subsequent segments, so as to obtain the notes corresponding to each word in the lyrics of the complete human voice audio.

[0200] 3. Take the determined lyrics, the durations corresponding to the lyrics, and the notes corresponding to the lyrics as sample labels.

[0201] It should be noted that during the process of using the training samples (i.e., human voice audio) and sample labels (i.e., lyrics, the durations corresponding to the lyrics, and the notes corresponding to the lyrics) for model training, they can be segmented into multiple sentence-level "lyric and song pairs" according to the lyrics and timestamps; and then use this "lyric and song pair" for model training.

[0202] Among them, the song in the "lyric and song pair" refers to the human voice segment, and the lyrics refer to the lyric segment corresponding to this human voice segment, the duration corresponding to the lyric segment, and the notes corresponding to the lyric segment.

[0203] Step 304: Input the training samples into the autoregressive model in the song transcription model to make predictions, and obtain the predicted lyric duration and the predicted number of notes corresponding to the lyrics.

[0204] Specifically, for the autoregressive model part, this solution configures two Whisper models; the specific execution method is as follows:

[0205] First, the song transcription model (songtrans) takes the human voice audio as input; processes this human voice audio by fine-tuning the first Whisper model to predict the lyrics and the durations corresponding to the lyrics (Word duration prediction); thus obtaining the data pair of [lyrics + lyric duration].

[0206] Among them, the lyric duration can be understood as the audio frames corresponding to a character in the lyrics. For example, Figure 3 for each character in the lyrics "Looking up" in

[0207] corresponds to 21, 18, 44, and 22 audio frames respectively.

[0208] Among them, the number of notes can be understood as the number of notes corresponding to a character in the lyrics. For example, Figure 3 For each character in the lyrics "Looking up" in Figure 3 , the corresponding numbers of notes are 1, 1, 2, and 1 respectively.

[0209] Based on the above steps, it can be seen that this method fine-tunes two Whisper models to achieve end-to-end prediction of two data pairs, [lyrics + lyrics duration] and [lyrics + number of notes], so as to achieve the prediction of lyrics and the alignment of lyrics and note sequences.

[0210] It should be noted that in the training of the autoregressive model, the first Whisper model is supervised and trained with the true lyrics and true lyrics duration obtained in the previous data annotation process. And the second Whisper model is supervised and trained with the true lyrics and the number of notes corresponding to the true lyrics obtained in the previous annotation process.

[0211] Specifically, according to the lyrics and the corresponding duration in the sample label, as well as the lyrics and the corresponding duration predicted by the first Whisper model, calculate the cross-entropy loss function, and use the cross-entropy loss function of each character in the lyrics to fine-tune the first Whisper model.

[0212] And, according to the lyrics and the corresponding number of notes in the sample label, as well as the lyrics and the corresponding number of notes predicted by the second Whisper model, calculate the cross-entropy loss function, and use the cross-entropy loss function of each character in the lyrics to fine-tune the second Whisper model.

[0213] Step 306: Input the training samples, the predicted lyrics duration and predicted number of notes corresponding to the lyrics obtained in step 304 into the non-autoregressive model for prediction to obtain the notes corresponding to each lyric.

[0214] Specifically, for the non-autoregressive model part, it is divided into two parts: determining the note duration and determining the pitch. Through the note duration and pitch, the notes corresponding to the lyrics can be determined; the specific implementation method is as follows:

[0215] The specific implementation method for determining the note duration is:

[0216] 1. Obtain the frame features of each audio frame of the human voice audio through the encoder part of a Whisper model;

[0217] 2. For the frame features corresponding to each character, predict the score corresponding to each frame through a linear layer and the Softmax function;

[0218] 3. Determine the audio frames corresponding to each word based on the duration (i.e., the number of audio frames) predicted for each word in step 304.

[0219] 4. Based on the predicted number of notes, select the frame with the highest score as the boundary frame and obtain the duration of each note.

[0220] For example, if a word corresponds to a certain number of notes, then based on the duration of this word (i.e., the number of audio frames), the duration of the notes corresponding to this word (i.e., the number of audio frames) can be determined; specifically, refer to Figure 3 , Figure 3 in NoteDur: 21, 18, 20, which are the note durations of the three characters "raise", "head", and "look up" in the lyrics "raise the head and look up" respectively.

[0221] If a word corresponds to two or more notes, then it is necessary to determine the boundary frame from multiple audio frames. This method can select the frame with the highest score as the boundary frame; through this boundary frame, the audio frames corresponding to each note can be divided from multiple audio frames, so as to determine the duration of each note; specifically, refer to Figure 3 , Figure 3 in the lyrics "raise the head and look up", "look up" corresponds to two notes, and Figure 3 in Note Dur: 19, 25, which are the note durations of the two notes corresponding to "look up" respectively.

[0222] The specific implementation method for determining the pitch is as follows:

[0223] 1. From the frame features corresponding to multiple audio frames, determine the frame features corresponding to the duration of each predicted note, and process them through an average pooling layer to obtain the note feature for each note.

[0224] 2. For each note feature, predict the pitch corresponding to each note through a linear layer and a Softmax function.

[0225] It should be noted that in the training of the non-autoregressive model, the non-autoregressive model is supervised and trained with the notes corresponding to the real lyrics obtained in the previous data annotation process.

[0226] Specifically, according to the notes in the lyrics of the sample label and the notes predicted by the non-autoregressive model, calculate the binary cross-entropy loss function, and use the binary cross-entropy loss function of each character in the lyrics to fine-tune the non-autoregressive model.

[0227] Based on the above steps, it can be seen that the speech processing model training method in this specification trains a singing transcription model that combines autoregressive and non-autoregressive methods. This model simplifies the tedious and complex process of lyrics and note annotation; it can achieve good results in lyrics and note transcription tasks and can align lyrics with notes.

[0228] This vocal transcription model is based on a unified music annotation model trained on the Whisper model. This model can directly annotate notes and lyrics, realizing automatic annotation of vocal data. The model achieves good performance and avoids cumbersome data preprocessing. At the overall solution level, this model is a model that realizes the alignment of lyrics and notes. From a methodological perspective, this model is based on Whisper for end-to-end data processing, while determining the notes corresponding to the lyrics.

[0229] Furthermore, to address the difficulties in aligning lyrics and notes and accurately obtain the lyrics-note pairs for a song, this method designs a data processing workflow based on optimized existing tools. The SongTrans transcription model simplifies the tedious and complex process of lyric and note annotation, and utilizes the correspondence information between lyrics and notes output by the SongTrans model to annotate the lyrics with notes. This scheme optimizes music annotation tools, designs a song data annotation workflow, and achieves large-scale music data annotation. This data processing workflow separates the accompaniment and segments the vocals using silent frames, and achieves the alignment of notes and lyrics.

[0230] See Figure 4 , Figure 4 A flowchart of a speech processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0231] Step 402: Determine the speech to be processed, wherein the speech to be processed contains character audio;

[0232] Step 404: Process the speech to be processed using a speech processing model to obtain the character corresponding to the character audio and the note corresponding to the character. The speech processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the speech to be processed to obtain the character corresponding to the character audio, as well as the duration of the character audio and the number of notes. The note processing unit is used to process the speech to be processed, the duration of the character audio, and the number of notes to obtain the note corresponding to the character.

[0233] It should be noted that this method can use the musical notes corresponding to the characters to perform target tasks, such as model training tasks, song scoring tasks, and song selection tasks.

[0234] The speech processing method in one or more embodiments of this specification can utilize a character processing unit in a speech processing model to process the speech to be processed, obtain the character corresponding to the character audio in the speech to be processed, the duration of the character audio corresponding to the character, and the number of notes, and utilize a note processing unit to process the speech to be processed, the character audio duration, and the number of notes to accurately obtain the note corresponding to the character; this achieves accurate determination of the correspondence between characters and notes in the speech to be processed, which facilitates accurate processing of characters and their corresponding notes during the execution of speech processing tasks, and improves the accuracy of task results.

[0235] See Figure 5 , Figure 5 A flowchart of a song processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0236] Step 502: Determine the song to be processed, wherein the song to be processed contains character audio;

[0237] Step 504: Process the song to be processed using a song processing model to obtain the lyrics characters corresponding to the character audio and the notes corresponding to the lyrics characters. The song processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the song to be processed to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and the number of notes corresponding to the lyrics characters. The note processing unit is used to process the song to be processed, the character audio duration, and the number of notes to obtain the notes corresponding to the lyrics characters.

[0238] In one or more embodiments provided in this specification, processing the song to be processed using a song processing model to obtain the lyrics characters corresponding to the character audio and the musical notes corresponding to the lyrics characters includes:

[0239] The song to be processed is input into the song processing model, and the character processing unit in the song processing model is used to process the song to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and number of notes corresponding to the lyrics characters.

[0240] Using the note processing unit in the song processing model, the song to be processed, the duration of the character audio, and the number of notes are processed to obtain the notes corresponding to the lyrics characters.

[0241] The song processing method in one or more embodiments of this specification, during the processing of the song to be processed, can utilize the character processing unit in the song processing model to process the song to be processed, obtain the lyric characters corresponding to the character audio in the song to be processed, the duration of the character audio corresponding to the lyric characters, and the number of notes, and utilize the note processing unit to process the song to be processed, the character audio duration, and the number of notes, to accurately obtain the notes corresponding to the lyric characters; this achieves accurate determination of the correspondence between the lyric characters and notes in the song to be processed, which facilitates accurate processing of the lyric characters and the notes corresponding to the lyric characters during the execution of the song processing task, and improves the accuracy of the task results.

[0242] Corresponding to the above method embodiments, this specification also provides an embodiment of a speech processing model training device, which includes:

[0243] The data determination module is configured to determine the speech processing model to be trained, and to determine the training samples and training labels associated with the speech processing model to be trained. The speech processing model to be trained includes a character processing unit and a note processing unit. The training samples are sample speech containing character audio. The training labels are the target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes.

[0244] The character processing module is configured to process the sample speech according to the character processing unit to obtain the predicted character corresponding to the character audio, as well as the predicted character audio duration and the number of predicted notes corresponding to the predicted character.

[0245] The note processing module is configured to process the sample speech, the audio duration of the predicted character, and the number of predicted notes according to the note processing unit to obtain the predicted note corresponding to the predicted character;

[0246] The model training module is configured to train the speech processing model to be trained based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes, so as to obtain the trained speech processing model.

[0247] Optionally, the character processing module is further configured to:

[0248] The sample speech is input into the character processing unit, wherein the character processing unit includes a duration prediction subunit and a quantity prediction subunit, and the sample speech is a sample song;

[0249] The duration prediction subunit is used to perform speech recognition on the sample speech, predict the predicted character corresponding to the character audio, and the duration of the predicted character audio corresponding to the predicted character;

[0250] The quantity prediction subunit is used to identify the number of notes in the sample speech, predict the predicted character corresponding to the character audio, and the number of predicted notes corresponding to the predicted character.

[0251] Optionally, the note processing module is further configured to:

[0252] The sample speech, the predicted character audio duration, and the predicted number of notes are input into the note processing unit, wherein the note processing unit includes a note duration prediction subunit and a note frequency prediction subunit.

[0253] Using the note duration prediction subunit, the audio frame features of multiple audio frames in the sample speech are determined, and based on the audio frame features, the audio duration of the predicted character, and the number of predicted notes, the audio duration of the note corresponding to the predicted character is determined;

[0254] Using the note frequency prediction subunit, the note audio features of multiple note audios in the sample speech are determined according to the note audio duration, and the note frequency corresponding to the predicted character is determined according to the note audio features.

[0255] Based on the note frequency and the note audio duration, the predicted note corresponding to the predicted character is determined.

[0256] Optionally, the note processing module is further configured to:

[0257] The audio encoder in the note duration prediction subunit is used to encode the multiple audio frames in the sample speech to obtain the audio frame features corresponding to each audio frame.

[0258] Using the linear layer in the note duration prediction subunit, prediction is performed based on the audio frame features to obtain the prediction score for each audio frame;

[0259] Based on the predicted character audio duration, determine the associated audio frame related to the predicted character from the plurality of audio frames;

[0260] When the number of predicted notes is one, the duration of the note audio corresponding to the predicted character is determined based on the associated audio frame;

[0261] When there are multiple predicted notes, boundary audio frames are determined from the associated audio frames based on the prediction scores. The associated audio frames are then segmented using the boundary audio frames to obtain multiple associated audio frame segments. Based on each associated audio frame segment, the note audio duration of the multiple predicted notes corresponding to the predicted character is determined.

[0262] Optionally, the note processing module is further configured to:

[0263] Using the average pooling layer in the note frequency prediction subunit, note features are extracted from the audio frame features based on the note audio duration to obtain the note audio features of the multiple note audios in the sample speech;

[0264] Using the linear layer in the note frequency prediction subunit, note frequency recognition is performed based on the audio features of each note to determine the note frequency corresponding to the predicted character.

[0265] Optionally, the model training module is further configured to:

[0266] A first loss function is calculated based on the target character, the audio duration of the target character, the number of target notes, the predicted character, the audio duration of the predicted character, and the number of predicted notes;

[0267] Calculate the second loss function based on the target note and the predicted note;

[0268] The first loss function is used to adjust the model parameters of the character processing unit, and the second loss function is used to adjust the model parameters of the note processing unit to obtain the trained speech processing model.

[0269] Optionally, the data determination module is further configured to:

[0270] The speech processing model to be trained is associated with multiple speech samples to be selected, and the speech source information and speech associated characters corresponding to each speech sample are determined.

[0271] Based on the voice source information, the multiple voices to be processed are filtered in the first stage to obtain the first stage voice;

[0272] Speech recognition is performed on the audio contained in the first stage of speech to obtain the speech recognition characters corresponding to the audio.

[0273] Based on the similarity between the speech-associated characters and the speech-recognition characters, the first-stage speech is subjected to a second-stage filtering to obtain the second-stage speech, which is then used as the training sample.

[0274] Optionally, the data determination module is further configured to:

[0275] The silent segments in the sample speech are identified, and the silent segments are used to segment the sample speech to obtain multiple sample speech segments, wherein the sample speech segments do not contain the silent segments;

[0276] Speech recognition is performed on each sample speech segment to obtain the target character corresponding to the character audio, and the duration of each sample speech segment is predicted to obtain the duration of the character audio corresponding to the target character;

[0277] From the plurality of sample speech segments, determine the preceding sample speech segment connected to the silence segment, and determine the target character corresponding to the character audio in the preceding sample speech segment as the associated character of the silence segment;

[0278] The audio duration of the segment corresponding to the silent segment is merged with the audio duration of the character corresponding to the associated character to obtain the audio duration of the target character.

[0279] Based on the audio duration of the target character, a musical note is marked for the target character to obtain the target note and the number of target notes corresponding to the target character;

[0280] The target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes are used as training labels.

[0281] Optionally, the data determination module is further configured to:

[0282] The sample speech containing character audio is used as the sample speech to be segmented, and the training labels are used as the training labels to be segmented.

[0283] Multiple character sentences are determined from the target characters, and based on each character sentence, the speech sample to be segmented is segmented to obtain multiple speech segments.

[0284] Based on each sample speech segment, the training label to be segmented is segmented to obtain sample label segments corresponding to each sample speech segment. The sample label segments include the target segment character corresponding to the character audio in the sample speech segment, the target segment note corresponding to the target segment character, the target segment character audio duration, and the number of target segment notes.

[0285] The plurality of sample speech segments are used as training samples, and the sample label segments corresponding to each sample speech segment are used as sample labels.

[0286] The speech processing model training apparatus in one or more embodiments of this specification provides a speech processing model including a character processing unit and a note processing unit. During model training, the character processing unit processes sample speech to predict the predicted character corresponding to the character audio in the sample speech, as well as the predicted character audio duration and the number of predicted notes. The note processing unit processes the sample speech, the predicted character audio duration, and the number of predicted notes to obtain the predicted note corresponding to the predicted character. Then, based on the target character, the target note corresponding to the target character, the target character audio duration, and the number of target notes, as well as the predicted character, the predicted character audio duration, the number of predicted notes, and the predicted note, the speech processing model to be trained is trained, thereby obtaining a speech processing model capable of accurately recognizing the notes corresponding to characters in speech. This achieves accurate determination of the correspondence between characters and notes in speech, facilitating accurate processing of characters and their corresponding notes during speech processing tasks, and improving the accuracy of task results.

[0287] The above is an illustrative scheme of a speech processing model training device according to this embodiment. It should be noted that the technical solution of this speech processing model training device and the technical solution of the speech processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the speech processing model training device, please refer to the description of the technical solution of the speech processing model training method described above.

[0288] Corresponding to the above method embodiments, this specification also provides embodiments of a voice processing device, which includes:

[0289] The voice determination module is configured to determine the voice to be processed, wherein the voice to be processed contains character audio.

[0290] The note determination module is configured to process the speech to be processed using a speech processing model to obtain the character corresponding to the character audio and the note corresponding to the character. The speech processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the speech to be processed to obtain the character corresponding to the character audio, as well as the duration of the character audio and the number of notes. The note processing unit is used to process the speech to be processed, the duration of the character audio, and the number of notes to obtain the note corresponding to the character.

[0291] The speech processing device in one or more embodiments of this specification, during the processing of a song to be processed, can utilize the character processing unit in the song processing model to process the song to be processed, obtain the lyric characters corresponding to the character audio in the song to be processed, the duration of the character audio corresponding to the lyric characters, and the number of notes, and utilize the note processing unit to process the song to be processed, the character audio duration, and the number of notes, accurately obtaining the notes corresponding to the lyric characters; thus, it achieves accurate determination of the correspondence between the lyric characters and notes in the song to be processed, facilitating accurate processing of the lyric characters and the notes corresponding to the lyric characters during the execution of the song processing task, and improving the accuracy of the task results.

[0292] The above is an illustrative scheme of a voice processing device according to this embodiment. It should be noted that the technical solution of this voice processing device and the technical solution of the above-described voice processing method belong to the same concept. For details not described in detail in the technical solution of the voice processing device, please refer to the description of the technical solution of the above-described voice processing method.

[0293] Corresponding to the above method embodiments, this specification also provides embodiments of a song processing apparatus, which includes:

[0294] The voice determination module is configured to determine the song to be processed, wherein the song to be processed contains character audio.

[0295] The note determination module is configured to process the song to be processed using a song processing model to obtain the lyrics characters corresponding to the character audio and the notes corresponding to the lyrics characters. The song processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the song to be processed to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and the number of notes corresponding to the lyrics characters. The note processing unit is used to process the song to be processed, the character audio duration, and the number of notes to obtain the notes corresponding to the lyrics characters.

[0296] Optionally, the note determination module is further configured to:

[0297] The song to be processed is input into the song processing model, and the character processing unit in the song processing model is used to process the song to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and number of notes corresponding to the lyrics characters.

[0298] Using the note processing unit in the song processing model, the song to be processed, the duration of the character audio, and the number of notes are processed to obtain the notes corresponding to the lyrics characters.

[0299] The song processing apparatus in one or more embodiments of this specification, during the processing of a song to be processed, can utilize the character processing unit in the song processing model to process the song to be processed, obtain the lyric characters corresponding to the character audio in the song to be processed, the duration of the character audio corresponding to the lyric characters, and the number of notes, and utilize the note processing unit to process the song to be processed, the character audio duration, and the number of notes, accurately obtaining the notes corresponding to the lyric characters; thus, it achieves accurate determination of the correspondence between the lyric characters and notes in the song to be processed, facilitating accurate processing of the lyric characters and the notes corresponding to the lyric characters during the execution of the song processing task, and improving the accuracy of the task results.

[0300] The above is an illustrative scheme of a song processing device according to this embodiment. It should be noted that the technical solution of this song processing device and the technical solution of the song processing method described above belong to the same concept. For details not described in detail in the technical solution of the song processing device, please refer to the description of the technical solution of the song processing method described above.

[0301] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0302] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0303] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0304] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.

[0305] The processor 620 is configured to execute computer-executable instructions that, when executed by the processor, implement the steps of any of the methods described above.

[0306] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to any one method embodiment, so the description is relatively simple; relevant parts can be referred to in the description of any one method embodiment.

[0307] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0308] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are relatively simple in description because they are substantially similar to any method embodiment; relevant parts can be referred to in the description of any method embodiment.

[0309] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0310] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solution of any of the above methods, and any details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of any of the above methods.

[0311] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0312] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0313] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0314] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0315] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A method for training a speech processing model, comprising: The speech processing model to be trained is determined, as well as the training samples and training labels associated with the speech processing model to be trained. The speech processing model to be trained includes a character processing unit and a note processing unit. The training samples are sample speech containing character audio. The training labels are the target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes. The character processing unit processes the sample speech to obtain the predicted character corresponding to the character audio, as well as the predicted character audio duration and the number of predicted notes. According to the note processing unit, the sample speech, the audio duration of the predicted character, and the number of predicted notes are processed to obtain the predicted note corresponding to the predicted character; The speech processing model to be trained is trained based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes to obtain the trained speech processing model.

2. The speech processing model training method according to claim 1, wherein processing the sample speech using the character processing unit to obtain the predicted character corresponding to the character audio, and the predicted character audio duration and predicted note number corresponding to the predicted character, comprises: The sample speech is input into the character processing unit, wherein the character processing unit includes a duration prediction subunit and a quantity prediction subunit; The duration prediction subunit is used to perform speech recognition on the sample speech, predict the predicted character corresponding to the character audio, and the duration of the predicted character audio corresponding to the predicted character; The quantity prediction subunit is used to identify the number of notes in the sample speech, predict the predicted character corresponding to the character audio, and the number of predicted notes corresponding to the predicted character.

3. The speech processing model training method according to claim 1, wherein processing the sample speech, the audio duration of the predicted character, and the number of predicted notes according to the note processing unit to obtain the predicted note corresponding to the predicted character includes: The sample speech, the predicted character audio duration, and the predicted number of notes are input into the note processing unit, wherein the note processing unit includes a note duration prediction subunit and a note frequency prediction subunit. Using the note duration prediction subunit, the audio frame features of multiple audio frames in the sample speech are determined, and based on the audio frame features, the audio duration of the predicted character, and the number of predicted notes, the audio duration of the note corresponding to the predicted character is determined; Using the note frequency prediction subunit, the note audio features of multiple note audios in the sample speech are determined according to the note audio duration, and the note frequency corresponding to the predicted character is determined according to the note audio features. Based on the note frequency and the note audio duration, the predicted note corresponding to the predicted character is determined.

4. The speech processing model training method according to claim 3, wherein the step of using the note duration prediction subunit to determine the audio frame features of multiple audio frames in the sample speech, and determining the note audio duration corresponding to the predicted character based on the audio frame features, the predicted character audio duration, and the number of predicted notes, comprises: The audio encoder in the note duration prediction subunit is used to encode the multiple audio frames in the sample speech to obtain the audio frame features corresponding to each audio frame. Using the linear layer in the note duration prediction subunit, prediction is performed based on the audio frame features to obtain the prediction score for each audio frame; Based on the predicted character audio duration, determine the associated audio frame related to the predicted character from the plurality of audio frames; When the number of predicted notes is one, the duration of the note audio corresponding to the predicted character is determined based on the associated audio frame; When there are multiple predicted notes, boundary audio frames are determined from the associated audio frames based on the prediction scores. The associated audio frames are then segmented using the boundary audio frames to obtain multiple associated audio frame segments. Based on each associated audio frame segment, the note audio duration of the multiple predicted notes corresponding to the predicted character is determined.

5. The speech processing model training method according to claim 3, wherein the step of using the note frequency prediction subunit to determine the note audio features of multiple note audios in the sample speech based on the note audio duration, and determining the note frequency corresponding to the predicted character based on the note audio features, comprises: Using the average pooling layer in the note frequency prediction subunit, note features are extracted from the audio frame features based on the note audio duration to obtain the note audio features of the multiple note audios in the sample speech; Using the linear layer in the note frequency prediction subunit, note frequency recognition is performed based on the audio features of each note to determine the note frequency corresponding to the predicted character.

6. The speech processing model training method according to claim 1, wherein training the speech processing model to be trained based on the training labels, the predicted characters, the audio duration of the predicted characters, the number of predicted notes, and the predicted notes to obtain the trained speech processing model includes: A first loss function is calculated based on the target character, the audio duration of the target character, the number of target notes, the predicted character, the audio duration of the predicted character, and the number of predicted notes; Calculate the second loss function based on the target note and the predicted note; The first loss function is used to adjust the model parameters of the character processing unit, and the second loss function is used to adjust the model parameters of the note processing unit to obtain the trained speech processing model.

7. The speech processing model training method according to claim 1, wherein determining the training samples associated with the speech processing model to be trained includes: The speech processing model to be trained is associated with multiple speech samples to be selected, and the speech source information and speech associated characters corresponding to each speech sample are determined. Based on the voice source information, the multiple voices to be processed are filtered in the first stage to obtain the first stage voice; Speech recognition is performed on the audio contained in the first stage of speech to obtain the speech recognition characters corresponding to the audio. Based on the similarity between the speech-associated characters and the speech-recognition characters, the first-stage speech is subjected to a second-stage filtering to obtain the second-stage speech, which is then used as the training sample.

8. The speech processing model training method according to claim 1, wherein determining the training labels associated with the speech processing model to be trained includes: The silent segments in the sample speech are identified, and the silent segments are used to segment the sample speech to obtain multiple sample speech segments, wherein the sample speech segments do not contain the silent segments; Speech recognition is performed on each sample speech segment to obtain the target character corresponding to the character audio, and the duration of each sample speech segment is predicted to obtain the duration of the character audio corresponding to the target character; From the plurality of sample speech segments, determine the preceding sample speech segment connected to the silence segment, and determine the target character corresponding to the character audio in the preceding sample speech segment as the associated character of the silence segment; The audio duration of the segment corresponding to the silent segment is merged with the audio duration of the character corresponding to the associated character to obtain the audio duration of the target character. Based on the audio duration of the target character, a musical note is marked for the target character to obtain the target note and the number of target notes corresponding to the target character; The target character corresponding to the character audio, the target note corresponding to the target character, the duration of the target character audio, and the number of target notes are used as training labels.

9. The speech processing model training method according to claim 1, wherein determining the training samples and training labels associated with the speech processing model to be trained includes: The sample speech containing character audio is used as the sample speech to be segmented, and the training labels are used as the training labels to be segmented. Multiple character sentences are determined from the target characters, and based on each character sentence, the speech sample to be segmented is segmented to obtain multiple speech segments. Based on each sample speech segment, the training label to be segmented is segmented to obtain sample label segments corresponding to each sample speech segment. The sample label segments include the target segment character corresponding to the character audio in the sample speech segment, the target segment note corresponding to the target segment character, the target segment character audio duration, and the number of target segment notes. The plurality of sample speech segments are used as training samples, and the sample label segments corresponding to each sample speech segment are used as sample labels.

10. A speech processing method, comprising: The speech to be processed is determined, wherein the speech to be processed contains character audio; The speech to be processed is processed using a speech processing model to obtain the character corresponding to the character audio and the note corresponding to the character. The speech processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the speech to be processed to obtain the character corresponding to the character audio, as well as the duration of the character audio and the number of notes. The note processing unit is used to process the speech to be processed, the duration of the character audio, and the number of notes to obtain the note corresponding to the character.

11. A song processing method, comprising: Identify the song to be processed, wherein the song to be processed contains character audio; The song to be processed is processed using a song processing model to obtain the lyrics characters corresponding to the character audio and the notes corresponding to the lyrics characters. The song processing model includes a character processing unit and a note processing unit. The character processing unit is used to process the song to be processed to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and the number of notes corresponding to the lyrics characters. The note processing unit is used to process the song to be processed, the character audio duration, and the number of notes to obtain the notes corresponding to the lyrics characters.

12. The song processing method according to claim 11, wherein processing the song to be processed using a song processing model to obtain the lyrics characters corresponding to the character audio and the musical notes corresponding to the lyrics characters includes: The song to be processed is input into the song processing model. The character processing unit in the song processing model is used to process the song to obtain the lyrics characters corresponding to the character audio, as well as the character audio duration and number of notes corresponding to the lyrics characters. Using the note processing unit in the song processing model, the song to be processed, the duration of the character audio, and the number of notes are processed to obtain the notes corresponding to the lyrics characters.

13. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.

14. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.

15. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.