Method and device for synthesizing singing sound, electronic equipment and program product

By adding ventilation marks to the music score file and dividing them into segments, the corresponding audio clips are generated to synthesize the singing voice, which solves the problems of holding breath and dragging the sound in traditional methods, and achieves a more natural singing synthesis effect.

CN120356453APending Publication Date: 2025-07-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410090755.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing singing vocal synthesis technology fails to fully consider the breathing sound in real singing, which leads to the phenomenon of holding breath or dragging the sound when synthesizing the singing, affecting the auditory experience.

Method used

By obtaining a score file with a breathing mark, dividing it into multiple score segments, and generating audio clips corresponding to these clips, and finally synthesize the singing based on the audio clips to simulate the breathing rhythm and tone of the live-action singing.

Benefits of technology

The synthesized singing is smoother and more natural, avoiding the phenomenon of holding breath and dragging the sound in traditional methods, and enhancing the audience's experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356453A_ABST
    Figure CN120356453A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for synthesizing singing, electronic equipment and a program product. The method comprises the steps of obtaining a music score file with a ventilation identifier, and segmenting the music score file into a plurality of music score segments based on the ventilation identifier; the method further includes generating a plurality of audio segments corresponding to the plurality of music score segments, and synthesizing a song corresponding to the music score file based on the plurality of audio segments. According to the embodiment of the invention, the music score file is segmented according to the ventilation identifier to obtain the plurality of music score segments, the corresponding audio segments are generated, and the plurality of audio segments are synthesized into the singing sound corresponding to the music score file, so that the phenomenon that the synthesized singing sound has unnatural suffocation or dragging sound and the like is avoided; the tone and intonation of a real person during singing are better restored, and the auditory experience of audiences is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of singing synthesis technology, and more specifically, to methods, apparatuses, electronic devices, and program products for synthesizing singing voices. Background Art

[0002] Singing synthesis technology is a technology that uses computer algorithms and sound processing techniques to generate human singing voices. It is based on the principle of audio signal processing and aims to create high-quality singing voices by simulating the voices and expressions of human singers.

[0003] Currently, singing synthesis technology has made great progress, can generate high-quality singing voices, and has been widely used in fields such as music production and virtual singers. Through singing synthesis technology, people can easily create singing voices that are very similar to real human voices, thus providing more possibilities for music production and creation. With the continuous progress of technology and the expansion of the application scope, singing synthesis technology is expected to play a greater role in the future. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, and program product for synthesizing singing voices.

[0005] According to a first aspect of the disclosure, a method for synthesizing singing voices is provided. The method includes obtaining a musical score file with breathing marks. The method further includes splitting the musical score file into multiple musical score segments based on the breathing marks. The method also includes generating multiple audio segments corresponding to the multiple musical score segments. In addition, the method further includes synthesizing a singing voice corresponding to the musical score file based on the multiple audio segments.

[0006] In a second aspect of the disclosure, an apparatus for synthesizing singing voices is provided. The apparatus includes a musical score file obtaining module configured to obtain a musical score file with breathing marks. The apparatus further includes a musical score file splitting module configured to split the musical score file into multiple musical score segments based on the breathing marks. The apparatus also includes an audio segment generating module configured to generate multiple audio segments corresponding to the multiple musical score segments. In addition, the apparatus further includes a singing voice synthesizing module configured to synthesize a singing voice corresponding to the musical score file based on the multiple audio segments.

[0007] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the electronic device to perform the method according to the first aspect.

[0008] In a fourth aspect of the present disclosure, a computer program product is provided. Computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions are executed by a processor to implement the method according to the first aspect.

[0009] The summary of the invention is to introduce the selection of concepts in a simplified form, which will be further described in the following detailed implementation. The summary of the invention is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings

[0010] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0011] Figure 1 A schematic diagram showing an example environment in which some embodiments of the present disclosure can be implemented;

[0012] Figure 2 A flowchart showing a method for synthesizing singing voices according to some embodiments of the present disclosure;

[0013] Figure 3 A schematic diagram showing a method for synthesizing singing voices according to some embodiments of the present disclosure;

[0014] Figure 4 A schematic diagram showing an alternative method for synthesizing singing voices according to some embodiments of the present disclosure;

[0015] Figure 5 A schematic diagram showing a method for training an inhalation mark prediction model according to some embodiments of the present disclosure;

[0016] Figure 6 A schematic diagram showing a method for segmenting music score fragments according to some embodiments of the present disclosure;

[0017] Figure 7 A schematic diagram showing a method for adding symbols to the head and tail of a music score fragment according to some embodiments of the present disclosure;

[0018] Figure 8 A block diagram showing a device for synthesizing singing voices according to some embodiments of the present disclosure; and

[0019] Figure 9 A block diagram showing an electronic device according to some embodiments of the present disclosure.

[0020] In all the drawings, the same or similar reference numerals denote the same or similar elements. Detailed Description of the Embodiments

[0021] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) shall comply with the requirements of corresponding laws, regulations and related provisions.

[0022] It is understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure shall be informed to the user and the user's authorization shall be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, when receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0024] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not constitute a limitation on the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0026] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not used to limit the protection scope of the present disclosure.

[0027] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects unless otherwise specified. There may also be other explicit and implicit definitions hereinafter.

[0028] In current singing synthesis technology, traditional methods usually rely on rests (such as long rests) extracted from a musical score file to segment the musical score, but do not fully consider the breathing sounds generated during actual singing. This results in some unnatural phenomena when synthesizing the singing voice. For example, there may be a drawn-out tone in some parts, while in other parts, there may be a breath-holding phenomenon, which significantly affects the overall auditory experience of the listener.

[0029] In an embodiment of the present disclosure, a musical score file with breathing marks is obtained, the musical score file is segmented into multiple musical score segments according to the breathing marks, multiple audio segments corresponding to the multiple musical score segments are generated, and finally, based on the multiple audio segments, the singing voice corresponding to the musical score file is synthesized. The musical score file is segmented according to the actual breathing rhythm of the singer, and the corresponding musical score segments are converted into corresponding audio segments to synthesize a complete singing voice segment, which can better simulate the tone and intonation during real singing. This method of synthesizing the singing voice can avoid the phenomena of drawn-out notes or breath-holding when using traditional methods to synthesize the singing voice, making the synthesized singing voice smoother and more natural, thereby improving the user experience.

[0030] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which some embodiments of the present disclosure can be implemented. As Figure 1 shown, the exemplary environment 100 can be set as a computing device 120, where the computing device 120 can be set as a computing system, a single server, a distributed server, or a cloud-based server, etc., can also be set as a user terminal, a mobile device, a computer, etc., and can also be set as a combination of the above devices.

[0031] Refer to Figure 1 , the computing device 120 obtains a musical score file at 102. In some embodiments, the computing device 120 can obtain the musical score file through wired transmission, wireless transmission, Bluetooth transmission, or infrared transmission, etc. A musical score file is a file used to record musical melodies and rhythms, and usually uses a staff notation or numbered musical notation and other symbol systems to record musical elements such as pitch, duration, and intensity. The musical scores involved in the present disclosure have various forms, such as staff notation, numbered musical notation, or Gongche notation, etc. In the musical score file, the notes are recorded on the musical score, and different symbols and markings are used to represent musical elements such as pitch, duration, and dynamics. These files can be used in fields such as music creation, rehearsal, performance, and teaching. A musical score is a musical symbol language that represents music through symbols, numbers, or patterns, etc., and can record various forms of music, thereby helping the author or reader better understand the musical work and providing specific guidance and help for music creators and performers.

[0032] In some embodiments, the format of the score file is a musical instrument digital interface (MIDI) format, a music XML format (MusicXML), etc. The MIDI format is a digital music interface format, which is used to record the notes and rhythm information of music, and can be used for applications such as music production and automatic performance. MusicXML is a music file format based on extensible markup language (XML), which can be used for music typesetting, publishing, and digital music production. In some embodiments, the score file includes a score file containing Chinese lyrics or a score file containing foreign lyrics. In some embodiments, the score file can be any entire song or any song segment in the entire song.

[0033] Continue to refer Figure 1 After the computing device 120 obtains the score file, it parses the score file at 104 to find that the score file is marked with a breathing port, which is a sign of the breathing sound of a real human singer during singing. When singing, the size and expression of the breathing sound are also one of the important criteria for judging a singer's singing skills. Some singers will reduce the breathing sound or control the expression of the breathing sound through special breathing exercises and vocal training to improve their singing level.

[0034] In some embodiments, the breathing port can be obtained by model reasoning, for example, the breathing port can be predicted by a breathing prediction mark model, and the breathing mark can be marked on the score file. In some embodiments, some breathing marks are manually marked. These score files that already have breathing marks are often used by singers to help themselves better plan their breath and breathing, remind themselves of the details that need to be paid attention to in singing, and the performance methods that need to be paid attention to in performance. In some embodiments, the breathing port is the pause indicated by the rest, and the breathing port is the personal breathing point mastered by human singers in order to interpret songs.

[0035] Continue to refer Figure 1 , the computing device 120 divides the score file into a plurality of score segments according to the ventilation port at 106. Figure 1 After obtaining the plurality of music score segments, the computing device performs singing synthesis model reasoning 108, and inputs the segmented plurality of music score segments into the singing synthesis model for reasoning to obtain audio segments 110. In some embodiments, before inputting the music score segments into the singing synthesis model for reasoning, symbols are added to the head and tail of each audio segment, for example, a breathing symbol is added to the head so that the synthesized audio segment has a more natural breathing rhythm.

[0036] In some embodiments, the singing synthesis model may be a combination of single or multiple various singing synthesis models, including but not limited to. In some embodiments, before performing the singing synthesis model inference 108, the data of the sheet music file is extracted and preprocessed. For example, features such as lyrics, phonemes, and notes are extracted from the sheet music file and these features are cleaned to obtain clean data for subsequent model training.

[0037] Continuing to refer to Figure 1 , after generating multiple audio segments, the computing device splices the audio segments at 112 to splice multiple audio segments into a complete audio segment. In some embodiments, when splicing the audio segments, different strategies are selected to splice the audio. For example, if the end of the previous segment of the current segment is a long note, then the breathing sound of the current segment is covered on a part of the long note of the previous segment. Returning to Figure 1 , after processing according to the splicing strategy, the computing device 120 outputs the complete singing at 114. At this time, the listener can vividly experience the complete synthesized singing that is consistent with the naturalness of a live performance and has a comfortable rhythm, improving the user experience.

[0038] Embodiments of the present disclosure obtain a sheet music file with breathing marks, split the sheet music file into multiple sheet music segments according to the breathing marks, generate multiple audio segments corresponding to the multiple sheet music segments, and then synthesize the singing corresponding to the sheet music file based on the multiple audio segments. The sheet music file is segmented according to the actual breathing rhythm of the singer and input into the singing synthesis model for inference to obtain audio segments and synthesize them into complete singing. In this way, the synthesized singing can better simulate the tone and intonation during a live performance, and the synthesized singing is also more fluent and natural, thus improving the listener's experience.

[0039] It should be understood that the architecture and functions in the example environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure. Embodiments of the present disclosure can also be applied to other environments with different structures and / or functions.

[0040] The following will be combined with Figures 2 to 9 to describe in detail the process according to the embodiments of the present disclosure. For ease of understanding, the specific data mentioned in the following description are all exemplary and are not used to limit the protection scope of the present disclosure. It can be understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.

[0041] Figure 2The flowchart of method 200 for synthesizing singing voice according to some embodiments of the present disclosure is shown. At block 202, a musical score file with breath-taking marks is obtained. In some embodiments, the musical score file has been manually marked with breath-taking marks. Alternatively, the breath-taking marks are obtained through model inference. For example, the breath-taking marks can be obtained through a breath-taking mark prediction model and marked in the musical score file. In some embodiments, the breath-taking marks may include rests.

[0042] At block 204, based on the breath-taking marks, the musical score file is segmented into multiple musical score segments. The musical score file with breath-taking marks is segmented into multiple musical score segments according to the breath-taking marks to facilitate the subsequent singing voice synthesis model inference work. In some embodiments, a breath-taking symbol is added to the head of each musical score segment, and a silence symbol is added to the tail of the musical score segment. In this way, it can be ensured that the synthesized singing voice finally has a breath-taking effect, improving the naturalness and fluency of the synthesized singing voice, and also matching the training stage of the singing voice synthesis model.

[0043] At block 206, multiple audio segments corresponding to the multiple musical score segments are generated. In some embodiments, the musical score segments are input into the singing voice synthesis model for inference to obtain the audio segments. In some embodiments, the musical score segments are parsed before inference. For example, a musical score parsing tool or algorithm is used to extract the musical symbols and information in the musical score segments, so that the computing device can understand and process the data format.

[0044] In some embodiments, after parsing the musical score segments, phoneme features of pitch, duration, and intensity are extracted from the parsed data. In voice synthesis, a phoneme refers to the smallest speech unit or the smallest speech segment that constitutes a syllable, and is divided from the perspective of timbre. A note is a symbol used to record the progress of sounds of different lengths, and the rhythm and melody of music can be arranged according to the length and duration of the notes. In voice synthesis, lyrics can cooperate with the melody and notes to jointly form a complete singing voice.

[0045] In some embodiments, the phoneme sequence of the musical score segment is determined according to the extracted phoneme features. A phoneme sequence refers to a sequence formed by arranging phonemes in a speech signal in a certain order. The phoneme sequence includes the interval information of adjacent phonemes, and the interval information of adjacent phonemes refers to the time interval or duration difference between two adjacent phonemes. In some embodiments, the phoneme features are converted into spectral features. For example, it is converted into the Mel spectrum in the frequency domain. This can better utilize the local information of the speech signal. The Mel frequency is an approximation of the way the human ear perceives frequency. The conversion between the Mel frequency and the linear frequency can be completed through the Mel frequency scale formula. The Mel spectrogram better simulates the human ear's perception of sound by using the Mel scale on the frequency axis of the spectrogram. In some embodiments, the corresponding musical score segment is generated according to the converted spectral features and the phoneme sequence. In some embodiments, the musical score segment is segmented into individual words, and then the phonemes of this individual word are determined according to the part of speech and the meaning of the individual word. For example, the word "huan" in "return something" is a verb meaning "return" or "give back", so the pronunciation of this character is "huan" rather than "huai".

[0046] In block 208, based on multiple audio segments, a singing voice corresponding to the musical score file is synthesized. In some embodiments, before splicing the segments, the audio segments are spliced according to a predefined strategy. For example, if the tail of the current segment presents as a long note, then we choose to cover the breathing part of the next synthesized audio segment onto the long note segment to maintain the natural transition of the sound. For example, if both the current synthesized audio segment and its previous synthesized audio segment are continuous singing voices without pauses, then we can splice these two segments by fading in and fading out. For example, when the previous synthesized audio segment is a rest segment, we cover the breathing part of the current synthesized audio segment onto the rest segment to maintain the coherence of the sound. In this way, it can be ensured that the synthesized audio maintains the original breathing effect of a real person singing after splicing, and at the same time makes the length of the synthesized audio consistent with the length of the original song.

[0047] In this embodiment, according to the obtained musical score file with breathing marks, the musical score file is segmented into multiple segments, and multiple audio segments are generated based on these musical score segments. Then these multiple audio segments are spliced into a complete segment, and finally a complete singing voice is output. The musical score file is segmented according to the actual breathing rhythm of the singer. In this way, the synthesized singing voice can better simulate the tone and intonation during real-person singing, making the synthesized singing voice more fluent and natural, thus enhancing the experience of the listener users.

[0048] Figure 3 A schematic diagram for synthesizing the singing voice 300 according to some embodiments of the present disclosure is shown. Refer to Figure 3, the user inputs the score file of the singing voice to be synthesized into the computing device by inputting a score 302 with a breathing port. In some embodiments, the breathing ports in these score files with breathing ports are manually marked. In some embodiments, the user can transmit the score file with a breathing port to the computing device through wireless, wired, or application programming interface (API), etc. The computing device splits the score file at 304. The computing device splits the score file into multiple score segments according to the marked breathing ports in the score file. Then, at 306, segment symbols are added. In some embodiments, a breathing symbol is added at the head of the split segment, and a silence symbol is added at the tail. In this way, each score segment can have a good breathing effect, so that the finally synthesized singing voice sounds comfortable and natural in rhythm, and the practice of adding symbols at the head and tail is also more conducive to the smooth progress of subsequent singing voice synthesis model inference.

[0049] Continue to refer to Figure 3 , the computing device continues to perform singing voice synthesis model inference 308. In some embodiments, before inputting each score segment into the singing voice synthesis model inference, the score file is parsed. For example, a score parsing tool is used to extract musical symbols and information in the score file, and phonemes, notes, lyrics, etc. are extracted from the parsed data, and these data are cleaned, format-converted, and feature-extracted. After model inference, segment audio corresponding to each score segment is output at 310.

[0050] As Figure 3 shown, at 312, policy processing, the generated audio segments corresponding to the score segments will be processed based on a predetermined splicing policy. In some embodiments, if the currently synthesized audio segment and the previously synthesized audio segment of the currently synthesized audio segment are continuous and non-stop singing voices, then these two synthesized audio segments will be spliced in a fade-in and fade-out superposition manner. In some embodiments, if the tail of the previously synthesized audio segment of the currently synthesized audio segment is a long note, then the breathing sound part of the currently synthesized audio segment will be covered on this long note segment. In some embodiments, if the previously synthesized audio segment of the currently synthesized audio segment is a rest segment, then the breathing segment part of the currently synthesized audio segment will be covered on the rest segment. In this way, it can be ensured that the breathing effect is retained when the audio segments are spliced, so that the synthesized audio segments have the effect of a real person singing, and it can also ensure that the length of the synthesized singing voice is consistent with the original length of the song.

[0051] At 314, all the processed synthesized audio segments are spliced to obtain a complete audio segment. At 316, the computing device will output the complete synthesized singing voice. In this way, the phenomena of breath-holding or continuous long notes that occur in traditional synthesized singing voices are avoided, and the user can experience the listening feeling of the synthesized singing voice like a real person singing.

[0052] Figure 4 A schematic diagram optionally used for synthesizing singing voice 400 according to some embodiments of the present disclosure is shown. Referring to Figure 4 , phonemes 404, notes 406, and lyrics 408 are input into a breath prediction model 410 to obtain phoneme-level predicted breath points 412. A phoneme generally refers to the smallest distinguishable sound unit in music, such as pitch, duration, etc. A note refers to a specific musical symbol, such as a whole note, a half note, etc. Lyrics are the text part of a song. In some embodiments, before extracting phoneme, note, and lyric information, a music score parsing file is used to parse the music score file to facilitate subsequent feature extraction. In some embodiments, phonemes, notes, and lyrics are input into the breath prediction model at the phoneme level. For example, inputting [phone1, phone2, phone3, phone4, phone5, phone6, phone7, phone8] into the breath identification prediction model will output a prediction sequence of the same length [0, 0, 0, 1, 0, 0, 1, 0], where 1 represents a breath identification and 0 represents no breath. Analyze the output prediction sequence and identify the music score file according to the breath points therein.

[0053] Continuing to refer to Figure 4 , at 414, a music score with breath marks is input. The music score file with breath marks is input into a computing device. At 416, the music score is segmented. The computing device will segment the music score file according to the breath marks marked in the music score file to obtain multiple music score segments. At 418, symbols are added. A breath symbol is added at the head of each music score segment, and a silence symbol is added at the tail of each music score segment. At 420, segment-based inference. Each segment with the added breath symbol and silence symbol is input into a singing voice synthesis model for inference. At 422, the synthesized audio segments are output to obtain synthesized audio segments corresponding to each music score segment.

[0054] Continuing to refer to Figure 4, at 424, policy processing. The obtained synthesized audio segments are processed according to predefined policies. In some embodiments, both the currently synthesized audio segment and the previous synthesized audio segment of the currently synthesized audio segment are continuous and uninterrupted singing voices, and a fade-in and fade-out superimposed splicing method will be adopted for these two synthesized audio segments. In some embodiments, if the tail of the previous synthesized audio segment of the currently synthesized segment is a long note, then the breathing sound part of the currently synthesized audio segment will be covered on this long note segment. In some embodiments, if the previous synthesized audio segment of the currently synthesized audio segment is a rest segment, then the synthesized breathing segment part of the currently synthesized audio segment will be covered on the rest segment. In this way, it can ensure that the breathing effect after the synthesized audio segments are spliced is optimized. At 426, all the synthesized audio segments after policy processing are spliced. At 428, the complete synthesized song is output. In this way, the duration of the complete synthesized audio segment after such processing is consistent with the audio duration corresponding to the score file, ensuring that the original length of the synthesized song remains unchanged, and finally a synthesized singing voice with an ideal inhalation effect is obtained.

[0055] Figure 5 A schematic diagram for training the breathing label prediction model 500 according to some embodiments of the present disclosure is shown. Refer to Figure 5 , the structure of the breathing label prediction model 520 at least includes a transformer 508, a multi-layer convolution 510, and a linear layer 512. The transformer uses the attention mechanism to improve the model training speed. The transformer consists of an input encoder and an output decoder, and these encoders and decoders are connected by several self-attention layers. These layers use the attention mechanism to calculate the relationship between the input and the output, thereby allowing the transformer model to process sequences in parallel.

[0056] The multi-layer convolutional layer refers to a model that contains multiple convolutional layers in a convolutional neural network. Each convolutional layer will add some non-linear operations, such as activation functions, batch normalization, etc., to increase the complexity and expressive ability of the model. The linear layer, also known as the fully connected layer or the dense layer, each neuron in the linear layer is connected to all neurons in the previous layer to achieve a linear combination or linear transformation of the previous layer. In some embodiments, before inputting the score file into the breathing label prediction model, a score parsing tool will be used to parse the score file to obtain the phonemes, notes, and lyric features associated with the score file.

[0057] Continue to refer to Figure 5, the phonemes 502, notes 504, and lyrics 506 will be input into the breathing mark prediction model to achieve phoneme-level prediction of breathing points at 514. In some embodiments, the phonemes, notes, and lyrics will be input into the breathing mark prediction model at the phoneme level, and a prediction sequence of the same length will be obtained, where the prediction sequence contains breathing prediction marks. In some embodiments, the extracted phoneme features, note features, and lyric features will be input into the breathing mark prediction model in the form of vector embedding to obtain a zero-one vector of the same length as the phonemes as the predicted value of the breathing point. In some embodiments, the phonemes, notes, and lyrics will pass through these three modules respectively to obtain a zero-one vector of the same length as the phonemes as the predicted value of the breathing point. For example, [phone1, phone2, phone3, phone4, phone5, phone6, phone7, phone8] is input into the breathing mark prediction model, and the model will output a prediction sequence [0, 0, 0, 1, 0, 0, 1, 0] of the same length. Among them, the number 1 represents the breathing mark, and the number 0 represents the non-breathing state.

[0058] Continue to refer to Figure 5 , the breathing mark prediction model is trained using the file annotated with breathing marks. Specifically, the sheet music file without breathing marks is input into the breathing mark prediction model for training, and the breathing mark prediction model will output a phoneme sequence with breathing marks. The generated phoneme sequence is compared with the label sequence with breathing prediction marks to adjust the breathing mark prediction model. For example, after the sheet music file without breathing marks is input into the model, the sequence generated by the model is [0, 0, 0, 1, 0, 0, 0], and the label sequence is [0, 0, 1, 1, 0, 0, 0]. The loss between the generated sequence and the label sequence is compared, for example, the loss between the generated sequence and the label sequence is calculated by mean square error, mean absolute error, or cross-entropy loss.

[0059] In some embodiments, the parameters of the breathing mark prediction model 520 are adjusted according to the loss. For example, if the loss is too large or too small, model parameters such as the learning rate and regularization coefficient can be adjusted to optimize the performance of the model, and better model parameters can be obtained through repeated iterative training. In some embodiments, if the loss satisfies the corresponding loss convergence condition, that is, the loss function value gradually stabilizes and tends to be stable, it is determined that the model stops training, and at this time, the model can perform the next prediction and inference. In some embodiments, before training the model, data cleaning and data annotation of the training samples of the model will be performed. Predicting the breathing points in the sheet music file through the breathing mark prediction model can save labor and time costs.

[0060] Figure 6 A schematic diagram for segmenting the sheet music segment 600 according to some embodiments of the present disclosure is shown. AsFigure 6 As shown, in the sheet music file 610, there are rest symbols 602-1 and 602-2. After manual annotation or inference by a breath mark prediction model, breath marks 604-1, 604-2, rest symbol 602-1, and rest symbol 602-1 are obtained. In some embodiments, the sheet music file is segmented according to the breath marks. Alternatively, the sheet music segments can also be segmented according to the rest symbols simultaneously.

[0061] In some embodiments, the rest symbol is also a breath mark point. In some embodiments, the breath mark is the rhythm point of the pause and breath when a human singer is singing, and the breath mark includes the rest symbol. The sheet music file 610 is segmented into 4 sheet music segments according to the breath marks and the rest symbols, namely sheet music segment 610-1, sheet music segment 610-2, sheet music segment 610-3, and sheet music segment 610-4. Segmenting the sheet music file according to the breath points to synthesize the singing voice can avoid the phenomena of breath holding or continuous long notes that occur in traditional singing voice synthesis, and improve the user experience.

[0062] Figure 7 The figure shows a schematic diagram of adding symbols 700 to the head and tail of a sheet music segment in some embodiments of the present disclosure. Refer to Figure 7 , add a breath symbol segment 720 to the head of the sheet music segment 710 segmented according to the breath mark and add a mute symbol segment 730 to the tail. In some embodiments, the lengths of the breath symbol segment and the mute symbol segment are determined according to the habits and rhythms of live singing. For example, if the singer tends to take a deep breath and sing at this breath point, the length of the breath symbol segment will be relatively longer. By adding the breath symbol and the mute symbol, it can be ensured that there is a certain breath effect after each sheet music segment segmented according to the breath mark is synthesized into an audio segment, and the segmented sheet music segments are also more suitable for processing by a computing device.

[0063] Figure 8 The figure shows a block diagram of a device 800 for synthesizing a singing voice in some embodiments of the present disclosure. As Figure 8 shown, the device 800 includes a sheet music file acquisition module 802 configured to acquire a sheet music file with breath marks. The device 800 further includes a sheet music file segmentation module 804 configured to segment the sheet music file into multiple sheet music segments based on the breath marks. The device 800 further includes an audio segment generation module 806 configured to generate multiple audio segments corresponding to the multiple sheet music segments. The device 800 further includes a singing voice synthesis module 808 configured to synthesize a singing voice corresponding to the sheet music file based on the multiple audio segments.

[0064] Figure 9A block diagram of an electronic device 900 showing some embodiments of the present disclosure. The device 900 may be the device or apparatus described in the embodiments of the present disclosure. As Figure 9 shown, the device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or computer program instructions loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The CPU / GPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 908. An input / output (I / O) interface 905 is also connected to the bus 904. Although not shown in Figure 9 , the device 900 may also include a coprocessor.

[0065] Multiple components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0066] Each of the methods or processes described above may be executed by the CPU / GPU 901. For example, in some embodiments, the method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the CPU / GPU 901, one or more steps or actions of the methods or processes described above may be executed.

[0067] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present disclosure.

[0068] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0069] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0070] The computer program instructions for performing the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages and conventional procedural programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of this disclosure.

[0071] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. The computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0072] The computer - readable program instructions can also be loaded onto a computer, other programmable data - processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data - processing apparatus, or other device to produce a computer - implemented process, whereby the instructions executed on the computer, other programmable data - processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0074] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the technical field to understand the embodiments disclosed herein.

[0075] Some example implementations of the present disclosure are listed below.

[0076] Example 1. A method for synthesizing singing voice, comprising:

[0077] Obtaining a musical score file with a breathing mark;

[0078] Based on the breathing mark, splitting the musical score file into a plurality of musical score segments;

[0079] Generating a plurality of audio segments corresponding to the plurality of musical score segments; and

[0080] Based on the plurality of audio segments, synthesizing a singing voice corresponding to the musical score file. Example 2. The method according to Example 1, wherein obtaining a musical score file with a breathing mark includes:

[0081] Receiving the musical score file marked with the breathing mark.

[0082] Example 3. The method according to any one of Examples 1-2, wherein obtaining a musical score file with a breathing mark includes:

[0083] The ventilation identifier is determined by a ventilation identifier prediction model, and the ventilation identifier prediction model includes at least a transformer layer, multiple convolutional layers, and a linear layer.

[0084] Example 4. The method according to any one of Examples 1-3, wherein determining the ventilation identifier by a ventilation identifier prediction model includes:

[0085] By parsing the music score file, phoneme features, note features, and lyric features associated with the music score file are extracted;

[0086] Based on the phoneme features, the note features, and the lyric features, a phoneme sequence of the music score file is determined, and the phoneme sequence includes at least a ventilation identifier; and

[0087] Based on the phoneme sequence, the music score file with the ventilation identifier is obtained.

[0088] Example 5. The method according to any one of Examples 1-4 further includes:

[0089] Training the ventilation identifier prediction model based on a plurality of music score files annotated with ventilation identifiers.

[0090] Example 6. The method according to any one of Examples 1-5, wherein training the ventilation identifier prediction model based on a plurality of music score files annotated with ventilation identifiers includes:

[0091] Inputting a music score file not annotated with a ventilation identifier into the ventilation identifier prediction model;

[0092] Generating a phoneme sequence including a ventilation identifier; and

[0093] Based on the loss between the generated phoneme sequence and the annotated phoneme sequence, adjusting the parameters of the ventilation identifier prediction model, and the annotated phoneme sequence is obtained based on a music score file annotated with a ventilation identifier.

[0094] Example 7. The method according to any one of Examples 1-6 further includes:

[0095] In response to the loss satisfying a loss convergence condition, generating the ventilation identifier prediction model.

[0096] Example 8. The method according to any one of Examples 1-7, wherein inputting a music score file not annotated with a ventilation identifier into the ventilation identifier prediction model includes:

[0097] Parsing the music score file not annotated with a ventilation identifier;

[0098] Extracting the phoneme features, note features, and lyric features of the unannotated music score file; and

[0099] Input the embeddings of the phoneme features, the note features, and the lyric features into the breath mark prediction model.

[0100] Example 9. The method according to any one of Examples 1-8 further includes:

[0101] Adding a breath mark segment to the head of each of the plurality of musical score segments; and

[0102] Adding a silence mark segment to the tail of each of the plurality of musical score segments.

[0103] Example 10. Synthesizing a singing voice corresponding to the musical score file based on the plurality of audio segments in the method according to any one of Examples 1-9 includes:

[0104] Stitching the plurality of audio segments based on a predefined stitching strategy; and

[0105] Synthesizing a singing voice corresponding to the musical score file based on the stitched plurality of audio segments.

[0106] Example 11. In the method according to any one of Examples 1-10, stitching the plurality of audio segments based on a predefined stitching strategy includes:

[0107] In response to the tail of the first audio segment being a long note, covering the breath mark segment of the second audio segment onto a partial long note segment of the first audio segment, where the first audio segment is before the second audio segment;

[0108] In response to the third audio segment being a singing voice and the fourth audio segment being a singing voice and the singing voices being coherent, superimposing the breath mark segment of the fourth audio segment onto the silence mark segment of the third audio segment, where the third audio segment is before the fourth audio segment; and

[0109] In response to the tail of the fifth audio segment being a rest, partially covering the breath mark segment of the sixth audio segment onto a partial rest segment of the fifth audio segment, where the fifth audio segment is before the sixth audio segment.

[0110] Example 12. In the method according to any one of Examples 1-11, generating the plurality of audio segments corresponding to the plurality of musical score segments includes:

[0111] Parsing the musical score segments in the plurality of musical score segments;

[0112] Extracting phoneme features including at least pitch, duration, and intensity based on the parsed musical score segments;

[0113] Based on the phoneme features, determine the phoneme sequence of the music score segment, where the phoneme sequence indicates the interval information of adjacent phonemes;

[0114] Based on the phoneme features, determine the spectral features of the music score segment; and

[0115] Determine the audio segment of the music score segment based on at least the spectral features and the phoneme sequence.

[0116] Example 13. The method according to any one of Examples 1-12,

[0117] further comprising:

[0118] Segment the lyrics of the music score segment into single words;

[0119] Based on the single word, determine the part of speech of the single word; and

[0120] Based on the single word and the part of speech, determine the phoneme sequence of the single word.

[0121] Example 14. An apparatus for synthesizing singing voice, comprising:

[0122] A music score file acquisition module, configured to acquire a music score file with a breath mark;

[0123] A music score file segmentation module, configured to segment the music score file into a plurality of music score segments based on the breath mark;

[0124] An audio segment generation module, configured to generate a plurality of audio segments corresponding to the plurality of music score segments; and

[0125] A singing voice synthesis module, configured to synthesize a singing voice corresponding to the music score file based on the plurality of audio segments.

[0126] Example 15. The apparatus according to Example 14, wherein the music score file acquisition module comprises:

[0127] A music score file receiving module, configured to receive the music score file marked with the breath mark.

[0128] Example 16. The apparatus according to any one of Examples 14-15, wherein the music score file acquisition module comprises:

[0129] A breath mark determination module, configured to determine the breath mark through a breath mark prediction model, where the breath mark prediction model at least comprises a transformer layer, a multi-layer convolutional layer, and a linear layer.

[0130] Example 17. The apparatus according to any one of Examples 14-16, wherein the ventilation identification determination module comprises:

[0131] A feature extraction module, configured to extract phoneme features, note features, and lyric features associated with the music score file by parsing the music score file;

[0132] A phoneme sequence determination module, configured to determine a phoneme sequence of the music score file based on the phoneme features, the note features, and the lyric features, the phoneme sequence at least including a ventilation identification; and

[0133] A music score file obtaining module with ventilation identification, configured to obtain the music score file with the ventilation identification based on the phoneme sequence.

[0134] Example 18. The apparatus according to any one of Examples 14-17, further comprising:

[0135] A ventilation identification prediction model training module, configured to train the ventilation identification prediction model based on a plurality of music score files annotated with ventilation identifications.

[0136] Example 19. The apparatus according to any one of Examples 14-18, wherein the ventilation identification prediction model training module comprises:

[0137] A music score file input module without ventilation identification annotation, configured to input a music score file without ventilation identification annotation into the ventilation identification prediction model;

[0138] A phoneme sequence generation module with ventilation identification, configured to generate a phoneme sequence including ventilation identifications; and

[0139] A parameter adjustment module of the ventilation identification prediction model, configured to adjust parameters of the ventilation identification prediction model based on a loss between the generated phoneme sequence and the annotated phoneme sequence, the annotated phoneme sequence being obtained based on a music score file annotated with ventilation identification.

[0140] Example 20. The apparatus according to any one of Examples 14-19, further comprising:

[0141] A ventilation identification prediction model generation module, configured to generate the ventilation identification prediction model in response to the loss satisfying a loss convergence condition.

[0142] Example 21. The apparatus according to any one of Examples 14-20, wherein the music score file input module without ventilation identification annotation comprises:

[0143] A music score file parsing module, configured to parse the music score file without ventilation identification annotation;

[0144] A feature extraction module, configured to extract phoneme features, note features, and lyric features of the unannotated music score file; and

[0145] A feature embedding module, configured to input the embeddings of the phoneme features, the note features, and the lyric features into the breath mark prediction model.

[0146] Example 22. The apparatus according to any one of Examples 14-21 further includes:

[0147] A breath mark segment adding module, configured to add a breath mark segment at the head of each of the multiple music score segments; and

[0148] A silence mark segment adding module, configured to add a silence mark segment at the tail of each of the multiple music score segments.

[0149] Example 23. The apparatus according to any one of Examples 14-22, wherein the singing voice synthesis module includes:

[0150] An audio splicing module, configured to splice the multiple audio segments based on a predefined splicing strategy; and

[0151] A singing voice synthesis module corresponding to the music score file, configured to synthesize a singing voice corresponding to the music score file based on the spliced multiple audio segments.

[0152] Example 24. The apparatus according to any one of Examples 14-23, wherein the audio splicing module includes:

[0153] A breath mark covering module, configured to, in response to the tail of a first audio segment being a long note, cover a breath mark segment of a second audio segment onto a partial long note segment of the first audio segment, where the first audio segment is before the second audio segment;

[0154] A breath mark segment superposition module, configured to, in response to a third audio segment being a singing voice and a fourth audio segment being a singing voice and the singing voices being coherent, superpose the breath mark segment of the fourth audio segment onto a silence mark segment of the third audio segment, where the third audio segment is before the fourth audio segment; and

[0155] A breath mark partial covering module, configured to, in response to the tail of a fifth audio segment being a rest, partially cover the breath mark segment of a sixth audio segment onto a partial rest segment of the fifth audio segment, where the fifth audio segment is before the sixth audio segment.

[0156] Example 25. The apparatus according to any one of Examples 14-24, wherein the audio segment generation module comprises:

[0157] A musical score segment parsing module configured to parse a musical score segment from the plurality of musical score segments;

[0158] A feature extraction module configured to extract phoneme features including at least pitch, duration, and intensity based on the parsed musical score segment;

[0159] A phoneme sequence determination module configured to determine a phoneme sequence of the musical score segment based on the phoneme features, the phoneme sequence indicating interval information of adjacent phonemes;

[0160] A spectral feature determination module configured to determine spectral features of the musical score segment based on the phoneme features; and

[0161] An audio segment determination module configured to determine an audio segment of the musical score segment based on at least the spectral features and the phoneme sequence.

[0162] Example 26. The apparatus according to any one of Examples 14-25, further comprising:

[0163] A lyrics segmentation module configured to segment the lyrics of the musical score segment into individual words;

[0164] A part-of-speech determination module configured to determine the part of speech of the individual word based on the individual word; and

[0165] An individual word phoneme sequence determination module configured to determine a phoneme sequence of the individual word based on the individual word and the part of speech

[0166] Example 27. An electronic device, comprising:

[0167] A processor; and

[0168] A memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the electronic device to perform actions, the actions including:

[0169] Obtaining a musical score file with a breathing mark;

[0170] Based on the breathing mark, splitting the musical score file into a plurality of musical score segments;

[0171] Generating a plurality of audio segments corresponding to the plurality of musical score segments; and

[0172] Based on the plurality of audio segments, synthesizing a singing voice corresponding to the musical score file.

[0173] Example 28. The electronic device according to Example 27, wherein obtaining the music score file with a breathing mark includes:

[0174] Receiving the music score file marked with the breathing mark.

[0175] Example 29. The electronic device according to any one of Examples 27-28, wherein obtaining the music score file with a breathing mark includes:

[0176] Determining the breathing mark through a breathing mark prediction model, the breathing mark prediction model at least includes a transformer layer, a multi-layer convolutional layer, and a linear layer.

[0177] Example 30. The electronic device according to any one of Examples 27-29, wherein determining the breathing mark through a breathing mark prediction model includes:

[0178] By parsing the music score file, extracting phoneme features, note features, and lyric features associated with the music score file;

[0179] Based on the phoneme features, the note features, and the lyric features, determining a phoneme sequence of the music score file, the phoneme sequence at least includes a breathing mark; and

[0180] Based on the phoneme sequence, obtaining the music score file with the breathing mark.

[0181] Example 31. The electronic device according to any one of Examples 27-30, further includes:

[0182] Training the breathing mark prediction model based on a plurality of music score files marked with breathing marks.

[0183] Example 32. The electronic device according to any one of Examples 27-31, wherein training the breathing mark prediction model based on a plurality of music score files marked with breathing marks includes:

[0184] Inputting the music score file not marked with a breathing mark into the breathing mark prediction model;

[0185] Generating a phoneme sequence including a breathing mark; and

[0186] Based on the loss between the generated phoneme sequence and the marked phoneme sequence, adjusting the parameters of the breathing mark prediction model, the marked phoneme sequence is obtained based on the music score file marked with a breathing mark.

[0187] Example 33. The electronic device according to any one of Examples 27-32, the action further includes:

[0188] Generate the breath mark prediction model in response to the loss satisfying the loss convergence condition.

[0189] Example 34. The electronic device according to any one of Examples 27-33, wherein inputting the sheet music file without a marked breath mark into the breath mark prediction model includes:

[0190] Parse the sheet music file without a marked breath mark;

[0191] Extract the phoneme features, note features, and lyric features of the unmarked sheet music file; and

[0192] Input the embeddings of the phoneme features, the note features, and the lyric features into the breath mark prediction model.

[0193] Example 35. The electronic device according to any one of Examples 27-34, the operation further includes:

[0194] Add a breath mark segment to the head of each music score segment among the multiple music score segments; and

[0195] Add a silence mark segment to the tail of each music score segment among the multiple music score segments.

[0196] Example 36. The electronic device according to any one of Examples 27-35, wherein synthesizing the singing voice corresponding to the sheet music file based on the multiple audio segments includes:

[0197] Stitch the multiple audio segments based on a predefined stitching strategy; and

[0198] Synthesize the singing voice corresponding to the sheet music file based on the stitched multiple audio segments.

[0199] Example 37. The electronic device according to any one of Examples 27-36, wherein stitching the multiple audio segments based on a predefined stitching strategy includes:

[0200] In response to the tail of the first audio segment being a long note, cover the breath mark segment of the second audio segment onto a partial long note segment of the first audio segment, where the first audio segment is before the second audio segment;

[0201] In response to the third audio segment being a singing voice and the fourth audio segment being a singing voice and the singing voices being coherent, superimpose the breath mark segment of the fourth audio segment onto the silence mark segment of the third audio segment, where the third audio segment is before the fourth audio segment; and

[0202] In response to the tail of the fifth audio segment being a rest, covering a part of the breath mark segment of the sixth audio segment onto a part of the rest segment of the fifth audio segment, where the fifth audio segment is before the sixth audio segment.

[0203] Example 38. The electronic device according to any one of Examples 27-37, wherein generating a plurality of audio segments corresponding to the plurality of music score segments includes:

[0204] Parsing the music score segments among the plurality of music score segments;

[0205] Based on the parsed music score segments, extracting phoneme features including at least pitch, duration, and intensity;

[0206] Based on the phoneme features, determining a phoneme sequence of the music score segment, where the phoneme sequence indicates interval information of adjacent phonemes;

[0207] Based on the phoneme features, determining a spectral feature of the music score segment; and

[0208] Determining an audio segment of the music score segment based on at least the spectral feature and the phoneme sequence.

[0209] Example 39. The electronic device according to any one of Examples 27-38, the action further includes:

[0210] Segmenting the lyrics of the music score segment into individual words;

[0211] Based on the individual words, determining the part of speech of the individual words; and

[0212] Based on the individual words and the part of speech, determining the phoneme sequence of the individual words

[0213] Example 40. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of Examples 1 to 13.

[0214] Example 41. A computer program product, the computer program product being tangibly stored on a computer-readable medium and including computer-executable instructions, the computer-executable instructions causing the device to execute the method according to any one of Examples 1 to 13 when executed by the device.

[0215] Although the present disclosure has been described in language specific to structural features and / or method logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. A method for synthesizing singing voice, comprising: Obtaining a musical score file with breathing mark; Based on the breathing mark, splitting the musical score file into multiple musical score segments; Generating multiple audio segments corresponding to the multiple musical score segments; And Based on the multiple audio segments, synthesizing a singing voice corresponding to the musical score file.

2. The method according to claim 1, wherein obtaining a musical score file with breathing mark comprises: Receiving the musical score file marked with the breathing mark.

3. The method according to claim 1, wherein obtaining a musical score file with breathing mark comprises: Determining the breathing mark through a breathing mark prediction model, the breathing mark prediction model at least comprising a converter layer, a multi-layer convolutional layer and a linear layer.

4. The method according to claim 3, wherein determining the breathing mark through a breathing mark prediction model comprises: By parsing the musical score file, extracting phoneme features, note features and lyric features associated with the musical score file; Based on the phoneme features, the note features and the lyric features, determining a phoneme sequence of the musical score file, the phoneme sequence at least comprising a breathing mark; And Based on the phoneme sequence, obtaining the musical score file with the breathing mark.

5. The method according to claim 3, further comprising: Training the breathing mark prediction model based on multiple musical score files marked with breathing marks.

6. The method according to claim 5, wherein training the breathing mark prediction model based on multiple musical score files marked with breathing marks comprises: Inputting a musical score file without a marked breathing mark into the breathing mark prediction model; Generating a phoneme sequence including a breathing mark; And Based on the loss between the generated phoneme sequence and the marked phoneme sequence, adjusting the parameters of the breathing mark prediction model, the marked phoneme sequence being obtained based on a musical score file marked with a breathing mark.

7. The method according to claim 6, further comprising: In response to the loss satisfying a loss convergence condition, generating the breathing mark prediction model.

8. The method according to claim 6, wherein inputting a musical score file without a marked breathing mark into the breathing mark prediction model comprises: Parsing the musical score file without a marked breathing mark; Extracting phoneme features, note features and lyric features of the musical score file without a mark; And Inputting the embeddings of the phoneme features, the note features and the lyric features into the breathing mark prediction model.

9. The method according to claim 1, further comprising: Adding a breathing symbol segment at the head of each musical score segment among the multiple musical score segments; And Adding a silent symbol segment at the tail of each musical score segment among the multiple musical score segments.

10. The method according to claim 1, wherein synthesizing a singing voice corresponding to the musical score file based on the multiple audio segments comprises: Based on a predefined splicing strategy, splicing the multiple audio segments; And Based on the spliced multiple audio segments, synthesizing a singing voice corresponding to the musical score file.

11. The method according to claim 10, wherein splicing the plurality of audio segments based on a predefined splicing strategy includes: In response to a long note at the end of a first audio segment, covering a breathing symbol segment of a second audio segment onto a partial long note segment of the first audio segment, the first audio segment being before the second audio segment; In response to a third audio segment being a singing voice and a fourth audio segment being a singing voice and the singing voices being coherent, superimposing the breathing symbol segment of the fourth audio segment onto a silent symbol segment of the third audio segment, the third audio segment being before the fourth audio segment; and In response to a rest at the end of a fifth audio segment, partially covering the breathing symbol segment of a sixth audio segment onto a partial rest segment of the fifth audio segment, the fifth audio segment being before the sixth audio segment.

12. The method according to claim 1, wherein generating a plurality of audio segments corresponding to the plurality of musical score segments includes: Parsing the musical score segments in the plurality of musical score segments; Extracting phoneme features including at least pitch, duration, and intensity based on the parsed musical score segments; Determining a phoneme sequence of the musical score segment based on the phoneme features, the phoneme sequence indicating interval information of adjacent phonemes; Determining a spectral feature of the musical score segment based on the phoneme features; and Determining an audio segment of the musical score segment based at least on the spectral feature and the phoneme sequence.

13. The method according to claim 12, further comprising: Segmenting the lyrics of the musical score segment into individual words; Determining the part of speech of the individual word based on the individual word; and And Determining the phoneme sequence of the individual word based on the individual word and the part of speech.

14. An apparatus for synthesizing singing voices, comprising: A musical score file acquisition module configured to acquire a musical score file with breathing identifications; A musical score file segmentation module configured to segment the musical score file into a plurality of musical score segments based on the breathing identifications; An audio segment generation module configured to generate a plurality of audio segments corresponding to the plurality of musical score segments; and And A singing voice synthesis module configured to synthesize a singing voice corresponding to the musical score file based on the plurality of audio segments.

15. An electronic device, comprising: A processor; And A memory coupled to the processor, the memory having instructions stored therein, which when executed by the processor cause the electronic device to execute the method according to any one of claims 1 to 13.

16. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 13.