Singing voice synthesis method, computer device and storage medium

By adjusting the pitch interval of the score information of the synthesized songs, the problem of inconsistency in the timbre of synthetic songs in the prior art is solved, and a higher range matching and timbre coordination is achieved.

CN115691468BActive Publication Date: 2025-05-09TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211353720.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-05-09
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

While ensuring the coverage of the range, it is difficult to ensure the timbre coordination of the synthesized singing data, especially when the user's training data is limited and the template range is fixed.

Method used

By adjusting the pitch interval of the score information of the synthesized song, the range matching degree is improved according to the range coverage determined by the singing data of the target object, thereby ensuring the timbre coordination of the synthesized song data.

Benefits of technology

Without changing the range of the target object, adjusting the pitch of the music score significantly improves the range matching and timbre coordination of the synthetic singing data, and improves the synthesis effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691468B_ABST
    Figure CN115691468B_ABST
Patent Text Reader

Abstract

The present application discloses a singing voice synthesis method, a computer device and a storage medium, the method comprising: obtaining singing voice data of a target object, and performing range detection on the singing voice data to obtain a first range of the target object; obtaining initial music score information of a song to be synthesized, and obtaining a second range of the song to be synthesized according to the initial music score information; determining a comparison result of the coincidence between the first range and the second range, if the coincidence comparison result indicates that the first range and the second range do not match, adjusting the pitch interval of the initial music score information to obtain target music score information, and finally calling an acoustic synthesis model associated with the target object to perform singing voice data synthesis processing according to the target music score information to obtain synthesized singing voice data associated with the target object and the song to be synthesized. Through this method, the music score of the song to be synthesized can be adjusted to improve the range matching degree, so that the synthesized singing voice data has a better effect and sounds more coordinated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a singing synthesis method, a computer device and a computer-readable storage medium. Background Art

[0002] Singing synthesis is a way of outputting singing voices through machines. Users can use singing synthesis technology to create music and obtain songs of different styles. In the process of singing synthesis, the range of sounds that singing synthesis can cover is limited. At present, singing synthesis is mostly done by increasing the training data of singing synthesis models to ensure the range of sounds covered by the model. The advantage of this approach is that it solves the pronunciation problem from the data side, and the method is direct and effective; but the disadvantage is that it requires a large amount of singing data and is expensive. If you do not purchase singing data, you need to use a method similar to data enhancement to expand the data volume. This approach will reduce the quality of training data, resulting in a discount in the sound quality of model output, and may also increase time and price costs.

[0003] At the same time, when using training data for singing synthesis, since the user's training data is limited and the template's vocal range is fixed, there is often a mismatch between the two, resulting in an uncoordinated timbre. For example, if a female user selects a male version of a song for synthesis, the resulting voice will be deeper than the user's original female voice. Therefore, how to ensure the vocal range and make the synthesized song data coordinated has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of the present application provide a singing synthesis method, a computer device and a storage medium, which can adjust the music score of the song to be synthesized, thereby increasing the degree of range matching and making the synthesized singing data sound more harmonious.

[0005] In a first aspect, an embodiment of the present application discloses a singing synthesis method, the method comprising:

[0006] Acquire singing data of a target object, and perform range detection on the singing data to obtain a first range of the target object; acquire initial score information of a song to be synthesized, and acquire a second range of the song to be synthesized according to the initial score information;

[0007] Determining a comparison result of the degree of overlap between the first sound range and the second sound range;

[0008] If the coincidence comparison result indicates that the first range and the second range do not match, the pitch interval of the initial music score information is adjusted to obtain target music score information; wherein the target pitch adjustment parameter is determined according to the first range and the second range, or is determined according to a reference range;

[0009] The acoustic synthesis model associated with the target object is called so that the acoustic synthesis model performs singing data synthesis processing according to the target music score information to obtain synthesized singing data associated with the target object and the song to be synthesized.

[0010] In a second aspect, an embodiment of the present application discloses a singing voice synthesis device, the device comprising:

[0011] An acquisition unit is used to acquire singing data of a target object, and perform range detection on the singing data to obtain a first range of the target object; acquire initial score information of a song to be synthesized, and acquire a second range of the song to be synthesized according to the initial score information;

[0012] a determination unit, configured to determine a result of a comparison of the degree of overlap between the first range and the second range;

[0013] a processing unit, configured to adjust the pitch interval of the initial music score information to obtain target music score information if the coincidence comparison result indicates that the first range and the second range do not match;

[0014] The processing unit is also used to call the acoustic synthesis model associated with the target object, so that the acoustic synthesis model performs singing data synthesis processing according to the target music score information to obtain synthesized singing data associated with the target object and the song to be synthesized.

[0015] In a third aspect, an embodiment of the present application discloses a computer device, which includes a processor suitable for implementing one or more computer programs; and a computer storage medium, wherein the computer storage medium stores one or more computer programs, and the one or more computer programs are suitable for being loaded and executed by the processor to perform the above-mentioned singing synthesis method.

[0016] In a fourth aspect, the present application discloses a computer-readable storage medium, which stores one or more computer programs, and the one or more computer programs are suitable for being loaded by a processor and executing the above-mentioned singing synthesis method.

[0017] In a fifth aspect, an embodiment of the present application discloses a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device performs the above-mentioned singing synthesis method.

[0018] In the embodiment of the present application, the singing data of the target object is first obtained, and the singing data is detected in the range to obtain the first range of the target object; the initial score information of the song to be synthesized is obtained, and the second range of the song to be synthesized is obtained according to the initial score information; the range matching degree of the target object and the song to be synthesized is first determined, that is, the comparison result of the overlap between the first range and the second range is determined, and if the first range and the second range do not match, the pitch interval of the initial score information is adjusted to obtain the target score information. In this process, the target score information is obtained by adjusting the pitch interval of the initial score, and then the acoustic synthesis model associated with the target object is called to perform singing data synthesis processing according to the target score information to obtain the synthesized singing data associated with the target object and the song to be synthesized. Through this method, the vocal range coverage of the target object is determined based on the singing data of the target object, and then the pitch of the song to be synthesized is adjusted, so that the vocal range matching degree is high. Without changing the vocal range of the target object, the pitch of the music score is changed, which can ensure that the timbre of the synthesized singing data changes little, making the synthesized singing data better and sounding more coordinated. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 It is a network architecture diagram of a singing synthesis system disclosed in an embodiment of the present application;

[0021] Figure 2 It is a flow chart of a singing voice synthesis method disclosed in an embodiment of the present application;

[0022] Figure 3a It is a statistical diagram of the vocal range of a female voice disclosed in an embodiment of the present application;

[0023] Figure 3b It is a statistical diagram of the range of a song disclosed in the embodiment of the present application;

[0024] Figure 3c is a statistical diagram of the range of another song disclosed in the embodiment of the present application;

[0025] Figure 4 It is a flow chart of another singing voice synthesis method disclosed in an embodiment of the present application;

[0026] Figure 5 It is a schematic diagram of the architecture of an acoustic synthesis model disclosed in an embodiment of the present application;

[0027] Figure 6 It is a structural schematic diagram of a singing voice synthesis device disclosed in an embodiment of the present application;

[0028] Figure 7 It is a structural schematic diagram of a computer device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0030] In order to make the synthesized singing data more effective (high range matching, the song sounds more coordinated), the embodiment of the present application proposes a singing synthesis method, which can analyze the range of the target object and the song to be synthesized during the lyrics synthesis process, and adjust the pitch of the song to be synthesized when the range coverage of the two does not match, so that the matching degree of the two is increased, and then synthesized again, so that the synthesized singing data sounds more coordinated. The singing synthesis method provided in the embodiment of the present application can be implemented based on AI (Artificial Intelligence) technology. AI refers to the theory, method, technology and application system of using digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. AI technology is a comprehensive discipline, and the fields it involves are relatively wide; and the singing synthesis method provided in the embodiment of the present application mainly involves machine learning (ML) technology in AI technology. Machine learning generally includes artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by formula teaching.

[0031] In a feasible embodiment, the singing voice synthesis method provided in the embodiment of the present application can also be implemented based on cloud technology (Cloud technology) and / or blockchain technology. Specifically, it may involve one or more of cloud storage (Cloud storage), cloud database (Cloud Database), and big data (Big data) in cloud technology. For example, the data required to execute the singing voice synthesis method (such as the singing voice data of the target object, the initial music score information of the song to be synthesized, etc.) is obtained from the cloud database. For another example, the data required to execute the singing voice synthesis method can be stored in the form of blocks on the blockchain; the data generated by executing the singing voice synthesis method (such as candidate music score information, target music score information, and synthesized singing voice data, etc.) can be stored in the form of blocks on the blockchain; in addition, the singing voice synthesis device that executes the singing voice synthesis method can be a node device in the blockchain network.

[0032] See also Figure 1 , Figure 1 is a network architecture diagram of a singing synthesis system disclosed in an embodiment of the application, such as Figure 1 As shown, the singing voice synthesis system 100 may include at least a terminal device 101 and a computer device 102, wherein the terminal device 101 and the computer device 102 may be connected in communication, and the connection mode may include a wired connection and a wireless connection, which is not limited here. In the specific implementation process, the terminal device 101 is mainly an input or output device, and may mainly input the song data of the target object and the music score information of the song to be synthesized to the computer device 102 (optionally, the song data of the target object and the music score information of the song to be synthesized may also be obtained from a database); the terminal device 101 may also be used to output the synthesized singing voice data. The computer device 102 is mainly a singing synthesis device, which is used to obtain the song data of the target object and the music score information of the song to be synthesized, and to process the song data of the target object and the music score information of the song to be synthesized respectively to obtain the first range of the target object and the second range of the song to be synthesized; it is also used to adjust the pitch range of the initial music score information to obtain the target music score information, call the acoustic synthesis model associated with the target object to perform singing data synthesis processing according to the target music score information, and obtain synthesized singing data associated with the target object and the song to be synthesized.

[0033] In one possible implementation, the terminal device 101 mentioned above includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.; the computer device 102 mentioned above can be a server, and the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. Figure 1 The network architecture diagram of the singing voice synthesis system is only an exemplary representation and is not intended to be limiting. Figure 1 The computer device 102 can be deployed as a node in the blockchain network, or the computer device 102 can be connected to the blockchain network, so that the computer device 102 can upload the synthesized singing data to the blockchain network for storage to prevent internal data from being tampered with, thereby ensuring data security.

[0034] In combination with the above-mentioned singing synthesis system, the singing synthesis method of the embodiment of the present application may generally include: the computer device 102 first obtains the singing data of the target object, and performs a range detection on the singing data to obtain the first range of the target object; obtains the initial score information of the song to be synthesized, and obtains the second range of the song to be synthesized based on the initial score information; first determines the comparison result of the degree of overlap between the first range and the second range, if the first range and the second range do not match, adjusts the initial score information to obtain the target score information, in this process, by adjusting the pitch range of the initial score information, thereby obtaining the target score information, and then calls the acoustic synthesis model associated with the target object to perform singing data synthesis processing according to the target score information, and obtains the synthesized singing data associated with the target object and the song to be synthesized. Through this method, the range of the target object's vocal range is determined based on the target object's singing data, and then the pitch range of the initial music score information is adjusted to increase the range matching degree. Without changing the target object's vocal range, the pitch of the music score is changed to ensure that the timbre of the synthesized singing data changes little, making the synthesized singing data sound more coordinated.

[0035] It should be noted that in the specific implementation of the present application, the data related to the singing data of the target object, the initial score information of the song to be synthesized, etc., are all authorized by the user. When the above embodiments of the present application are applied to specific products or technologies, the data involved in the use need to obtain the user's permission or consent, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0036] See also Figure 2, is a flow chart of a singing voice synthesis method disclosed in an embodiment of the present application. The singing voice synthesis method can be executed by a computer device. The singing voice synthesis method includes but is not limited to steps S201-S205:

[0037] S201: Acquire singing voice data of a target object, and perform range detection on the singing voice data to obtain a first range of the target object.

[0038] The target object refers to the sound that needs to be synthesized in the lyrics synthesis process, which can be singer A or the user himself. The singing data is the audio data related to the target object. For the user himself, the singing data can be the data of the user's singing practice; for singer A, the singing data can be the relevant songs sung by singer A.

[0039] In a possible implementation, the singing data can be acquired by the computer device after the user triggers the "synthesis" button in the user interface, that is, the user selects the relevant singing data in the user interface and clicks the "synthesis" button, and then the computer device acquires the singing data of the target object. On the one hand, the singing data can be sent by the user device to the computer device, or it can be acquired from the database by the computer device after receiving the synthesis instruction.

[0040] In a possible implementation, the vocal range of the target object can be obtained by analyzing the singing data of the target object. Therefore, after the singing data is obtained, the singing data is subjected to a vocal range detection to obtain the first vocal range of the target object. In this process, the fundamental frequency data in the singing data is first extracted (the fundamental frequency extraction technology refers to the fundamental frequency curve of the human voice in the user's dry voice), and then the extracted fundamental frequency data is analyzed to obtain the first vocal range of the target object.

[0041] Among them, various fundamental frequency extraction tools can be used for fundamental frequency extraction. There are many tools for fundamental frequency extraction, and the methods used in extraction include traditional digital signal processing and deep learning methods. Extraction methods based on traditional digital signal processing, such as Yin and Pyin, are highly applicable and have low computational complexity. Although the fundamental frequency extraction methods based on deep learning, such as CREPE, may have low applicability in scenarios outside the training set, they have high accuracy in scenarios within the training set (such as specific noise scenarios). The Pyin toolkit supports frame-by-frame extraction. In the calculation of each frame, multiple candidate peak and valley values ​​are found as candidate points, which effectively avoids local optimality, and the subsequent use of hidden Markov models further improves the smoothness of the fundamental frequency curve. The Yin algorithm finds the minimum period of the speech waveform by calculating the autocorrelation function.

[0042] Vocal range is a range of values, which is the range of the sound that a user can make when singing. Different species and individuals have different vocal ranges. Figure 3aAs shown in FIG. 1 , it is a statistical diagram of the vocal range of a female voice, in which the horizontal axis is the pitch and the vertical axis is the frequency of the sound (in Hz). Figure 3a It can be seen that the female vocal range is mainly distributed between 52-70, that is, G3-G4; Figure 3b As shown in the figure, it is a statistical diagram of the range of a song, where the horizontal axis is the pitch and the vertical axis is the frequency of the sound (unit Hz). Figure 3b As can be seen, the song's range is from 52 to 70.

[0043] S202: Acquire initial music score information of the song to be synthesized, and acquire the second musical range of the song to be synthesized according to the initial music score information.

[0044] In order to obtain the range of the song to be synthesized, the relevant information of the data to be synthesized can be analyzed to obtain the range of the song to be synthesized. In a possible implementation, the initial score information of the song to be synthesized can be directly obtained from the database, and then the second range of the song to be synthesized can be obtained based on the initial score information. The score information contains the pitch of the lyrics, and the range of the song can be known based on the pitch.

[0045] In another possible implementation, after the user selects the name of the song to be synthesized, the computer device obtains the original audio data of the song to be synthesized according to the identification information (unique identification identifier) ​​of the song, extracts the fundamental frequency data of the audio data, and determines the range of the song to be synthesized according to the fundamental frequency data. This method is the same as the method for obtaining the range of the target object, and will not be described here.

[0046] S203: Determine a comparison result of the degree of overlap between the first range and the second range.

[0047] In a possible implementation, after determining the first range of the target object and the second range of the song to be synthesized, the first range and the second range are matched to determine whether the ranges of the two are adapted, which may include: first determining the target range overlap between the first range and the second range; then comparing the target range overlap with the overlap threshold to obtain the overlap comparison result. If the target range overlap is greater than the overlap threshold (for example, 90%, which is adjustable), it can be considered that the first range and the second range match; if the target range overlap is less than or equal to the overlap threshold, it is considered that the first range and the second range do not match.

[0048] In another implementation, determining the comparison result of the overlap between the first range and the second range may further include: determining the pitch span value of the target object according to the first range, determining the cross pitch span value between the target object and the song to be synthesized according to the first range and the second range, and determining the comparison result of the overlap between the first range and the second range according to the cross pitch span value and the pitch span value of the target object. In this process, determining the cross pitch span value between the target object and the song to be synthesized according to the first range and the second range may further include: determining the upper limit value of the pitch of the target object according to the first range, and determining the upper limit value of the pitch of the song to be synthesized according to the second range; if the upper limit value of the pitch of the song to be synthesized is less than the upper limit value of the pitch of the target object, determining the difference between the upper limit value of the pitch of the song to be synthesized and the lower limit value of the pitch of the target object as the cross pitch span value; if the upper limit value of the pitch of the song to be synthesized is greater than or equal to the upper limit value of the pitch of the target object, determining the difference between the upper limit value of the pitch of the target object and the lower limit value of the pitch of the song to be synthesized as the cross pitch span value.

[0049] Further, after determining the cross pitch span value and the pitch span value of the target object, the target range overlap between the first range and the second range is determined, which may specifically include determining the ratio between the cross pitch span value and the pitch span value, and then comparing the ratio with a set value, and taking the maximum value between the ratio and the set value as the target range overlap between the first range and the second range. Finally, the target range overlap is compared with the overlap threshold to obtain the overlap comparison result. If the target range overlap is greater than the overlap threshold (for example, 90%, which is adjustable), it can be considered that the first range and the second range match; if the target range overlap is less than or equal to the overlap threshold, it is considered that the first range and the second range do not match.

[0050] The calculation of target range overlap is described by taking an example. Assuming that the first range of the target object is [x1, x2] and the second range of the song to be synthesized is [y1, y2], the target range overlap ratio (ratio) can be calculated using the following formula (1):

[0051]

[0052] Among them, ratio1 is the initial range overlap determined based on the first range and the second range. In order to ensure that the range overlap value is not negative, the larger formula is used to update ratio1 to obtain the target range overlap. x2-x1 is the pitch span value of the target object, and y2-x1 and x2-y1 are the cross pitch span values ​​between the target object and the song to be synthesized under different conditions.

[0053] For example, the first range of the target object is 52 to 62, and the second range of the song to be synthesized can be 36 to 50. Since 50 is less than 62, according to formula (1), x2-x1=10, y2-x1=-2, ratio1=-0.2, ratio=0. Based on this, it can be concluded that the target range matching degree of the first range and the second range is 0. For another example, the first range of the target object is 52 to 62, and the second range of the song to be synthesized can be 54 to 68. Since 68 is greater than 62, according to formula (1), x2-x1=10, x2-y1=8, ratio1=0.8, ratio=0.8. Based on this, it can be concluded that the target range matching degree of the first range and the second range is 0.8.

[0054] S204: If the comparison result of the degree of coincidence indicates that the first range and the second range do not match, the pitch interval of the initial music score information is adjusted to obtain the target music score information.

[0055] Among them, whether the first range and the second range match is very important. If they match, the synthesized singing data will have a better effect. If they do not match, the synthesized singing data will not have an ideal effect. For example, Figure 3a The female voice range is between 52-70. Figure 3b The range of the songs in the song is from 52 to 70, that is to say, Figure 3a The timbre of the female voice in the film can be fully mastered Figure 3b For example, Figure 3c The following is a statistical chart of the range of another song. Figure 3c It can be seen that the range of the song is from 40 to 65, so Figure 3a The bass part of the female voice in the song is not within the range of the song, so Figure 3a The female voice and Figure 3c If we synthesize this song in , the synthesis effect is not ideal.

[0056] In a possible implementation, if the coincidence comparison result indicates that the first range and the second range do not match, the pitch interval of the initial score information can be adjusted according to the target pitch adjustment parameter to obtain the target score information. The target pitch adjustment parameter is determined according to the first range and the second range, or according to the reference range.

[0057] If the target pitch adjustment parameter is determined according to the first range and the second range, the pitch interval of the initial score information is adjusted according to the target pitch adjustment parameter, and the target score information includes: determining the pitch average of the target object according to the first range, determining the pitch average of the song to be synthesized according to the second range, determining the deviation parameter according to the pitch average of the target object and the pitch average of the song to be synthesized, determining the deviation parameter as the target pitch adjustment parameter, adding the pitch information contained in the initial score information to the target pitch adjustment parameter, and obtaining the target score information. This scheme obtains different target score information for different objects (users). Even for the same song to be synthesized, the target score information corresponding to different objects is different, because different objects have different ranges, and the obtained pitch averages are different, so that the deviation parameters are different, and correspondingly, the obtained target score information is also different.

[0058] If the target pitch adjustment parameter is determined according to the reference range, the pitch interval of the initial music score information is adjusted according to the target pitch adjustment parameter, and the target music score information includes: obtaining the reference range, the reference range includes the bass range, the middle range and the treble range, and obtaining the lower limit value of the pitch of the bass range, the lower limit value of the pitch of the middle range and the lower limit value of the pitch of the treble range. Determine the first pitch adjustment parameter according to the lower limit value of the pitch of the bass range and the lower limit value of the pitch of the song to be synthesized, determine the second pitch adjustment parameter according to the lower limit value of the pitch of the middle range and the lower limit value of the pitch of the song to be synthesized, and determine the third pitch adjustment parameter according to the lower limit value of the pitch of the treble range and the lower limit value of the pitch of the song to be synthesized. Then adjust the pitch interval of the initial music score information according to the first pitch adjustment parameter, the second pitch adjustment parameter and the third pitch adjustment parameter to obtain the target music score information.

[0059] Among them, adjusting the pitch interval of the initial music score information according to the first pitch adjustment parameter, the second pitch adjustment parameter and the third pitch adjustment parameter to obtain the target music score information can include: adding the pitch information contained in the initial music score information and the first pitch adjustment parameter to obtain the first music score information, adding the pitch information contained in the initial music score information and the second pitch adjustment parameter to obtain the second music score information, and adding the pitch information contained in the initial music score information and the third pitch adjustment parameter to obtain the third music score information. Calculate the first range overlap between the first range and the range corresponding to the first score information, calculate the second range overlap between the first range and the range corresponding to the second score information, and calculate the third range overlap between the first range and the range corresponding to the third score information; then determine the maximum range overlap among the first range overlap, the second range overlap and the third range overlap; determine the score information corresponding to the maximum range overlap as the target score information, wherein the score information corresponding to the maximum range overlap is one of the first score information, the second score information or the third score information.

[0060] It should be noted that the bass range, middle range and high range can be obtained according to the custom setting of the second range of the song to be synthesized, or can be determined according to a set algorithm, that is, after the range of the song to be synthesized is determined, the corresponding bass range, middle range and high range can be obtained according to the algorithm. The calculation method of the first range overlap, the second range overlap and the third range overlap can be referred to formula (1).

[0061] S205: calling the acoustic synthesis model associated with the target object so that the acoustic synthesis model performs singing data synthesis processing according to the target music score information to obtain synthesized singing data associated with the target object and the song to be synthesized.

[0062] In a possible implementation, after determining the target score information of the song to be synthesized, the acoustic synthesis model associated with the target object is called to perform singing data synthesis processing according to the target score information, so as to obtain the synthesized singing data associated with the target object and the song to be synthesized. The target score information includes phoneme information and note sequence. The phoneme information and note sequence of the song to be synthesized are taken as input, and the acoustic synthesis model processes the phoneme information and note sequence to obtain the synthesized singing data. The acoustic synthesis model is trained in advance, and the corresponding acoustic synthesis model is unique for different objects, because user identification information is added during the training process of the acoustic synthesis model.

[0063] In an embodiment of the present application, a computer device first obtains singing data of a target object, and performs a range detection on the singing data to obtain a first range of the target object; obtains initial score information of a song to be synthesized, and obtains a second range of the song to be synthesized based on the initial score information; first determines the range of the target object and the song to be synthesized, and then matches the ranges of the two. If the first range and the second range do not match, the initial score information is adjusted according to a target pitch adjustment parameter to obtain candidate score information, and the target score information is determined based on the candidate score information. In this process, different score templates are obtained by adjusting the pitch interval of the song, and then the acoustic synthesis model associated with the target object is called to perform singing data synthesis processing based on the target score information to obtain synthesized singing data associated with the target object and the song to be synthesized. Through this method, the vocal range coverage of the target object is determined based on the singing data of the target object, and then the pitch of the song to be synthesized is adjusted, so that the vocal range matching degree is high. Without changing the vocal range of the target object, the pitch of the music score is changed, which can ensure that the timbre of the synthesized singing data changes little, making the synthesized singing data sound more coordinated.

[0064] Based on the above description, the present application embodiment discloses another singing synthesis method, see Figure 4 , is a flowchart of another singing voice synthesis method disclosed in an embodiment of the present application. Assuming that the target object is user A and the song to be synthesized is song B, the flowchart of the corresponding singing voice synthesis method includes but is not limited to the following steps:

[0065] S401: Acquire singing voice data of user A, and perform vocal range detection on the singing voice data to obtain a first vocal range of user A.

[0066] S402: Acquire initial music score information of song B, and acquire the second musical range of song B according to the initial music score information.

[0067] S403: Whether the first range and the second range match.

[0068] If the first range and the second range match, the initial score information of the song to be synthesized is directly used as the target score information, and step S405 is executed; if the first range and the second range do not match, step S404 is executed.

[0069] S404: Adjust the initial music score information according to the target pitch adjustment parameter to obtain the target music score information.

[0070] When adjusting the initial score information, there are two solutions. In actual application scenarios, you can choose one of the two solutions.

[0071] S405: Call the acoustic synthesis model associated with user A to perform singing data synthesis processing according to the target music score information to obtain synthesized singing data associated with user A and song B.

[0072] The above steps S401-S405 and Figure 2 Steps S201-204 in are the same and will not be repeated here.

[0073] In this section, the training of the acoustic synthesis model is described:

[0074] The training singing data of the target object is obtained, and the training singing data is parsed to obtain the phoneme information, note sequence and object identification of user A; the initial acoustic synthesis model is trained based on the phoneme information, note sequence and object identification, and the parameters are continuously fine-tuned to obtain the acoustic synthesis model of user A. The training singing data can be the same as the singing data in step S401, or can be obtained from a database separately. The time length of the training singing data should be long enough (for example, a total of 15 minutes) to ensure that more note sequences can be extracted.

[0075] In a possible implementation, after obtaining the training singing voice data of user A, the training singing voice data is detected to obtain the vocal range of user A; if the vocal range of user A does not meet the preset conditions, the first vocal range information is adjusted to obtain the adjusted vocal range; the note sequence is determined based on the adjusted vocal range. This process is equivalent to data enhancement processing, which mainly ensures that the vocal range coverage of user A's data is at least more than one octave, that is, 12 semitones, so as to ensure that most songs can be sung on the basis of vocal range adaptation (because the vocal range span of the absolute majority of songs is more than 12 semitones). If the vocal range of user A's training singing voice data is less than 12 semitones, it is necessary to perform data enhancement on this part of the training data. The data enhancement here refers to the lifting and lowering of the original audio dry sound by 6 semitones (pitch change without speed change processing) to broaden the pitch coverage of user A. This process uses a lifting and lowering algorithm (which can change the audio pitch (that is, the fundamental frequency) without changing the speed and the pitch change without speed change algorithm of the phoneme, and the commonly used algorithms are psola).

[0076] Among them, the acoustic synthesis model can be developed based on the FastSpeech framework, which consists of two parts: an encoder and a decoder. The FastSpeech model inputs a phoneme sequence and outputs a Mel spectrum. The main structure is a Multi-Head self-attention and one-dimensional convolution based on Tramsformer, which is called FFT Block (Feed-Forward Transformer). N FFT Blocks are stacked on the phoneme end and the spectrum end. When training the acoustic synthesis model, the input information is: object identification (different identifications for different objects), note sequence (can be obtained based on the singing data of the object) and phoneme information, and the output is a Mel spectrum. Finally, a vocoder is connected to convert the acoustic features (usually Mel spectrum) into a playable speech waveform. Unlike ordinary synthesis models, the model for personalized singing synthesis needs to add object identification on the input side, so that the model can learn features related to the object's timbre and decouple information such as singing content from the timbre. Among them, the architecture of the acoustic synthesis model can be as follows Figure 5 As shown, it can be seen that the input is phoneme information, note sequence and object identification, and then passes through the embedding layer (a way to convert discrete variables into continuous vector representation. In neural networks, embedding can not only reduce the spatial dimension of discrete variables, but also represent the variable meaningfully), and then passes through the encoder, decoder, linear layer and vocoder, and the output is a speech waveform.

[0077] In the embodiment of the present application, the acoustic synthesis model is mainly explained. Combined with the acoustic synthesis model, singing synthesis can be performed for different objects to truly reflect personalization. In this process, the music score of the song to be synthesized is adjusted to solve the problem of vocal range adaptation.

[0078] Based on the above method embodiment, the present application embodiment also provides a structural schematic diagram of a singing synthesis device. Figure 6 , which is a structural schematic diagram of a singing synthesis device provided in an embodiment of the present application. Figure 6 The singing voice synthesis device 600 shown can operate an acquisition unit 601, a determination unit 602 and a processing unit 603, specifically:

[0079] The acquisition unit 601 is used to acquire the singing data of the target object, and perform range detection on the singing data to obtain the first range of the target object; acquire the initial score information of the song to be synthesized, and acquire the second range of the song to be synthesized according to the initial score information;

[0080] A determination unit 602, configured to determine a result of a comparison of a degree of overlap between the first range and the second range;

[0081] The processing unit 603 is used to adjust the pitch interval of the initial music score information to obtain target music score information if the overlap comparison result indicates that the first range and the second range do not match; call the acoustic synthesis model associated with the target object so that the acoustic synthesis model performs singing data synthesis processing according to the target music score information to obtain synthesized singing data associated with the target object and the song to be synthesized.

[0082] In a possible implementation, when the determination unit 602 determines the result of the comparison of the degree of overlap between the first sound range and the second sound range, it is specifically configured to:

[0083] In a possible implementation, when the determination unit 602 determines the target sound range overlap between the first sound range and the second sound range, it is specifically configured to:

[0084] Determine a pitch span value of the target object according to the first sound range;

[0085] Determine a cross pitch span value between the target object and the song to be synthesized according to the first range and the second range;

[0086] A result of comparing the degree of overlap between the first range and the second range is determined according to the cross pitch span value and the pitch span value of the target object.

[0087] In a possible implementation, when the determination unit 602 determines the cross pitch span value between the target object and the song to be synthesized according to the first range and the second range, it is specifically used to:

[0088] Determine the upper limit value of the pitch of the target object according to the first sound range, and determine the upper limit value of the pitch of the song to be synthesized according to the second sound range;

[0089] If the pitch upper limit value of the song to be synthesized is less than the pitch upper limit value of the target object, determining the difference between the pitch upper limit value of the song to be synthesized and the pitch lower limit value of the target object as the crossover pitch span value;

[0090] If the pitch upper limit value of the song to be synthesized is greater than or equal to the pitch upper limit value of the target object, the difference between the pitch upper limit value of the target object and the pitch lower limit value of the song to be synthesized is determined as the crossover pitch span value.

[0091] In a possible implementation, when the determination unit 602 determines the result of the comparison of the overlap between the first range and the second range according to the cross pitch span value and the pitch span value of the target object, it is specifically configured to:

[0092] A ratio between the crossover pitch span value and the pitch span value of the target object is determined, and a result of a comparison of the overlap between the first range and the second range is determined according to the ratio.

[0093] In a possible implementation, when the processing unit 603 adjusts the pitch interval of the initial music score information to obtain the target music score information, it is specifically used to:

[0094] Determine the average pitch of the target object according to the first sound range, determine the average pitch of the song to be synthesized according to the second sound range; and determine a deviation parameter according to the average pitch of the target object and the average pitch of the song to be synthesized;

[0095] The pitch information contained in the initial music score information is added to the deviation parameter to obtain the target music score information.

[0096] In a possible implementation, when the processing unit 603 adjusts the pitch interval of the initial music score information to obtain the target music score information, it is specifically used to:

[0097] Acquire a reference range, wherein the reference range includes a bass range, a middle range, and a treble range;

[0098] Determine a first pitch adjustment parameter according to the lower limit value of the pitch of the bass range and the lower limit value of the pitch of the song to be synthesized, determine a second pitch adjustment parameter according to the lower limit value of the pitch of the middle range and the lower limit value of the pitch of the song to be synthesized, and determine a third pitch adjustment parameter according to the lower limit value of the pitch of the treble range and the lower limit value of the pitch of the song to be synthesized;

[0099] The pitch interval of the initial music score information is adjusted according to the first pitch adjustment parameter, the second pitch adjustment parameter, and the third pitch adjustment parameter to obtain target music score information.

[0100] In a possible implementation, the processing unit 603 adjusts the pitch interval of the initial music score information according to the first pitch adjustment parameter, the second pitch adjustment parameter, and the third pitch adjustment parameter to obtain the target music score information, specifically for:

[0101] Adding the pitch information contained in the initial music score information to the first pitch adjustment parameter to obtain first music score information, adding the pitch information contained in the initial music score information to the second pitch adjustment parameter to obtain second music score information, and adding the pitch information contained in the initial music score information to the third pitch adjustment parameter to obtain third music score information;

[0102] Target music score information is determined from the first music score information, the second music score information, and the third music score information.

[0103] In a possible implementation, when the determining unit 602 determines the target music score information from the first music score information, the second music score information, and the third music score information, it is specifically used to:

[0104] Calculating a first range overlap between the first range and the range corresponding to the first score information, calculating a second range overlap between the first range and the range corresponding to the second score information, and calculating a third range overlap between the first range and the range corresponding to the third score information;

[0105] Determine the maximum sound range overlap among the first sound range overlap, the second sound range overlap and the third sound range overlap; determine the music score information corresponding to the maximum sound range overlap as the target music score information, wherein the music score information corresponding to the maximum sound range overlap is the first music score information, the second music score information or the third music score information.

[0106] It can be understood that the functions of each functional unit of the singing synthesis device provided in the embodiment of the present application can be specifically implemented according to the method in the above method embodiment, and its specific implementation process can refer to the relevant description in the above method embodiment, which will not be repeated here.

[0107] In a feasible embodiment, the singing synthesis device provided in the embodiment of the present application can be implemented in software. The singing synthesis device can be stored in a memory. It can be software in the form of programs and plug-ins, and includes a series of units, including an acquisition unit, a processing unit, and a determination unit; wherein the acquisition unit, the processing unit, and the determination unit are used to implement the singing synthesis method provided in the embodiment of the present application.

[0108] In other feasible embodiments, the singing voice synthesis device provided in the embodiment of the present application can also be implemented by a combination of software and hardware. As an example, the singing voice synthesis device provided in the embodiment of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the singing voice synthesis method provided in the embodiment of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.

[0109] In the embodiment of the present application, the acquisition unit 601 acquires the singing data of the target object, and performs a range detection on the singing data to obtain the first range of the target object; the initial score information of the song to be synthesized is obtained, and the second range of the song to be synthesized is obtained according to the initial score information; the determination unit 602 determines the comparison result of the overlap between the first range and the second range, and if the first range and the second range do not match, the processing unit 603 adjusts the pitch interval of the initial score information to obtain the initial score information. In this process, the initial score information is obtained by adjusting the pitch interval of the initial score information, and then the acoustic synthesis model associated with the target object is called to perform singing data synthesis processing according to the target score information to obtain the synthesized singing data associated with the target object and the song to be synthesized. Through this method, the range coverage of the target object is determined based on the singing data of the target object, and then the pitch of the song to be synthesized is adjusted, so that the range matching degree is high, and the pitch of the score is changed without changing the range of the target object, so that the timbre of the synthesized singing data can be guaranteed to change little, so that the synthesized singing data sounds more coordinated.

[0110] See also Figure 7 , Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device described in the embodiment of the present application includes: a processor 701, a communication interface 702, and a memory 703. Among them, the processor 701, the communication interface 702, and the memory 703 can be connected via a bus or other means, and the embodiment of the present application takes the connection via a bus as an example.

[0111] Among them, the processor 701 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power on and off instructions sent by the user to the computer device, and control the computer device to perform power on and off operations; for another example, the CPU can transmit various interactive data between the internal structures of the computer device, and so on. The communication interface 702 can optionally include a standard wired interface, a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), which is controlled by the processor 701 to send and receive data. The memory 703 (Memory) is a memory device in the computer device for storing programs and data. It can be understood that the memory 703 here can include both the built-in memory of the computer device and the extended memory supported by the computer device. The memory 703 provides a storage space, which stores the operating system of the computer device, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc., and this application does not limit this.

[0112] In the embodiment of the present application, the processor 701 performs the following operations by running the executable program code in the memory 703:

[0113] Acquire singing data of a target object, and perform range detection on the singing data to obtain a first range of the target object; acquire initial score information of a song to be synthesized, and acquire a second range of the song to be synthesized according to the initial score information;

[0114] Determining a comparison result of the degree of overlap between the first sound range and the second sound range;

[0115] If the comparison result of the degree of coincidence indicates that the first range and the second range do not match, adjusting the pitch interval of the initial music score information to obtain target music score information;

[0116] The acoustic synthesis model associated with the target object is called so that the acoustic synthesis model performs singing data synthesis processing according to the target music score information to obtain synthesized singing data associated with the target object and the song to be synthesized.

[0117] In a possible implementation, when the processor 701 determines the result of the comparison of the degree of overlap between the first sound range and the second sound range, it is specifically configured to:

[0118] Determine a pitch span value of the target object according to the first sound range;

[0119] Determine a cross pitch span value between the target object and the song to be synthesized according to the first range and the second range;

[0120] A result of comparing the degree of overlap between the first range and the second range is determined according to the cross pitch span value and the pitch span value of the target object.

[0121] In a possible implementation, when the processor 701 determines the cross pitch span value between the target object and the song to be synthesized according to the first sound range and the second sound range, it is specifically configured to:

[0122] Determine the upper limit value of the pitch of the target object according to the first sound range, and determine the upper limit value of the pitch of the song to be synthesized according to the second sound range;

[0123] If the pitch upper limit value of the song to be synthesized is less than the pitch upper limit value of the target object, determining the difference between the pitch upper limit value of the song to be synthesized and the pitch lower limit value of the target object as the crossover pitch span value;

[0124] If the pitch upper limit value of the song to be synthesized is greater than or equal to the pitch upper limit value of the target object, the difference between the pitch upper limit value of the target object and the pitch lower limit value of the song to be synthesized is determined as the crossover pitch span value.

[0125] In a possible implementation, when the processor 701 determines the result of the comparison of the overlap between the first range and the second range according to the cross pitch span value and the pitch span value of the target object, it is specifically configured to:

[0126] A ratio between the crossover pitch span value and the pitch span value of the target object is determined, and a result of a comparison of the overlap between the first range and the second range is determined according to the ratio.

[0127] In a possible implementation, the processor 701 adjusts the pitch interval of the initial music score information to obtain the target music score information, specifically for:

[0128] Determine the average pitch of the target object according to the first sound range, determine the average pitch of the song to be synthesized according to the second sound range; and determine a deviation parameter according to the average pitch of the target object and the average pitch of the song to be synthesized;

[0129] The pitch information contained in the initial music score information is added to the deviation parameter to obtain the target music score information.

[0130] In a possible implementation, the processor 701 adjusts the pitch interval of the initial music score information to obtain the target music score information, specifically for:

[0131] Acquire a reference range, wherein the reference range includes a bass range, a middle range, and a treble range;

[0132] Determine a first pitch adjustment parameter according to the lower limit value of the pitch of the bass range and the lower limit value of the pitch of the song to be synthesized, determine a second pitch adjustment parameter according to the lower limit value of the pitch of the middle range and the lower limit value of the pitch of the song to be synthesized, and determine a third pitch adjustment parameter according to the lower limit value of the pitch of the treble range and the lower limit value of the pitch of the song to be synthesized;

[0133] The pitch interval of the initial music score information is adjusted according to the first pitch adjustment parameter, the second pitch adjustment parameter, and the third pitch adjustment parameter to obtain target music score information.

[0134] In a possible implementation, the processor 701 adjusts the pitch interval of the initial music score information according to the first pitch adjustment parameter, the second pitch adjustment parameter, and the third pitch adjustment parameter to obtain the target music score information, specifically for:

[0135] Adding the pitch information contained in the initial music score information to the first pitch adjustment parameter to obtain first music score information, adding the pitch information contained in the initial music score information to the second pitch adjustment parameter to obtain second music score information, and adding the pitch information contained in the initial music score information to the third pitch adjustment parameter to obtain third music score information;

[0136] Target music score information is determined from the first music score information, the second music score information, and the third music score information.

[0137] In a possible implementation, when the processor 701 determines the target music score information from the first music score information, the second music score information, and the third music score information, it is specifically configured to:

[0138] Calculating a first range overlap between the first range and the range corresponding to the first score information, calculating a second range overlap between the first range and the range corresponding to the second score information, and calculating a third range overlap between the first range and the range corresponding to the third score information;

[0139] Determine the maximum sound range overlap among the first sound range overlap, the second sound range overlap and the third sound range overlap; determine the music score information corresponding to the maximum sound range overlap as the target music score information, wherein the music score information corresponding to the maximum sound range overlap is the first music score information, the second music score information or the third music score information.

[0140] According to one aspect of the present application, an embodiment of the present application further provides a computer program product, the computer program product comprising a computer program, the computer program being stored in a computer-readable storage medium. The processor 701 reads the computer program from the computer-readable storage medium, and the processor 701 executes the computer program, so that the computer device 700 executes Figure 2 as well as Figure 4 Singing voice synthesis method.

[0141] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic, and the division of the modules described above is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0143] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A singing voice synthesis method, characterized in that: The method comprises: Acquire singing data of a target object, and perform range detection on the singing data to obtain a first range of the target object; acquire initial score information of a song to be synthesized, and acquire a second range of the song to be synthesized according to the initial score information; Determine a pitch span value of the target object according to the first sound range; Determine a cross pitch span value between the target object and the song to be synthesized according to the first range and the second range; Determining a comparison result of the degree of overlap between the first range and the second range according to the cross pitch span value and the pitch span value of the target object; If the comparison result of the degree of coincidence indicates that the first range and the second range do not match, adjusting the pitch interval of the initial music score information to obtain target music score information; The acoustic synthesis model associated with the target object is called so that the acoustic synthesis model performs singing data synthesis processing according to the target music score information to obtain synthesized singing data associated with the target object and the song to be synthesized.

2. The method according to claim 1, characterized in that The step of determining the cross pitch span value between the target object and the song to be synthesized according to the first sound range and the second sound range includes: Determine the upper limit value of the pitch of the target object according to the first sound range, and determine the upper limit value of the pitch of the song to be synthesized according to the second sound range; If the pitch upper limit value of the song to be synthesized is less than the pitch upper limit value of the target object, determining the difference between the pitch upper limit value of the song to be synthesized and the pitch lower limit value of the target object as the crossover pitch span value; If the pitch upper limit value of the song to be synthesized is greater than or equal to the pitch upper limit value of the target object, the difference between the pitch upper limit value of the target object and the pitch lower limit value of the song to be synthesized is determined as the crossover pitch span value.

3. The method according to claim 1, characterized in that The step of determining a comparison result of the degree of overlap between the first range and the second range according to the cross pitch span value and the pitch span value of the target object includes: A ratio between the crossover pitch span value and the pitch span value of the target object is determined, and a result of a comparison of the overlap between the first range and the second range is determined according to the ratio.

4. The method according to any one of claims 1 to 3, characterized in that: The step of adjusting the pitch interval of the initial music score information to obtain target music score information includes: Determine the average pitch of the target object according to the first sound range, determine the average pitch of the song to be synthesized according to the second sound range; and determine a deviation parameter according to the average pitch of the target object and the average pitch of the song to be synthesized; The pitch information contained in the initial music score information is added to the deviation parameter to obtain the target music score information.

5. The method according to any one of claims 1 to 3, characterized in that: The step of adjusting the pitch interval of the initial music score information to obtain target music score information includes: Acquire a reference range, wherein the reference range includes a bass range, a middle range, and a treble range; Determine a first pitch adjustment parameter according to the lower limit value of the pitch of the bass range and the lower limit value of the pitch of the song to be synthesized, determine a second pitch adjustment parameter according to the lower limit value of the pitch of the middle range and the lower limit value of the pitch of the song to be synthesized, and determine a third pitch adjustment parameter according to the lower limit value of the pitch of the treble range and the lower limit value of the pitch of the song to be synthesized; The pitch interval of the initial music score information is adjusted according to the first pitch adjustment parameter, the second pitch adjustment parameter, and the third pitch adjustment parameter to obtain target music score information.

6. The method according to claim 5, characterized in that The step of adjusting the pitch interval of the initial music score information according to the first pitch adjustment parameter, the second pitch adjustment parameter, and the third pitch adjustment parameter to obtain target music score information includes: Adding the pitch information contained in the initial music score information to the first pitch adjustment parameter to obtain first music score information, adding the pitch information contained in the initial music score information to the second pitch adjustment parameter to obtain second music score information, and adding the pitch information contained in the initial music score information to the third pitch adjustment parameter to obtain third music score information; Target music score information is determined from the first music score information, the second music score information, and the third music score information.

7. The method according to claim 6, characterized in that The determining target music score information from the first music score information, the second music score information and the third music score information comprises: Calculating a first range overlap between the first range and the range corresponding to the first score information, calculating a second range overlap between the first range and the range corresponding to the second score information, and calculating a third range overlap between the first range and the range corresponding to the third score information; Determine the maximum sound range overlap among the first sound range overlap, the second sound range overlap and the third sound range overlap; determine the music score information corresponding to the maximum sound range overlap as the target music score information, wherein the music score information corresponding to the maximum sound range overlap is the first music score information, the second music score information or the third music score information.

8. A computer device, characterized in that: The computer device comprises: a processor adapted to implement one or more computer programs; and, A computer storage medium storing one or more computer programs, wherein the one or more computer programs are suitable for being loaded by the processor and executing the singing synthesis method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more computer programs, and the one or more computer programs are suitable for being loaded by a processor and executing the singing synthesis method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Song synthesis method, device and equipment and storage medium

    CN111681637A

  • Information processing apparatus, electronic musical instrument, information processing system, information processing method, and storage medium

    CN115116414A