Singing matching method and device, electronic equipment, storage medium and program product
By acquiring the singer's training set and extracting features, the matching degree between the singer and the song is automatically determined, solving the problems of single matching degree and poor accuracy in existing technologies, and achieving a more efficient and accurate singing conversion effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-01
- Publication Date
- 2026-05-08
AI Technical Summary
In existing technologies, the methods for determining the matching degree between singers and songs are relatively simple and inaccurate, resulting in poor vocal conversion effects.
By acquiring the training set of singers, the features of each singing content sample under the target singing attributes are extracted, including vocal range, musical style and timbre features. Based on these features, the matching degree between the singer and the content to be sung is determined. Using preset detection technology and feature fusion method, the singing matching degree is automatically determined.
It enriches the methods for determining the matching degree of singing, improves the accuracy of the matching degree, reduces the time and human error of manual listening and annotation, and enhances the effect of singing conversion.
Smart Images

Figure CN121997055A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to a singing matching method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] In music technology, Song Voice Conversion (SVC) is a product of the combination of artificial intelligence and computer technology, used to convert the vocals of a song into the vocals of another singer.
[0003] However, in related technologies, the method for determining the matching degree between singers and songs during vocal conversion is relatively simple, and the accuracy of the determined matching degree is poor. Summary of the Invention
[0004] This disclosure provides a singing matching method, apparatus, electronic device, storage medium, and program product to enrich the ways of determining singing matching degree and improve the accuracy of the determined singing matching degree.
[0005] In a first aspect, embodiments of this disclosure provide a singing matching method, including:
[0006] Obtain the training set of the singer, and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer;
[0007] Based on the first singing features of each singing content sample, the second singing features of the singer under the target singing attribute are determined;
[0008] Based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute, the singing matching degree between the singer and the content to be sung is determined.
[0009] Secondly, embodiments of this disclosure also provide a singing matching device, comprising:
[0010] The feature extraction module is used to obtain the training set of the singer and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer;
[0011] The feature determination module is used to determine the second singing feature of the singer under the target singing attribute based on the first singing feature of each singing content sample;
[0012] The matching degree determination module is used to determine the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, including:
[0014] One or more processors;
[0015] Memory, used to store one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the singing matching method as described in the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the singing matching method as described in embodiments of this disclosure.
[0018] Fifthly, embodiments of this disclosure also provide a computer program product that, when executed by a computer, causes the computer to implement the singing matching method as described in embodiments of this disclosure.
[0019] The singing matching method, apparatus, electronic device, storage medium, and program product provided in this disclosure acquire a training set of singers and extract first singing features of each singing content sample in the training set under a target singing attribute. The training set is the set of singing content samples used when training the singer. Based on the first singing features of each singing content sample, a second singing feature of the singer under the target singing attribute is determined. Based on the second singing feature of the singer under the target singing attribute and the third singing feature of the content to be sung under the target singing attribute, the singing matching degree between the singer and the content to be sung is determined. This disclosure utilizes the above technical solution to determine the singer's singing features based on the singing features of each singing content sample in the singer's training set. This eliminates the need for manual listening and annotation, enriching the methods for determining the singer's singing features, reducing the time spent on determining the singer's singing features, and improving the accuracy of the determined singing features. This, in turn, enriches the methods for determining the singing matching degree between the singer and the content to be sung, reduces the time spent on determining the singing matching degree between the singer and the content to be sung, and improves the accuracy of the determined singing matching degree. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A schematic flowchart of a singing matching method provided in an embodiment of this disclosure;
[0022] Figure 2 A flowchart illustrating another singing matching method provided in this embodiment of the present disclosure;
[0023] Figure 3 A schematic diagram illustrating an SVC singer matching process provided in an embodiment of this disclosure;
[0024] Figure 4 A pitch frequency distribution diagram of an SVC singer provided in this embodiment of the present disclosure;
[0025] Figure 5 A schematic diagram illustrating the process of determining the matching degree of a musical style, provided in an embodiment of this disclosure;
[0026] Figure 6 A schematic diagram illustrating the process of determining timbre matching degree according to an embodiment of this disclosure;
[0027] Figure 7 A structural block diagram of a singing matching device provided in an embodiment of this disclosure;
[0028] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0031] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0034] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0035] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0036] Figure 1 This is a flowchart illustrating a singing matching method provided in an embodiment of this disclosure. The method can be executed by a singing matching device, which can be implemented in software and / or hardware and can be configured in an electronic device, typically a computer, mobile phone, or tablet computer. The singing matching method provided in this embodiment is applicable to scenarios where the singing matching degree between a singer and the content to be sung is determined, such as when determining the singing matching degree between a singer and the content to be sung during vocal conversion.
[0037] In SVC (Single Voice Capture) scenarios, poor vocal quality is a frequent issue when generating singing voices. Taking vocal range characteristics as an example, problems such as being unable to reach high notes, muteness, voice cracking, strained vocals, and distortion may occur during vocal generation, especially when a male singer performs a female song or vice versa. This problem is often caused by a mismatch between the vocal ranges of the SVC virtual singer and the target song. For example, if the singer's vocal range is E2 to A#4, while the song's range is G3 to C6, a vocal range mismatch will occur.
[0038] In related technologies, obtaining the singing characteristics of an SVC singer requires controlling the SVC singer to sing a large number of songs and manually reviewing the performance to label the singer's singing characteristics. For example, to obtain the vocal range information of an SVC singer, it is necessary to control the SVC singer to sing a large number of songs and manually review the performance, such as manually checking whether the pronunciation of each note is normal, thereby determining the singer's vocal range. Singer matching is then performed based on the labeled singing characteristics.
[0039] However, the methods for determining the matching degree between singers and songs in related technologies are relatively simple, and the accuracy of the determined matching degree is poor.
[0040] In view of this, the present disclosure provides a singing matching method to enrich the determination of the singing matching degree between singers and songs, and to improve the accuracy of the determined singing matching degree.
[0041] like Figure 1 As shown, the singing matching method provided in this embodiment may include:
[0042] S101. Obtain the training set of the singer and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer.
[0043] In this context, the singer can be understood as the object whose performance matching degree with the content to be sung is to be calculated. In some examples, this singer may include virtual singers and / or non-virtual singers. The singer can be trained based on its training set. The singer's training set is the set of training samples used to train this singer. The training set may contain multiple performance content samples. These performance content samples are the samples used to train the singer's performance effect. For example, these performance content samples can be dry vocals of the same object or different objects singing different performance content. Performance content can be understood as the content being sung, such as songs and / or operas. Dry signal can be understood as the unprocessed raw vocal sound, such as human voice. This sound has not been processed by reverb, delay, chorus, or other effects, maintaining the purity and authenticity of the recording. Because it can provide the clearest audio signal, dry signal is often used in the recording and mixing process.
[0044] The target singing attribute can be understood as the singing attribute to be matched. In a singing conversion scenario, the target singing attribute can be an attribute related to the conversion effect. The target singing attribute can be preset by developers or users as needed. For example, the target singing attribute may include at least one of the following: vocal range attribute, musical style attribute, timbre attribute, and singing tempo attribute. There can be one or more target singing attributes. When there are multiple target singing attributes, the singing matching method provided in this embodiment can be executed for each target singing attribute separately. The first singing feature can be understood as the singing feature of the singing content sample under the target singing attribute. For example, the first singing feature can be the attribute feature of the target singing attribute of the singing content sample, such as the vocal range feature, musical style feature, timbre feature, and / or singing tempo feature of the singing content sample.
[0045] Specifically, a training set of singers can be obtained, and for each performance sample in the training set, the first performance feature under the target performance attribute can be extracted. The method for extracting the first performance feature is not limited; when the target performance attributes are different, the extraction methods for the first performance feature can be the same or different.
[0046] For example, when the target singing attribute is the range attribute, a preset pitch detection technique can be used to detect the note information in the singing content sample, serving as the first singing feature (i.e., range feature) of the corresponding singing content sample under the range attribute. The preset pitch detection technique can be flexibly set as needed. For instance, the preset pitch detection technique can be a basic-pitch detection technique. In this case, the basic-pitch technique can be used to detect the pitch information at various positions in the vocal part (such as dry vocals), such as detecting and determining the start time, end time, and corresponding pitch of each note in each singing content sample. This note pitch can be represented using a musical code (MIDI number) in the Musical Instrument Digital Interface (MIDI) protocol, which corresponds one-to-one with the note and can be used to identify this note. For example, when the MIDI number is 60, its corresponding note is C4 (i.e., middle C).
[0047] For example, when the target singing attribute is a genre attribute, a preset genre detection technique can be used to detect the genre characteristics of each singing content sample, which serves as its first singing feature under the genre attribute. The preset genre detection technique can be flexibly configured as needed. The genre attribute can include multiple attribute types (i.e., attribute values), such as multiple genre types. Taking a song as an example, the genre type could be pop, folk, rock, and / or metal, etc. The genre feature of a singing content sample can be considered as the type matching degree between the singing content sample and each attribute type of the genre attribute.
[0048] For example, when the target singing attribute is timbre, a preset voiceprint detection method can be used to detect the voiceprint features of each singing content sample, which can then be used as its first singing feature under the timbre attribute. The preset voiceprint detection method can be flexibly set as needed; for example, it can use a pre-trained voiceprint detection model to detect the voiceprint features of each singing content sample. This embodiment does not limit this approach.
[0049] S102. Determine the second singing feature of the singer under the target singing attribute based on the first singing feature of each singing content sample.
[0050] The second singing feature can be understood as the singing feature of the singer under the target singing attribute. For example, the second singing feature can be the attribute feature of the target singer's target singing attribute, such as the target singer's vocal range feature, musical style feature, timbre feature and / or singing speed feature, etc.
[0051] For example, after determining the first singing features of each singing content sample in the training set of the singer, the second singing features of the singer under the target singing attribute can be determined based on the first singing features of each singing content sample in the training set. For example, the first singing features of each singing content sample can be fused, and the fused singing features can be used as the second singing features of the singer under the target singing attribute, and so on.
[0052] It should be noted that the execution timing of S101-S102 is not limited. For example, it can be executed for a singer when there is content to be sung, such as when it is necessary to determine the singing matching degree between a singer and the content to be sung. Alternatively, S101-S102 can be executed separately for each singer in advance. For example, for each singer, after the training set of the singer is constructed, such as before training the singer, during training the singer, or after training the singer, S101-S102 can be executed in advance for the singer to obtain the second singing feature of the singer under the target singing attribute and store the second singing feature. Thus, when there is content to be sung, the matching degree between the singer and the content to be sung can be determined based on the pre-stored second singing feature of the singer.
[0053] S103. Determine the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute.
[0054] In this context, "content to be sung" can be understood as the content to be sung. Taking the SVC scenario as an example, the content to be sung can be the content that needs to be voice-converted, such as the song or opera that needs to be voice-converted. "Third singing feature" can be understood as the singing characteristics of the content to be sung under the target singing attribute. For example, the third singing feature can be the attribute features of the target singing attribute of the content to be sung, such as the vocal range, style, timbre, and / or tempo of the content to be sung. "Singing matching degree" can be understood as the degree of matching between the singer's singing characteristics under the target singing attribute and the singing characteristics of the content to be sung under the target singing attribute.
[0055] Specifically, when there is content to be sung, such as when a matching degree determination operation for the content to be sung is received, the third singing feature of the content to be sung under the target singing attribute can be obtained. Based on the second singing feature of the singer under the target singing attribute and the third singing feature of the content to be sung under the target singing attribute, the singing matching degree between the singer and the content to be sung can be determined. For example, the similarity between the second singing feature of the singer under the target singing attribute and the third singing feature of the content to be sung under the target singing attribute can be calculated as the singing matching degree between the singer and the content to be sung, and so on.
[0056] In some implementations, after determining the degree of matching between the singer and the content to be sung, this degree of matching can be displayed so that users can view and / or determine whether to use this singer to perform the content to be sung based on this degree of matching.
[0057] In other implementations, after determining the matching degree between the singer and the content to be sung, it is possible to determine whether to use the singer to sing the content based on the pre-set singer selection rules and the matching degree. For example, based on the matching degree between multiple singers and the content to be sung under the target singing attributes, at least one singer can be selected from these multiple singers to sing the content to be sung.
[0058] Optionally, the singer is a candidate singer from the singer set. After determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute, the method further includes: selecting at least one candidate singer from the singer set based on the singing matching degree as the target singer of the content to be sung; and performing voice conversion on the original singing data of the content to be sung based on the target singer to obtain the target singing data of the content to be sung.
[0059] The singer set can be a collection of candidate singers. The singer set may include at least one pre-trained candidate singer. The target singer can be a candidate singer determined to perform the content to be sung. The original singing data can be the singing data of the content to be sung before sound conversion; for example, the original singing data may include the accompaniment of the content to be sung and the original dry vocals before conversion. This original dry vocals can be sung by a singer other than the target singer. The target singing data can be the singing data obtained after sound conversion of the content to be sung; for example, the target singing data may include the accompaniment of the content to be sung and the converted target dry vocals. This target dry vocals can be sung by the target singer. Sound conversion can be understood as converting the vocal portion of the singing content, which may include, but is not limited to, vocal conversion.
[0060] Specifically, the matching degree between each candidate singer and the content to be sung under the target singing attributes can be determined based on the second singing characteristics of each candidate singer in the singer set and the third singing characteristics of the content to be sung.
[0061] After obtaining the matching degree between each candidate singer in the singer set and the content to be sung under the target singing attributes, at least one candidate singer can be selected from the singer set as the target singer for the content to be sung, based on the matching degree between each candidate singer and the content to be sung under the target singing attributes. For example, candidate singers whose matching degree meets a set condition can be selected from the singer set, such as selecting the M candidate singers with the highest matching degree, or selecting candidate singers whose matching degree is greater than a preset matching degree threshold, etc., as the target singers for the content to be sung. Here, M is a positive integer, and its specific value is not limited; for example, M can be set to 1, 3, or 5, etc.; the preset matching degree threshold can be flexibly set as needed, such as setting the preset matching degree threshold to 0.8, 0.9, or 0.95, etc.
[0062] After determining the target singer of the content to be sung, the voice of the content to be sung can be converted based on this target singer. For example, after determining the target singer of the content to be sung, the voice of the content to be sung can be converted automatically or in response to the user's voice conversion operation, based on the target singer, to obtain the target singing data of the content to be sung.
[0063] When performing voice conversion, for example, the raw dry voice can be extracted from the original singing data of the content to be sung, and based on singing conversion technology, this raw dry voice can be converted into the target dry voice sung by the target singer. This target dry voice is then used to replace the raw dry voice in the original singing data of the content to be sung, thereby obtaining the target singing data of the content to be sung.
[0064] The singing matching method provided in this embodiment obtains a training set of the singer and extracts the first singing features of each singing content sample in the training set under the target singing attribute. This training set is the set of singing content samples used when training the singer. Based on the first singing features of each singing content sample, the second singing features of the singer under the target singing attribute are determined. Based on the second singing features of the singer under the target singing attribute and the third singing features of the content to be sung under the target singing attribute, the singing matching degree between the singer and the content to be sung is determined. This embodiment utilizes the above technical solution to determine the singer's singing features based on the singing features of each singing content sample in the singer's training set. This eliminates the need for manual listening and annotation, enriching the methods for determining the singer's singing features, reducing the time spent on determining the singer's singing features, and improving the accuracy of the determined singing features. This, in turn, enriches the methods for determining the singing matching degree between the singer and the content to be sung, reduces the time spent on determining the singing matching degree between the singer and the content to be sung, and improves the accuracy of the determined singing matching degree.
[0065] In some embodiments, the target singing attribute may include a genre attribute. The attribute value of the genre attribute can be the attribute type of the genre attribute, such as a genre type. This genre type can be, for example, pop, folk, rock, and / or metal, etc. Optionally, the target singing attribute includes a genre attribute, which includes multiple attribute types, and the second singing feature is the type matching degree between the singer and each of the attribute types. The type matching degree can be understood as the matching degree between the singer and a certain attribute type of the target singing attribute, which can be used to characterize the probability that the singer is suitable for singing content of the corresponding attribute type, i.e., the probability that the singer is suitable for singing content of the corresponding genre type. When the target singing attribute is a genre attribute, this type matching degree can be the matching degree between the target singer and each genre type.
[0066] There are no restrictions on how the type match between a singer and a certain attribute type is determined. For example, for each attribute type, the arithmetic mean of the type match between each performance content sample in the singer's training set and that attribute type can be calculated as the type match between the singer and that attribute type.
[0067] After obtaining the type matching degree between the singer and each attribute type of the target singing attribute, the singing matching degree between the singer and the content to be sung under each attribute type can be calculated based on the type matching degree between the singer and each attribute type of the target singing attribute, and the type matching degree between the content to be sung and each attribute type of the target singing attribute. For example, for each type matching degree of the target singing attribute, the geometric mean between the type matching degree between the singer and this attribute type and the type matching degree between the content to be sung and this attribute type can be calculated as the singing matching degree between the singer and the content to be sung under this attribute type. Optionally, determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes: for each attribute type of the target singing attribute, calculating the geometric mean between the first type matching degree of the singer and the second type matching degree of the content to be sung as the singing matching degree between the singer and the content to be sung under the attribute type, wherein the first type matching degree is the type matching degree between the singer and the attribute type, and the second type matching degree is the type matching degree between the content to be sung and the attribute type. Here, the first type of matching degree can be understood as the type matching degree between the singer and the current attribute type of the target singing attribute; the second type of matching degree can be understood as the type matching degree between the content to be sung and the current attribute type of the target singing attribute.
[0068] In some embodiments, the target singing attribute includes a timbre attribute, and determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes: calculating the similarity between the second singing feature and the third singing feature of the content to be sung under the target attribute, as the singing matching degree between the singer and the content to be sung under the target singing attribute.
[0069] In this embodiment, the target singing attribute may include timbre attribute to further enrich the comprehensiveness of the singing matching degree and improve the accuracy of the determined singing matching degree between the singer and the content to be sung under the timbre attribute.
[0070] When the target singing attribute is a timbre attribute, the first singing feature and / or the second singing feature under the target singing attribute can be a timbre feature, such as a voiceprint feature; the singing matching degree between the singer and the content to be sung under the target singing attribute can be, for example, the similarity between the singer's timbre feature and the timbre feature of the content to be sung.
[0071] Specifically, when the target singing attribute is timbre, the similarity between the singer's second singing feature under the timbre attribute and the third singing feature of the content to be sung under the target attribute can be calculated. For example, by comparing voiceprints, the similarity between the singer's voiceprint features and the voiceprint features of the content to be sung can be determined as the singing matching degree between the singer and the content to be sung under the timbre attribute.
[0072] Figure 2 This is a flowchart illustrating another singing matching method provided in an embodiment of this disclosure. The scheme in this embodiment can be combined with one or more optional schemes in the above embodiments. Optionally, the target singing attribute includes a vocal range attribute, and determining the singer's second singing feature under the target singing attribute based on the first singing feature of each singing content sample includes: according to the first singing feature of each singing content sample, counting the first occurrence frequency value of each attribute value of the target singing attribute in each singing content sample; generating an occurrence frequency curve of the target singing attribute based on the first occurrence frequency value; and determining the range of the first attribute value of the singer under the target singing attribute based on the derivative value of each attribute value in the occurrence frequency curve, as the singer's second singing feature under the target singing attribute.
[0073] Correspondingly, such as Figure 2 As shown, the singing matching method provided in this embodiment may include:
[0074] S201. Obtain the training set of the singer and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer, and the target singing attribute includes the vocal range attribute.
[0075] Taking the vocal range attribute as an example, after obtaining the singer's training set, a preset pitch detection technique can be used to detect the note information of each note in the vocal range sample for each vocal content sample in the training set. For example, the note start time, note end time, and note pitch of each note in the vocal range sample can be detected as the vocal range feature (i.e., the first vocal feature) of the vocal range sample under the vocal range attribute.
[0076] S202. Based on the first singing feature of each singing content sample, count the first occurrence count of each attribute value of the target singing attribute in each singing content sample.
[0077] When the target singing attribute is a range attribute value, the attribute value can be different musical notes. The first occurrence count value can be understood as the total number of times each note appears in each singing content sample.
[0078] Specifically, after extracting the note information of each note in each singing content sample, the total number of times each note appears in each singing content sample can be counted, which is used as the first occurrence value of this note in each singing content sample.
[0079] S203. Generate the occurrence frequency curve of the target singing attribute based on the first occurrence frequency value.
[0080] When the target singing attribute is a vocal range attribute, the occurrence frequency curve of the target singing attribute can be understood as a curve representing the total number of occurrences of each note in each singing content sample. For example, the horizontal axis of this occurrence frequency curve can be a note or note pitch, and the vertical axis of this occurrence frequency curve can be the first occurrence frequency value.
[0081] For example, after statistically obtaining the first occurrence count of each note in each singing content sample, a note occurrence count curve of the range attribute can be generated based on this first occurrence count. For instance, a distribution map of the occurrence count of each note can be drawn based on the first occurrence count of each note in each singing content sample, and a note occurrence count curve of the range attribute can be generated based on this occurrence count distribution map, which serves as the occurrence count curve of the range attribute.
[0082] In some implementations, after statistically obtaining the first occurrence count of each note in each singing content sample, it is not necessary to perform data cleaning on this first occurrence count. Instead, the occurrence count curve of the vocal range attribute can be directly generated based on the statistically obtained first occurrence count of each note. This improves the generation speed of the occurrence count curve and thus improves the speed of determining the singer's singing characteristics under the vocal range attribute.
[0083] In other implementations, after statistically obtaining the first occurrence count of each note in each singing content sample, data cleaning can be performed on this first occurrence count, and an occurrence count curve of the vocal range attribute can be generated based on the data-cleaned first occurrence count to improve the accuracy of the generated occurrence count curve, thereby improving the accuracy of the determined singing characteristics of the singer under the vocal range attribute. Optionally, before generating the occurrence count curve of the target singing attribute based on the first occurrence count, the method further includes: determining the valid occurrence count and invalid occurrence count among each of the first occurrence counts based on a preset attribute value range, wherein the attribute value corresponding to the valid occurrence count is within the preset attribute value range, and the attribute value corresponding to the invalid occurrence count is outside the preset attribute value range; obtaining the target occurrence count value among each of the invalid occurrence counts that meets a preset condition, and determining the valid occurrence count value less than the target occurrence count value as an invalid occurrence count value; generating the occurrence count curve of the target singing attribute based on the first occurrence count includes: generating the occurrence count curve of the target singing attribute based on the remaining valid occurrence count values.
[0084] The preset attribute value range can be a pre-defined range of attribute values. In some examples, when the target singing attribute is a vocal range attribute, this preset attribute value range can be set based on the vocal range that humans can sing. For example, this preset attribute value range can be the vocal range that humans can sing. For instance, the preset attribute range can be set to E2~C6 (i.e., MIDI number 40~84). The valid occurrence count value can be the first occurrence count value corresponding to the attribute value (such as a note) within the preset attribute value range. The invalid occurrence count value can be the first occurrence count value corresponding to the attribute value outside the preset attribute value range. The target occurrence count value can be understood as the occurrence count value that meets the preset conditions. This preset condition is not limited. For example, this preset condition can be set to the maximum occurrence count value, the minimum occurrence count value, the arithmetic mean or median of the invalid occurrence count values, etc. In this case, for instance, the target occurrence count value can be the maximum, minimum, arithmetic mean or median of the invalid occurrence count values.
[0085] For example, after statistically obtaining the first occurrence count of each note (i.e., attribute value) in each singing content sample, notes within a preset attribute value range can be initially marked as valid notes, and the first occurrence count corresponding to the valid notes can be initially marked as valid occurrence count values; notes outside the preset attribute value range can be marked as invalid notes, and the first occurrence count corresponding to the invalid notes can be determined as invalid occurrence count values. Invalid occurrence count values that meet preset conditions are obtained from each invalid occurrence count value as target occurrence count values. Valid occurrence count values less than this target occurrence count value are obtained from each valid occurrence count value and marked as invalid occurrence count values. After marking, each invalid occurrence count can be cleared, and based on the remaining valid occurrence count values and the valid notes corresponding to the remaining valid occurrence count values, a note occurrence count curve for the range attribute can be generated as the occurrence count curve for the target singing attribute.
[0086] S204. Based on the derivative values of each attribute value in the occurrence curve, determine the range of first attribute values of the singer under the target singing attribute, and use it as the second singing feature of the singer under the target singing attribute.
[0087] The first attribute value range can be understood as the range of attribute values of the singer under the target singing attribute. Taking the target singing attribute as the vocal range attribute as an example, the first attribute range can be the vocal range range of the singer.
[0088] After generating the occurrence frequency curve of the target singing attribute, the derivative of this occurrence frequency curve can be obtained to obtain the derivative value of each attribute value in this occurrence frequency curve. Based on the derivative value of each attribute value, the attribute range of the singer under the target singing attribute can be determined as the second singing feature of the singer under the target singing attribute.
[0089] In some implementations, considering that when the target singing attribute is a vocal range attribute, the frequency of occurrence of each note within the singer's vocal range is generally a normal distribution, the note whose frequency of occurrence rises faster in the frequency curve can be taken as the lowest note in the singer's vocal range, and the note whose frequency of occurrence falls faster in the frequency curve can be taken as the highest note in the singer's vocal range, thereby determining the singer's note range and further improving the accuracy of the determined vocal range.
[0090] Optionally, determining the range of first attribute values for the singer under the target singing attribute based on the derivative values at each attribute value in the occurrence frequency curve includes: obtaining the minimum and maximum attribute values in the occurrence frequency curve where the absolute value of the derivative is greater than or equal to a preset threshold, based on the derivative values at each attribute value in the occurrence frequency curve; and determining the range of first attribute values for the singer under the target singing attribute based on the minimum and maximum attribute values. Here, the absolute value of the derivative can be the absolute value of the derivative.
[0091] For example, after determining the derivative values of each attribute value in the occurrence frequency curve, such as after generating the derivative curve of the occurrence frequency curve, the minimum attribute value (e.g., the lowest note) whose absolute derivative value is greater than or equal to a preset threshold can be obtained as the starting point of the first attribute value range; and the maximum attribute value (e.g., the highest note) whose absolute derivative value is greater than or equal to the preset threshold can be obtained as the ending point of the first attribute value range. Thus, the first attribute value range of the singer under the target singing attribute can be obtained. The preset threshold can be flexibly set as needed, and this embodiment does not limit it.
[0092] S205. Based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute, determine the singing matching degree between the singer and the content to be sung.
[0093] In this embodiment, taking the vocal range attribute as the target singing attribute as an example, the singing matching degree between the singer and the content to be sung under the vocal range attribute can be determined based on the singer's vocal range range (i.e., the second singing feature) and the vocal range range of the content to be sung (i.e., the third singing feature).
[0094] In this embodiment, the method for determining the singing match degree between the singer and the content to be sung under the vocal range attribute is not limited. For example, the overlapping portion between the singer's vocal range and the vocal range of the content to be sung can be obtained, and the proportion of this overlapping portion within the vocal range of the content to be sung can be calculated as the singing match degree between the singer and the content to be sung under the vocal range attribute (i.e., vocal range match degree). Optionally, the third singing feature is the second attribute value range of the content to be sung under the target singing attribute. Determining the singing match degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes: obtaining the sub-attribute value range overlapping between the first attribute value range and the second attribute value range; calculating the proportion of attribute values within the sub-attribute value range within the second attribute value range as the singing match degree between the singer and the content to be sung under the target singing attribute.
[0095] The second attribute value range can be understood as the range of attribute values for the content to be sung under the target singing attribute. Taking the target singing attribute as the vocal range attribute as an example, the second attribute value range can be the vocal range of the content to be sung. The sub-attribute value range can be the attribute value range corresponding to the overlapping part between the first attribute value range and the second attribute value range.
[0096] The calculation method for the proportion of attribute values within the sub-attribute value range within the second attribute value range is not limited. For example, the ratio between the interval length corresponding to the sub-attribute value range and the interval length corresponding to the second attribute value range can be calculated as the proportion of attribute values within the sub-attribute value range within the second attribute value range; or the ratio between the number of attribute values (such as the number of musical notes) contained in the sub-attribute value range and the number of attribute values contained in the second attribute value range can be calculated as the proportion of attribute values within the sub-attribute value range within the second attribute value range, and so on.
[0097] The method for determining the range of the second attribute values of the content to be sung is not limited. For example, the range of the second attribute values of the content to be sung can be determined based on the frequency of occurrence of each attribute value of the target singing attribute in the original dry vocal data of the content to be sung. Optionally, before obtaining the overlapping sub-attribute value ranges between the first and second attribute value ranges, the method further includes: extracting the dry vocal data of the content to be sung from the original vocal data, and statistically analyzing the second frequency of occurrence of each attribute value of the target singing attribute in the content to be sung based on the dry vocal data; and determining the range of the second attribute values of the content to be sung under the target singing attribute based on the second frequency of occurrence.
[0098] The second occurrence count value can be understood as the number of times each attribute value of the target singing attribute (such as musical notes) appears in the original dry vocal data of the content to be sung.
[0099] Specifically, the raw vocal data of the content to be sung can be stripped of its raw vocals. For example, the raw vocal data can be separated into vocal and accompaniment. The separated vocals can then be de-harmonicized and de-reverbized to obtain the raw vocal data of the content to be sung. After obtaining the raw vocal data, a preset pitch detection technique can be used to detect the note information of each note in the vocal data. Based on this note information, the second occurrence count of each note in the vocal data can be counted, and the range of the second attribute value of the content to be sung under the target singing attribute can be determined based on this second occurrence count.
[0100] When determining the range of the second attribute value of the content to be sung under the target singing attribute based on the second occurrence count value, for example, the minimum attribute value (such as the lowest note) and the maximum attribute value (such as the highest note) in the original dry vocal data of the content to be sung, which have an occurrence count value greater than or equal to a preset occurrence count threshold, can be directly used as the two endpoints of the range of the second attribute value. This will give the range of the second attribute value of the content to be sung under the target singing attribute, thereby increasing the generation speed of the occurrence count curve and thus increasing the speed of determining the singer's singing characteristics under the vocal range attribute.
[0101] Furthermore, after statistically obtaining the first occurrence count of each note in each sample of singing content, the second occurrence count can be cleaned, and the range of the second attribute value of the content to be sung under the target singing attribute can be determined based on the cleaned second occurrence count. This improves the accuracy of the generated occurrence count curve and, consequently, the accuracy of the determined singing characteristics of the singer under the vocal range attribute. For example, the second occurrence counts whose corresponding attribute values are within a preset attribute value range can be initially marked as valid occurrence counts, while those whose corresponding attribute values are outside the preset attribute value range can be marked as invalid occurrence counts. The occurrence counts that meet preset conditions (such as the maximum occurrence count) among the invalid occurrence counts are obtained. The valid occurrence counts that are less than the occurrence counts that meet the preset conditions are remarked as invalid occurrence counts. The minimum and maximum attribute values corresponding to the remaining valid occurrence counts after remarking are used as the two endpoints of the second attribute value range, thereby obtaining the range of the second attribute value of the content to be sung under the target singing attribute.
[0102] The singing matching method provided in this embodiment can automatically determine the vocal range, reduce the time spent determining the vocal range, enrich the methods for determining the vocal range matching degree, and improve the accuracy of the determined vocal range matching degree.
[0103] Figure 3 This disclosure provides a schematic diagram of a singer matching process according to an embodiment. For example... Figure 3 As shown, in some optional embodiments, taking the target singing attribute as the vocal range attribute and the singer as an SVC singer as an example, the vocal range matching process of the SVC singer in the SVC scenario can be described as follows:
[0104] A1. Determine the vocal range (i.e., the second singing characteristic) of the SVC singer.
[0105] Specifically, a training set of SVC singers can be obtained; pitch detection technology can be used to detect the pitch of each training song (i.e., the singing content sample) in the training set to obtain the pitch information of each training song; based on the obtained pitch information, the total number of occurrences of each note (i.e., attribute value) in each training song can be counted (i.e., the first occurrence value); the total occurrence value of each note can be cleaned to obtain the singer's vocal range (i.e., the second singing feature).
[0106] The training set can contain multiple training songs, each of which can be dry audio data. Pitch information for the vocal parts can be extracted using basic-pitch. The main pitch information can include the note start time (star_time_s), note end time (end_time_s), and note pitch (pitch_midi). The unit for note pitch (pitch_midi) can be a MIDI number, for example, MIDI number 60 = Note name C4 (middle C).
[0107] In the notes output by basic pitch, such as Figure 4 As shown, due to the presence of noise, pitch may be distributed over a wide range. Considering that the typical vocal range of humans (i.e., the preset attribute range) is generally E2 to C6 (MIDI number 40 to MIDI number 84), notes outside this range can be considered invalid noise data caused by overtones and can be discarded. After identifying the notes outside this range, the frequency of occurrence of the note that appears most frequently outside the range (i.e., meets the preset conditions) can be used as a filtering threshold. For notes within the range, frequencies below this filtering threshold are removed. After data removal, a note distribution curve (i.e., frequency curve) for the SVC singer can be generated based on the remaining frequency values. The highest and lowest notes whose absolute derivative values are greater than or equal to the preset threshold are taken as the two endpoints of the SVC singer's vocal range, thus determining the SVC singer's vocal range.
[0108] A2. Determine the vocal range (i.e., the third singing characteristic) of the target song (i.e., the content to be sung).
[0109] Specifically, the dry vocal data of the target song can be obtained by stripping the dry vocal data; pitch detection technology can be used to detect the pitch of the raw dry vocal data of the target song to obtain the pitch information of the target song; based on the obtained pitch information, the total number of occurrences of each note in the target song (i.e., the second occurrence value) can be counted; and the total occurrence value of each note can be cleaned to obtain the range of the target song.
[0110] The method for pitch detection of the target song is similar to that for pitch detection of the training song, and will not be elaborated here.
[0111] For the target song, the vocal parts need to be separated (i.e., dry vocal stripping) and input into the basic-pitch input to extract pitch information. For example, a three-model stripping technique can be used, sequentially passing the vocal-accompaniment separation model, the harmonic de-shaping model, and the reverberation de-shaping model to obtain the clean vocal dry part of the target song.
[0112] During data cleaning, for example, notes outside the range of human vocal range can be identified, and the frequency of the note occurring most frequently outside the range (i.e., meeting a preset condition) can be used as a filtering threshold. The frequency of notes outside the range is then removed, and for notes within the range, frequencies below this filtering threshold are removed. After data cleaning, the highest and lowest notes from the remaining data can be used as the two endpoints of the target song's vocal range, thereby determining the target song's vocal range.
[0113] A3. Calculate the vocal range matching degree (i.e., singing matching degree) between the SVC singer and the target song based on the singer's vocal range and the song's vocal range. Determine whether this vocal range matching degree meets the preset matching degree conditions, such as whether this vocal range matching degree is greater than the preset vocal range matching degree threshold. If it is determined that it meets the preset matching degree conditions, use this SVC singer to perform vocal conversion on the target song based on SVC technology to obtain the content sung by this SVC singer (i.e., the target singing data).
[0114] In calculating the vocal range matching degree, for example, the complement of the singer's vocal range within the song's vocal range can be taken, and the degree of matching between the two can be calculated based on this complement. For instance, the portion of the song's vocal range that is not included in the singer's vocal range can be calculated; the proportion of this excluded portion in the song's vocal range can be calculated, and the difference between 1 and this proportion can be used as the vocal range matching degree between the SVC singer and the target song.
[0115] After calculating the vocal range matching degree between each SVC singer and the target song, SVC singers that meet the preset matching degree conditions (such as vocal range matching degree being greater than the preset matching degree threshold) can be selected to perform vocal conversion on the target song to ensure that there are no problems such as muteness, voice cracking, vocal strain, or distortion during vocal conversion.
[0116] Furthermore, the singer matching method provided in this embodiment can also be used for multi-dimensional matching of SVC singers and target songs. In this case, in addition to vocal range, matching dimensions such as musical style and timbre can also be included.
[0117] Figure 5 This is a schematic diagram illustrating the process of determining the matching degree of a musical style, as provided in an embodiment of this disclosure. Figure 5 As shown, taking an SVC singer as an example, the process of determining the genre matching degree can be described as follows:
[0118] B1. Determine the singer's musical style information (i.e., the second singing characteristic) of the SVC singer.
[0119] Specifically, a training set of SVC singers can be obtained; style detection technology can be used to detect the style of each training song in the training set to obtain the style information of each training song, such as the type matching degree between each training song and each style type; the style information of each training song can be cleaned, such as calculating the arithmetic mean of the type matching degree between each training song and this style type for each style type, which is used as the type matching degree between the SVC singer and this style type. Thus, the singer's style information (i.e., the second singing feature) of the SVC singer can be obtained.
[0120] B2. Determine the musical style information (i.e., the third singing characteristic) of the target song.
[0121] Specifically, the music style detection technology is used to detect the music style of the target song, and obtain the music style information of the target song, such as the type matching degree between the target song and each music style type.
[0122] B3. Calculate the style matching degree between the SVC singer and the target song based on the singer's style information and the song's style information.
[0123] Specifically, for each style (i.e. genre), the geometric mean between the probability that the singer is suitable for that style (i.e., genre matching degree) and the probability that the song is suitable for that style can be calculated. That is, the two are multiplied together and then the square root is taken (assuming that only two possible values are multiplied together) to obtain the genre matching degree between the SVC singer and the target song in that style.
[0124] For example, assuming the singer's genre information is: Pop 0.75, Folk 0.4, Rock 0.18, Metal 0.02, and the song's genre information is: Pop 0.8, Rock 0.3, Metal 0.1, then the pop style matching value between the SVC singer and the target song is (0.75 × 0.8). 1 / 2 ≈0.775; Folk style match value is (0.4×0) 1 / 2 =0; the rock style match value is (0.18 × 0.3). 1 / 2 ≈0.232; Metal style matching value is (0.02×0.1) 1 / 2 ≈0.141.
[0125] Figure 6 This is a schematic diagram illustrating a process for determining timbre matching degree according to an embodiment of the present disclosure, as shown below. Figure 6 As shown, taking an SVC singer as an example, the process of determining the timbre matching degree can be described as follows: Voiceprint technology is used to detect the voiceprints of the training songs in the singer's training set to determine the singer's voiceprint information; voiceprint technology is used to detect the voiceprints of the target song to determine the song's voiceprint information; the singer's voiceprint information and the song's voiceprint information are compared to determine the voiceprint similarity between the SVC singer and the target song, which serves as the timbre matching degree between the SVC singer and the target song for user reference. This timbre matching degree can be used to filter SVC singers whose timbre most closely resembles that of the target song.
[0126] When there are multiple target singing attributes, the overall matching degree between the SVC singer and the target song can be determined by weighting. For example, assuming that the importance of the vocal range attribute is higher than that of the style attribute, and the importance of the style attribute is higher than that of the timbre attribute, the weight of the vocal range matching degree can be set higher than that of the style matching degree, and the weight of the timbre matching degree can be set lower than that of the style matching degree. Finally, a weighted average is taken to obtain the overall matching degree between the SVC singer and the target song.
[0127] It should be noted that, in this embodiment of the disclosure, the singing features of the target song can also be used as input to perform singer matching through machine learning. For example, the model can be trained using input model training set information (such as pitch), target song information (such as pitch), and labeled data of the audio performance after SVC, to obtain a matching model. Therefore, when a target song exists, the original singing data of the target song can be input into this matching model to calculate the singing matching degree between the target song and each SVC singer.
[0128] The singing matching method provided in this embodiment, taking vocal range attributes as an example, ensures that the training songs for SVC singers are generally in their optimal state during training. For example, songs with unsuitable vocal ranges are discarded or have their pitch adjusted to ensure the quality of the training set. Therefore, analyzing the vocal range characteristics of the training set reveals the singer's vocal range information. Pitch analysis techniques can then be used to quantify pitch information, and the singer's vocal range can be calculated through data analysis. The target song's vocal track is separated using track splitting techniques, and pitch analysis is then used to obtain the song's vocal range. The SVC singer's vocal range is matched with the target song's vocal range, and the vocal range matching degree is calculated. Based on this matching degree, it can be determined whether the target song is suitable for the SVC singer, thus eliminating SVC singers with poor performance. For example, the calculated vocal range matching degrees for each SVC singer can be displayed to help users select the best-performing SVC singer, thereby improving the singing conversion effect.
[0129] Figure 7 This is a structural block diagram of a singing matching device provided in an embodiment of this disclosure. The device can be implemented by software and / or hardware, and can be configured in an electronic device, typically a computer, mobile phone, or tablet computer. It can determine the singing matching degree between the singer and the content to be sung by executing a singing matching method, such as determining the singing matching degree between the singer and the content to be sung during vocal conversion. Figure 7 As shown, the singing matching device provided in this embodiment may include: a feature extraction module 701, a feature determination module 702, and a matching degree determination module 703, wherein,
[0130] The feature extraction module 701 is used to obtain the training set of the singer and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer;
[0131] Feature determination module 702 is used to determine the second singing feature of the singer under the target singing attribute based on the first singing feature of each singing content sample;
[0132] The matching degree determination module 703 is used to determine the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute.
[0133] The singing matching device provided in this embodiment obtains a training set of singers through a feature extraction module and extracts the first singing features of each singing content sample in the training set under the target singing attribute. This training set is the set of singing content samples used when training the singer. A feature determination module determines the second singing features of the singer under the target singing attribute based on the first singing features of each singing content sample. Finally, a matching degree determination module determines the singing matching degree between the singer and the content to be sung based on the second singing features of the singer under the target singing attribute and the third singing features of the content to be sung under the target singing attribute. This embodiment utilizes the above technical solution to determine the singer's singing features based on the singing features of each singing content sample in the singer's training set. This eliminates the need for manual listening and annotation, enriching the methods for determining the singer's singing features, reducing the time spent on determining the singer's singing features, and improving the accuracy of the determined singing features. This, in turn, enriches the methods for determining the singing matching degree between the singer and the content to be sung, reduces the time spent on determining the singing matching degree between the singer and the content to be sung, and improves the accuracy of the determined singing matching degree.
[0134] Optionally, the target singing attribute includes a vocal range attribute, and the feature determination module 702 includes: a frequency statistics unit, used to count the first occurrence frequency of each attribute value of the target singing attribute in each singing content sample based on the first singing feature of each singing content sample; a curve generation unit, used to generate an occurrence frequency curve of the target singing attribute based on the first occurrence frequency value; and a range determination unit, used to determine the range of the first attribute value of the singer under the target singing attribute based on the derivative value of each attribute value in the occurrence frequency curve, as the second singing feature of the singer under the target singing attribute.
[0135] Furthermore, the singing matching device may further include: a first number determination module, used to determine, before generating the occurrence count curve of the target singing attribute based on the first occurrence count value, the valid occurrence count value and invalid occurrence count value among each of the first occurrence count values, based on a preset attribute value range, wherein the attribute value corresponding to the valid occurrence count value is within the preset attribute value range, and the attribute value corresponding to the invalid occurrence count value is outside the preset attribute value range; a second number determination module, used to obtain the target occurrence count value among each of the invalid occurrence count values that meets a preset condition, and determine the valid occurrence count value less than the target occurrence count value as an invalid occurrence count value; the curve generation unit may specifically be used to: generate the occurrence count curve of the target singing attribute based on the remaining valid occurrence count values.
[0136] Optionally, the range determination unit is specifically used to: obtain the minimum and maximum attribute values in the occurrence frequency curve where the absolute value of the derivative is greater than or equal to a preset threshold, based on the derivative values at each attribute value in the occurrence frequency curve; and determine the first attribute value range of the singer under the target singing attribute according to the minimum and maximum attribute values.
[0137] Optionally, the third singing feature is the range of second attribute values of the content to be sung under the target singing attribute, and the matching degree determination module 703 can be specifically used to: obtain the sub-attribute value range that overlaps between the first attribute value range and the second attribute value range; calculate the proportion of attribute values located within the sub-attribute value range in the second attribute value range, as the singing matching degree between the singer and the content to be sung under the target singing attribute.
[0138] Furthermore, the singing matching device may further include: a frequency statistics module, used to extract dry vocal data of the content to be sung from the original singing data of the content to be sung before obtaining the sub-attribute value range that overlaps between the first attribute value range and the second attribute value range, and to count the second occurrence frequency value of each attribute value of the target singing attribute in the content to be sung based on the dry vocal data; and a range determination module, used to determine the second attribute value range of the content to be sung under the target singing attribute based on the second occurrence frequency value.
[0139] Optionally, the target singing attribute includes a musical style attribute, which includes multiple attribute types, and the second singing feature is the type matching degree between the singer and each of the attribute types.
[0140] Optionally, the matching degree determination module 703 may be specifically used to: for each attribute type of the target singing attribute, calculate the geometric mean between the first type matching degree of the singer and the second type matching degree of the content to be sung, as the singing matching degree between the singer and the content to be sung under the attribute type, wherein the first type matching degree is the type matching degree between the singer and the attribute type, and the second type matching degree is the type matching degree between the content to be sung and the attribute type.
[0141] Optionally, the target singing attribute includes a timbre attribute, and the matching degree determination module 703 may be specifically used to: calculate the similarity between the second singing feature and the third singing feature of the content to be sung under the target attribute, as the singing matching degree between the singer and the content to be sung under the target singing attribute.
[0142] Optionally, the singer is a candidate singer from a set of singers, and the singing matching device may further include: an object selection module, used to select at least one candidate singer from the set of singers as the target singer of the content to be sung based on the singing matching degree between the singer and the content to be sung, according to the singing matching degree, after determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute; and a voice conversion module, used to perform voice conversion on the original singing data of the content to be sung based on the target singer to obtain the target singing data of the content to be sung.
[0143] The singing matching device provided in this disclosure can execute the singing matching method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the singing matching method. Technical details not described in detail in this embodiment can be found in the singing matching method provided in any embodiment of this disclosure.
[0144] The following is for reference. Figure 8 This illustration shows a structural diagram of an electronic device (e.g., a terminal device or a server) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0145] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0146] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0147] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0148] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0149] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0150] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0151] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a training set of the singer and extract a first singing feature of each singing content sample in the training set under a target singing attribute, wherein the training set is the set of singing content samples used when training the singer; determine a second singing feature of the singer under the target singing attribute based on the first singing feature of each singing content sample; and determine the singing matching degree between the singer and the content to be sung based on the second singing feature and a third singing feature of the content to be sung under the target singing attribute.
[0152] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0154] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of modules do not, in some cases, constitute a limitation on the unit itself.
[0155] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0156] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0157] According to one or more embodiments of this disclosure, Example 1 provides a singing matching method, including:
[0158] Obtain the training set of the singer, and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer;
[0159] Based on the first singing features of each singing content sample, the second singing features of the singer under the target singing attribute are determined;
[0160] Based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute, the singing matching degree between the singer and the content to be sung is determined.
[0161] According to one or more embodiments of this disclosure, Example 2 describes the method described in Example 1, wherein the target singing attribute includes a vocal range attribute, and the step of determining the singer's second singing feature under the target singing attribute based on the first singing feature of each singing content sample includes:
[0162] Based on the first singing feature of each singing content sample, the first occurrence count of each attribute value of the target singing attribute in each singing content sample is calculated.
[0163] Generate the occurrence frequency curve of the target singing attribute based on the first occurrence frequency value;
[0164] Based on the derivative values of each attribute value in the occurrence curve, the range of the first attribute value of the singer under the target singing attribute is determined, which serves as the second singing feature of the singer under the target singing attribute.
[0165] According to one or more embodiments of this disclosure, Example 3, based on the method of Example 2, further includes, before generating the occurrence frequency curve of the target singing attribute based on the first occurrence frequency value:
[0166] The valid occurrence count and invalid occurrence count are determined for each of the first occurrence count values based on a preset attribute value range, wherein the attribute value corresponding to the valid occurrence count value is within the preset attribute value range, and the attribute value corresponding to the invalid occurrence count value is outside the preset attribute value range;
[0167] Obtain the target occurrence count value that meets the preset condition from each of the invalid occurrence count values, and determine the valid occurrence count value that is less than the target occurrence count value as the invalid occurrence count value;
[0168] The step of generating the occurrence frequency curve of the target singing attribute based on the first occurrence frequency value includes:
[0169] Generate the occurrence curve of the target singing attribute based on the remaining valid occurrence counts.
[0170] According to one or more embodiments of this disclosure, Example 4, based on the method described in Example 2, the step of determining the range of first attribute values for the singer under the target singing attribute based on the derivative values at each attribute value in the occurrence curve includes:
[0171] Based on the derivative values of each attribute value in the occurrence frequency curve, the minimum and maximum attribute values in the occurrence frequency curve whose absolute derivative values are greater than or equal to a preset threshold are obtained.
[0172] The range of first attribute values for the singer under the target singing attribute is determined based on the minimum attribute value and the maximum attribute value.
[0173] According to one or more embodiments of this disclosure, Example 5 describes the method described in Example 2, wherein the third singing feature is the range of second attribute values of the content to be sung under the target singing attribute, and the step of determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes:
[0174] Obtain the sub-attribute value range that overlaps between the first attribute value range and the second attribute value range;
[0175] Calculate the percentage of attribute values within the range of the sub-attribute values in the range of the second attribute values, and use this percentage as the singing match degree between the singer and the content to be sung under the target singing attribute.
[0176] According to one or more embodiments of this disclosure, Example 6, based on the method of Example 5, further includes, before obtaining the sub-attribute value ranges that overlap between the first attribute value range and the second attribute value range:
[0177] Extract the dry vocal data of the content to be sung from the original vocal data of the content to be sung, and calculate the second occurrence count of each attribute value of the target vocal attribute in the content to be sung based on the dry vocal data;
[0178] The range of the second attribute value of the content to be sung under the target singing attribute is determined based on the second occurrence count value.
[0179] According to one or more embodiments of this disclosure, Example 7 describes the method described in Example 1, wherein the target singing attribute includes a musical style attribute, the musical style attribute includes multiple attribute types, and the second singing feature is the type matching degree between the singer and each of the attribute types.
[0180] According to one or more embodiments of this disclosure, Example 8 describes the method described in Example 7, wherein determining the singing match degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes:
[0181] For each attribute type of the target singing attribute, calculate the geometric mean between the singer's first type matching degree and the second type matching degree of the content to be sung, as the singing matching degree between the singer and the content to be sung under the attribute type, wherein the first type matching degree is the type matching degree between the singer and the attribute type, and the second type matching degree is the type matching degree between the content to be sung and the attribute type.
[0182] According to one or more embodiments of this disclosure, Example 9 describes the method according to Example 1, wherein the target singing attribute includes a timbre attribute, and the step of determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes:
[0183] Calculate the similarity between the second singing feature and the third singing feature of the content to be sung under the target attribute, and use it as the singing matching degree between the singer and the content to be sung under the target singing attribute.
[0184] According to one or more embodiments of this disclosure, Example 10, based on any one of Examples 1-9, wherein the singer is a candidate singer from a set of singers, after determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute, further includes:
[0185] Based on the singing matching degree, at least one candidate singer is selected from the set of singers as the target singer for the content to be sung;
[0186] Based on the original singing data of the target singer for the content to be sung, the voice is converted to obtain the target singing data of the content to be sung.
[0187] According to one or more embodiments of this disclosure, Example 11 provides a singing matching device, comprising:
[0188] The feature extraction module is used to obtain the training set of the singer and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer;
[0189] The feature determination module is used to determine the second singing feature of the singer under the target singing attribute based on the first singing feature of each singing content sample;
[0190] The matching degree determination module is used to determine the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute.
[0191] According to one or more embodiments of this disclosure, Example 12 provides an electronic device, including:
[0192] One or more processors;
[0193] Memory, used to store one or more programs.
[0194] When the one or more programs are executed by the one or more processors, the one or more processors implement the singing matching method as described in any of Examples 1-10.
[0195] According to one or more embodiments of the present disclosure, Example 13 provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the singing matching method as described in any of Examples 1-10.
[0196] According to one or more embodiments of this disclosure, Example 14 provides a computer program product that, when executed by a computer, causes the computer to implement the singing matching method as described in any of Examples 1-10.
[0197] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0198] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0199] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A singing matching method, characterized in that, include: Obtain the training set of the singer, and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer; Based on the first singing features of each singing content sample, the second singing features of the singer under the target singing attribute are determined; Based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute, the singing matching degree between the singer and the content to be sung is determined.
2. The method according to claim 1, characterized in that, The target singing attribute includes a vocal range attribute. Determining the singer's second singing feature under the target singing attribute based on the first singing feature of each singing content sample includes: Based on the first singing feature of each singing content sample, the first occurrence count of each attribute value of the target singing attribute in each singing content sample is calculated. Generate the occurrence frequency curve of the target singing attribute based on the first occurrence frequency value; Based on the derivative values of each attribute value in the occurrence curve, the range of the first attribute value of the singer under the target singing attribute is determined, which serves as the second singing feature of the singer under the target singing attribute.
3. The method according to claim 2, characterized in that, Before generating the occurrence frequency curve of the target singing attribute based on the first occurrence frequency value, the method further includes: The valid occurrence count and invalid occurrence count are determined for each of the first occurrence count values based on a preset attribute value range, wherein the attribute value corresponding to the valid occurrence count value is within the preset attribute value range, and the attribute value corresponding to the invalid occurrence count value is outside the preset attribute value range; Obtain the target occurrence count value that meets the preset condition from each of the invalid occurrence count values, and determine the valid occurrence count value that is less than the target occurrence count value as the invalid occurrence count value; The step of generating the occurrence frequency curve of the target singing attribute based on the first occurrence frequency value includes: Generate the occurrence curve of the target singing attribute based on the remaining valid occurrence counts.
4. The method according to claim 2, characterized in that, Determining the range of first attribute values for the singer under the target singing attribute based on the derivative values at each attribute value in the occurrence curve includes: Based on the derivative values of each attribute value in the occurrence frequency curve, the minimum and maximum attribute values in the occurrence frequency curve whose absolute derivative values are greater than or equal to a preset threshold are obtained. The range of first attribute values for the singer under the target singing attribute is determined based on the minimum attribute value and the maximum attribute value.
5. The method according to claim 2, characterized in that, The third singing feature is the range of second attribute values of the content to be sung under the target singing attribute. Determining the singing match degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes: Obtain the sub-attribute value range that overlaps between the first attribute value range and the second attribute value range; Calculate the percentage of attribute values within the range of the sub-attribute values in the range of the second attribute values, and use this percentage as the singing match degree between the singer and the content to be sung under the target singing attribute.
6. The method according to claim 5, characterized in that, Before obtaining the sub-attribute value range that overlaps between the first attribute value range and the second attribute value range, the method further includes: Extract the dry vocal data of the content to be sung from the original vocal data of the content to be sung, and calculate the second occurrence count of each attribute value of the target vocal attribute in the content to be sung based on the dry vocal data; The range of the second attribute value of the content to be sung under the target singing attribute is determined based on the second occurrence count value.
7. The method according to claim 1, characterized in that, The target singing attribute includes a musical style attribute, which includes multiple attribute types. The second singing feature is the type matching degree between the singer and each of the attribute types.
8. The method according to claim 7, characterized in that, The step of determining the singing match degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes: For each attribute type of the target singing attribute, calculate the geometric mean between the singer's first type matching degree and the second type matching degree of the content to be sung, as the singing matching degree between the singer and the content to be sung under the attribute type, wherein the first type matching degree is the type matching degree between the singer and the attribute type, and the second type matching degree is the type matching degree between the content to be sung and the attribute type.
9. The method according to claim 1, characterized in that, The target singing attribute includes a timbre attribute. Determining the singing match between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute includes: Calculate the similarity between the second singing feature and the third singing feature of the content to be sung under the target attribute, and use it as the singing matching degree between the singer and the content to be sung under the target singing attribute.
10. The method according to any one of claims 1-9, characterized in that, The singer is a candidate singer from the singer set. After determining the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute, the method further includes: Based on the singing matching degree, at least one candidate singer is selected from the set of singers as the target singer for the content to be sung; Based on the original singing data of the target singer for the content to be sung, the voice is converted to obtain the target singing data of the content to be sung.
11. A singing matching device, characterized in that, include: The feature extraction module is used to obtain the training set of the singer and extract the first singing feature of each singing content sample in the training set under the target singing attribute, wherein the training set is the set of singing content samples used when training the singer; The feature determination module is used to determine the second singing feature of the singer under the target singing attribute based on the first singing feature of each singing content sample; The matching degree determination module is used to determine the singing matching degree between the singer and the content to be sung based on the second singing feature and the third singing feature of the content to be sung under the target singing attribute.
12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the singing matching method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the singing matching method according to any one of claims 1-10.
14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the singing matching method according to any one of claims 1-10.