Training method for singing voice authenticity identification model, singing voice authenticity identification method, and related product
By training a singing feature extraction model and a target classifier model, and using singing features and auxiliary features to distinguish between AI-generated and real singing voices, the problem of identification in existing technologies has been solved, and efficient and accurate singing voice authentication has been achieved.
Patent Information
- Application Number
- PCT/CN2025/101611
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-28
- Filing Date
- 2025-06-18
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies struggle to effectively distinguish between AI-generated singing voices and real human singing voices, resulting in a heavy workload and a high risk of misjudgment, which impacts copyright and audience acceptance.
A singing feature extraction model and a target classifier model are used. By training singing segments with positive and negative samples, the distance between similar classes is reduced and the distance between dissimilar classes is increased. The target classifier model is trained by combining singing representation features and auxiliary features to identify the authenticity of singing.
It improves the efficiency and accuracy of voice authentication, enabling it to quickly distinguish between machine-generated and human-generated singing, thus filling a gap in voice authentication technology.
Smart Images

Figure CN2025101611_02012026_PF_FP_ABST
Abstract
Description
Song voice identification model training method, song voice identification method and related products
[0001] The present application claims priority to the Chinese patent application No. 202410861746.5, filed on June 28, 2024, and entitled "Song voice identification model training method, song voice identification method and related products", the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of audio technology, in particular to a song voice identification model training method, a song voice identification method and related products. BACKGROUND
[0003] With the development of artificial intelligence generated content (AIGC, Artificial Intelligence Generated Content) technology, AI generated song voices sound more and more natural, and it is already difficult for ordinary people to distinguish them. However, unauthorized imitation of a singer's synthesized songs often affects the audience size of the original songs and causes copyright or license disputes. In addition, from the mainstream consciousness, the general audience also cannot accept that the songs they hear are actually AI generated rather than sung by a real person.
[0004] Currently, most related audio identification technologies are for spoken voice and cannot be directly applied to song voice identification. Moreover, there are few technologies specifically for song voice identification, and professional personnel need to consider singing methods, skills, breath, pronunciation, and other acoustic factors in song voice to analyze whether the to-be-tested song is AI sung, which is a complex process. In the face of a large number of diverse songs released, relying on manual identification of a series of skill factors of each song is a heavy workload and is easy to mistake AI sung songs for real person songs. Therefore, there is an urgent need to provide an effective solution. SUMMARY
[0005] Embodiments of the present application provide a song voice identification model training method, a song voice identification method and related products to improve the efficiency of identifying the authenticity of song voices.
[0006] The first aspect of the embodiments of the present application provides a song voice identification model training method, the song voice identification model comprising a song voice feature extraction model and a target classifier model, and the training method comprising:
[0007] extracting each song voice segment from a machine sung song sample as a positive sample song voice segment and extracting each song voice segment from a real person sung song sample as a negative sample song voice segment; the song voice segment refers to an audio segment after removing accompaniment and song voice free segments in the song sample;
[0008] inputting the positive sample song segments and the negative sample song segments into the song feature extraction model respectively to extract song features corresponding to the positive sample song segments and the negative sample song segments respectively; the song feature extraction model is a feature extraction model trained towards a training target; the training target includes reducing distances between the positive sample song segments and increasing distances between the positive sample song segments and the negative sample song segments;
[0009] training an initial classifier model using the label information corresponding to the positive sample song segments and the negative sample song segments respectively and the song features until a convergence condition is met to stop training and obtain the target classifier model; the label information is used to mark whether a song segment is machine singing, and the target classifier model is used to output a prediction probability of a song segment in a to-be-tested song being a machine singing segment to determine whether the to-be-tested song is machine singing.
[0010] In an embodiment, after extracting each song segment from a machine singing song sample as a positive sample song segment, the training method further includes:
[0011] performing enhancement processing on the positive sample song segments to generate more positive sample song segments also belonging to machine singing.
[0012] In an embodiment, after extracting song features corresponding to the positive sample song segments and the negative sample song segments respectively, the training method further includes:
[0013] distinguishing at least song representation features from the song features corresponding to the positive sample song segments and the negative sample song segments respectively to train an initial classifier model; the song representation features are used at least to represent singing rhythm and / or emotional changes of a human voice signal in the song segments.
[0014] In an embodiment, the training of the initial classifier model using the label information corresponding to the positive sample song segments and the negative sample song segments respectively and the song features includes:
[0015] using, as song auxiliary features, features other than song representation features in the song features corresponding to the positive sample song segments and the negative sample song segments respectively; the song representation features are used at least to represent singing rhythm and / or emotional changes of a human voice signal in the song segments.
[0016] use the positive sample song segment, the negative sample song segment, and the label information corresponding to each of the positive sample song segment and the negative sample song segment, and the song auxiliary feature to jointly train a song representation model in the initial classifier model and the song feature extraction model to obtain the target classifier model and the trained song representation model;
[0017] The trained song representation model meets the training target, and a learning rate of the song representation model during training is less than a learning rate of the initial classifier model during training. The trained song representation model is configured to output a song representation feature corresponding to a song segment through a time domain signal of the song segment.
[0018] In an embodiment, if the song auxiliary feature includes multiple types of features, the training method further includes:
[0019] The multiple types of features are weighted and fused to obtain an auxiliary feature fusion result, and the auxiliary feature fusion result is used at least for training the initial classifier model.
[0020] In an embodiment, the obtaining of each song segment in the song sample set includes:
[0021] collecting positive sample songs sung by a machine and negative sample songs sung by a real person;
[0022] performing vocal separation on a current sample song with accompaniment in the positive sample songs and the negative sample songs to remove the accompaniment in the current sample song;
[0023] removing a non-song segment in the current sample song and splicing a song segment in the current sample song;
[0024] performing slicing on the spliced segment after removing the accompaniment and the non-song segment to obtain a song segment with a preset length.
[0025] In an embodiment, if a fusion probability obtained by fusing the prediction probabilities of each song segment in a sample song is greater than or equal to a target judgment threshold, the sample song is determined to be a song sung by a machine; and a determination process of the target judgment threshold includes:
[0026] adjusting an initial judgment threshold initially allocated to the target classifier model based on a proportion result between a prediction number and an actual number of each song sample until the initial judgment threshold is adjusted to a value that affects the proportion result to meet a preset result, and obtaining a target judgment threshold.
[0027] The predicted number refers to the number of songs in each song sample predicted by the target classifier model as machine singing based on the initial judgment threshold, and the real number refers to the number of songs actually sung by machines in each song sample.
[0028] The method of the first aspect of the application can be implemented by the content of the second aspect of the application when implemented.
[0029] The second aspect of the embodiment of the application provides a singing authentication method, which comprises:
[0030] Obtaining each singing segment in the to-be-tested song;
[0031] Inputting each singing segment into a singing authentication model to obtain a predicted probability that each singing segment is predicted as a machine singing segment; the singing authentication model is trained according to the training method of the first aspect or any specific implementation manner of the first aspect;
[0032] Fusing and calculating the predicted probabilities of each singing segment to obtain a fusion probability corresponding to the to-be-tested song;
[0033] If the fusion probability is greater than or equal to a target judgment threshold, it is determined that the to-be-tested song is a machine singing song.
[0034] In an implementation manner, if there are multiple singing authentication models, and there is a type difference between the song samples of each singing authentication model, the type difference does not include the distinction between machine singing and human singing, and the singing authentication method further comprises:
[0035] Obtaining the fusion probabilities of the to-be-tested song calculated by each singing authentication model respectively;
[0036] Comprehensively determining whether the to-be-tested song is a machine singing song based on the fusion probabilities of each singing authentication model.
[0037] In an implementation manner, after inputting each singing segment into a singing authentication model to obtain a predicted probability that each singing segment is predicted as a machine singing segment, the singing authentication method further comprises:
[0038] If the gap between multiple predicted probabilities exceeds a preset gap, it is determined that the to-be-tested song is a song sung by machines and humans together.
[0039] The third aspect of the embodiment of the application provides an electronic device, which comprises:
[0040] A central processing unit, a memory and an input-output interface;
[0041] The memory is a transitory storage memory or a persistent storage memory;
[0042] The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the method described in the first aspect of the embodiments or any specific implementation manner of the first aspect.
[0043] The fourth aspect of the embodiments of the present application provides a computer readable storage medium, including instructions, when the instructions run on a computer, make the computer execute the method described in the first aspect of the embodiments of the present application or any specific implementation manner of the first aspect.
[0044] The fifth aspect of the embodiments of the present application provides a computer program product containing instructions or a computer program, when the computer program product runs on a computer, makes the computer execute the method described in the first aspect of the embodiments of the present application or any specific implementation manner of the first aspect.
[0045] From the above technical solutions, the embodiments of the present application have at least the following advantages:
[0046] Selecting positive sample song fragments and negative sample song fragments with different labels helps to provide sufficient and reliable data support for the training of the song authentication model, and promotes the song authentication model to more deeply identify the song fragments in the to-be-tested song. Among them, the song feature extraction model is used to extract the song features of each song fragment, which can effectively replace the artificial learning of the acoustic differences between machine singing and human singing, so as to guarantee the efficiency and accuracy of the target classifier model in identifying the authenticity of the song, and at the same time, fill the technical gap of detecting machine-generated song. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art according to these drawings.
[0048] It should be noted that although each step in the flowchart (if any) involved in each embodiment is drawn in sequence according to the arrow, unless otherwise stated in this paper, the execution of these steps has no strict order limit, and these steps can be executed in other order. Moreover, at least one part of the steps in the flowchart involved in each embodiment can include multiple steps or multiple stages, which do not necessarily be executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed with at least one part of other steps or steps or stages in other steps or stages.
[0049] FIG. 1 is a schematic diagram of a system architecture according to an embodiment of the present application;
[0050] FIG. 2 and FIG. 3 are schematic diagrams of a training method of a singing voice authentication model according to an embodiment of the present application;
[0051] FIG. 4 and FIG. 5 are schematic diagrams of a singing voice authentication method according to an embodiment of the present application;
[0052] FIG. 6 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0054] The terms "first", "second", "third", "fourth" and the like (if any) in the description, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0055] In the following description, references are made to "one embodiment" or "one example" or similar expressions, which describe a subset of all possible embodiments, but it can be understood that "one embodiment" or "one example" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict. In the following description, the term "a plurality of" refers to at least two. If the threshold value (if any) mentioned in the present application is reached, in some specific examples, it can include the case where the former is greater than the threshold value. If "any" or "at least one" or similar expressions are mentioned, it can specifically refer to any one of the listed examples or any combination between these examples.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0057] Please refer to the system architecture diagram of the embodiment of the present application as shown in FIG. 1. The audio recognition method provided by the embodiment of the present application can be applied in the application environment as shown in FIG. 1, wherein the terminal 102 communicates with the server 101 through a network, and the data storage system 100 can store data required to be processed by the server 101. The data storage system 100 can be integrated on the server 101, or can be placed on a cloud or other network server. The terminal 102 can obtain a to-be-tested song and transmit the to-be-tested song to the server 101; the server 101 can output the prediction probability of each song fragment in the to-be-tested song being a machine singing fragment through a trained song voice identification model, and comprehensively judge whether the to-be-tested song is a machine singing song or a real person singing song; then, the server 101 can return the judgment result to the terminal 102, so that the user knows the authenticity (i.e., whether it is an AI singing song) of the to-be-tested song.
[0058] The terminal 102 described above can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 101 can be implemented by an independent server or a server cluster composed of multiple servers. It should be noted that the method provided by the embodiment of the present application can be implemented by the terminal device and the server as described above, or can be implemented on the server side or on the terminal device side, and the specific implementation can be determined according to the actual application scene, which is not limited here.
[0059] The method of the present application will be described in detail below.
[0060] Please refer to FIG. 2, the first aspect of the present application provides one specific embodiment of a training method of a song voice identification model, the song voice identification model includes a song voice feature extraction model and a target classifier model, and the embodiment includes the following operation steps:
[0061] Step 21, obtaining positive sample song fragments and negative sample song fragments.
[0062] Specifically, each song fragment can be extracted from a machine (which can be referred to as AI) singing song sample as a positive sample song fragment, and each song fragment can be extracted from a real person singing song sample as a negative sample song fragment. The song fragment refers to an audio fragment after removing the accompaniment and the song fragment in the song sample.
[0063] The song segment of the embodiment of the present application can specifically refer to a whole piece of pure song (i.e. audio without accompaniment and songless segment) that is not sliced, or a plurality of sub-segments obtained by slicing the whole piece of pure song; considering that there is often an interval between the lyrics before and after, using the whole piece of song will have space blank or time span too long, which is easy to cause waste of training resources or prolong the training time, so the song segment in the model training and use process generally uses the sub-segment after slicing. The number of positive sample song segments and negative sample song segments can be relatively flat to prevent any segment from being too much to affect the judgment accuracy of the song authentication model on the song; of course, the number of positive sample song segments can also be increased to improve the robustness of the model and better adapt to actual application.
[0064] In actual process, the data used for training and testing in the embodiment of the present application has been authorized, and the source of song samples sung by machine (or simply referred to as positive samples) can include:
[0065] 1) Video website. Taking a certain APP video as an example, AI singing videos can be obtained from the video library according to the video title, introduction and other information, screened according to a certain strategy, and then the audio thereof is extracted as positive samples. The selection data amount of positive samples can be increased according to the user-created collection.
[0066] 2) Song synthesis software. AI song synthesis software is used to synthesize AI song, for example, music melody information and song lyrics are input into the song synthesis software, and the audio of AI singing is obtained as positive sample data;
[0067] 3) AI singing songs found in the song library.
[0068] The source of song samples sung by real person (or simply referred to as negative sample) can include song library or musician open platform, and most of the songs obtained from these sources are sung by real person, and a small amount of AI singing songs can be excluded through model prediction or manual screening.
[0069] Step 22, input the positive sample song segment and the negative sample song segment into the song feature extraction model respectively to extract the song features corresponding to the positive sample song segment and the negative sample song segment respectively.
[0070] Among them, the song feature extraction model can be a feature extraction model trained towards the training target; the training target includes reducing the distance between the positive sample song segments and increasing the distance between the positive sample song segments and the negative sample song segments. In other words, the song feature extraction model can be a feature extraction model trained by adopting the idea of contrast learning.
[0071] The contrast learning idea is adopted here, which helps the singing feature extraction model to automatically learn the distinguishing feature representation of the sample under limited sample resources, so that the difference between the samples can be captured through the feature representation. For example, the singing feature extracted by the singing feature extraction model is input into the classifier model to make a binary classification judgment on the song as AI singing or human singing. The singing feature can represent the time information and / or pronunciation information in the singing, which helps the classifier model to fully learn and identify the authenticity of the singing, i.e., whether it is AI singing or not.
[0072] Step 23, using the label information and the singing feature corresponding to each of the positive sample singing segment and the negative sample singing segment, training the initial classifier model until the convergence condition is met to stop training, and obtaining a target classifier model.
[0073] The label information is used to mark whether the singing segment is machine singing. For example, the label of the positive sample song and the singing segment therein can be set to 1, and the label of the negative sample song and the singing segment therein can be set to 0. The target classifier model is used to output the prediction probability of the singing segment in the to-be-tested song as a machine singing segment, so as to judge whether the to-be-tested song is machine singing.
[0074] Generally, since the singing type in the data can be judged from the data source as human singing or AI generated, manual type labeling of the song and its singing segment is not required. For a small amount of incorrect labels, an existing model (such as the singing feature extraction model described above) or manual investigation can be used to determine whether the label needs to be corrected or whether the song with the incorrect label needs to be abandoned as a sample song.
[0075] In summary, selecting positive sample singing segments and negative sample singing segments with different labels helps to provide sufficient and reliable data support for the training of the singing authenticity identification model, and promotes the singing authenticity identification model to more deeply identify the singing segments in the to-be-tested song. The singing feature extraction model is used to extract the singing feature of each singing segment, which can effectively replace manual learning of the acoustic differences between machine singing and human singing, to ensure the efficiency and accuracy of the target classifier model in identifying the authenticity of the song. At the same time, it fills the technical gap in detecting machine-generated singing.
[0076] On the basis of the above example description, some specific possible implementation examples will be provided below. In actual application, the implementation contents of these examples can be combined or implemented individually according to the corresponding functional principles and application logic. If combined, the execution order of the combined examples can be determined according to the respective processing logic, and the specific execution order can be determined according to the actual scene.
[0077] Based on the example content of FIG. 2, the present application provides another embodiment of a training method of a singing voice authentication model, which includes the following operation steps:
[0078] Step 21, obtaining positive sample singing voice segments and negative sample singing voice segments.
[0079] In some specific examples, step 21 can specifically include the following preprocessing process: collecting machine-sung song samples and human-sung song samples, removing accompaniment and non-singing voice segments present in the song samples, and splicing the remaining segments in the song samples; slicing the spliced segments after removing the accompaniment and non-singing voice segments to obtain singing voice segments of a predetermined length.
[0080] The specific sources of positive sample songs and negative sample songs can be found in the related description in the above embodiments, which will not be repeated here. For example, for the collected positive sample songs and negative sample songs that both contain singing voices, the following preprocessing process can be performed to make them available for model training:
[0081] 1) Singing voice accompaniment separation: separate the singing voice from each sample song to obtain the audio corresponding to the pure singing voice;
[0082] 2) Removing non-singing voice parts and splicing: removing the non-singing voice parts in the audio by a voice activity detection (VAD) method or a method of detecting loudness, and splicing the parts containing singing voices together;
[0083] 3) Slicing: dividing the spliced audio into non-overlapping segments of a predetermined length, such as 5 seconds according to the required time length of the model; here, the front and back segments of the same song are set to be non-overlapping, which helps to reduce the loss of training resources caused by overlapping segments.
[0084] It should be noted that the execution order of the above-mentioned singing voice accompaniment separation, non-singing voice removal, splicing and slicing, etc. can not be limited, and can be set according to actual conditions, which is not limited here.
[0085] The singing voice segments obtained after the above preprocessing operations can be input into a singing voice feature extraction model to extract singing voice features in the singing voice segments. When the singing voice feature extraction model is specifically composed of a singing voice representation model and a singing voice transcription model, the corresponding extracted singing voice features can be singing voice representation features, singing voice auxiliary features, etc. The specific details can be found in the description of step 22 below.
[0086] To increase the number of positive sample song segments, improve the robustness of the model, and better adapt to practical applications, in some specific examples, after step 21, the training method of the embodiment of the application can further include: performing enhancement processing on the positive sample song segments to generate more positive sample song segments that also belong to machine singing.
[0087] For example, the existing positive sample songs or positive sample song segments can be subjected to at least one of the following data enhancement operations, such as audio cropping, time stretching, pitch conversion, and adding noise, to generate more positive sample song segments. These enhanced songs are similar to the original songs in content or style, but have slight differences, which can be used for the song feature extraction model to more thoroughly learn the feature representation of machine singing songs and improve the anti-fake ability of the classifier model.
[0088] Step 22, input the positive sample song segments and the negative sample song segments into the song feature extraction model to extract the respective song features corresponding to the positive sample song segments and the negative sample song segments.
[0089] In practical applications, to avoid investing a large amount of computing resources in some less practical song features, important song features (or main features) with greater practicality can be distinguished to train the initial classifier model. Therefore, in some specific examples, after step 22, the training method of the embodiment of the application can further include: distinguishing at least song representation features from the respective song features corresponding to the positive sample song segments and the negative sample song segments to train the initial classifier model; and the song representation features are used at least to represent the singing rhythm and / or emotional changes of the human voice signal in the song segments.
[0090] For example, the song feature extraction model can include a song representation model and a song transcription model that are both trained. In one implementation, the time domain signals of each song segment can be input into the song representation model to obtain the song representation features of each song segment; and the frequency domain signals of each song segment can be input into the song transcription model to obtain the song auxiliary features of each song segment.
[0091] As shown in FIG. 3, the singing authentication model (or AI singing recognition model) of the embodiment of the present application mainly consists of three parts: a singing representation model (i.e., an audio representation pre-training model shown in the figure), a singing transcription model, and a target classifier model. The entire singing authentication model receives the audio waveform of a singing segment as input to output the probability value of the singing segment being AI singing. Specifically, the singing representation model can be input with each singing segment of the same song to obtain the singing representation features of each singing segment. The singing representation features are more abstract features than the singing auxiliary features, and specifically can include at least one singing main feature such as emotional features, signal features of machine-synthesized sound, and rhythm features, or can be interpreted as including features such as singing techniques or vocalization methods, which can belong to the description of singing in the frequency domain and / or time domain. The singing representation features are input into the classifier as main features for training and recognition.
[0092] Step 23: using the label information and the singing features corresponding to the positive and negative sample singing segments respectively, the initial classifier model is trained until the convergence condition is met to stop training, and the target classifier model is obtained.
[0093] In some examples, the specific operation process of step 23 can include: using the features other than the singing representation features in the singing features corresponding to the positive and negative sample singing segments respectively as singing auxiliary features; and using the singing representation features to at least represent the singing rhythm and / or emotional changes of the human voice signal in the singing segment.
[0094] In addition to the main features output by the singing representation model, other features (i.e., singing auxiliary features) can also be used for auxiliary judgment. The singing auxiliary features specifically can include pronunciation features and / or duration features of the lyrics in the singing segment, etc. In the training process, before the audio data is input into the singing transcription model, the time-domain waveform of the singing segment can be converted into a frequency spectrum signal such as a waveform graph into a frequency spectrum graph using Fourier transform or the like, so that the singing transcription model can better learn the frequency domain characteristics in the singing, etc.
[0095] It should be noted that in the case where the singing representation model is used as a singing feature extraction model, the singing transcription model of the embodiment of the present application can or can not exist, i.e., the auxiliary singing transcription model can or can not be enabled, because the singing representation features extracted by the singing representation model as main features can also include or alternatively refer to the singing auxiliary features output by the singing transcription model.
[0096] Exemplarily, the audio representation pre-training model can be selected from MERT, MULE or CLMR, etc. These models are trained by a large amount of song data and can represent various music information in the audio. More general audio representation pre-training models such as ABT or M2D model can also be used to extract features from the singing segment. These models have good performance on both speech tasks and music tasks. Taking a recognition model using the MERT model for feature extraction as an example, the MERT model receives the time-domain waveform of the singing segment as input and outputs a feature with a shape of [25, T, 1024]. The size of the time dimension parameter T (which can represent how many features in 1s) is different for input audios of different lengths. Feature compression can be performed by averaging along this dimension to obtain a feature vector with a size of [25, 1024]. The values 25 and 1024 can be simply regarded as the row and column numbers of the vector.
[0097] In combination with the characteristics of singing voice, state features, pitch features, pronunciation features, duration features or intensity features, etc. can be designed and used to assist recognition. These features can be identified by a trained singing transcription (singing transposition) model. This is because the singing transcription model has the ability to identify features such as pitch, note duration, phoneme, and pronunciation. Generally, a singing transcription model can be composed of several convolution modules and bidirectional long short-term memory (Bi-directional Long Short-Term Memory, Bi-directional LSTM) modules. Therefore, the singing segment of the embodiments of the present application can be trained on a singing transcription task first, and then used for a singing authentication task. The singing transcription task refers to automatically identifying the lyrics information and the pitch and duration information of each note from the singing audio. The data used to train the singing transcription model can include singing audio and corresponding note start point information, pitch information, lyrics information, phoneme information, etc. This part of data can be obtained from open source datasets and company internal data authorization. In actual situations, before inputting the audio data into the singing transcription model, the audio waveform should be converted into a spectrogram using Fourier transform or the like, so that the singing transcription model can better learn the feature description of the singing voice in the frequency domain, thereby making up for or assisting the recognition ability of the audio representation pre-training model for some features.
[0098] The following will specifically illustrate the song auxiliary features. In some examples, the following five categories of song auxiliary features can be selectively adopted or combined, i.e., the song auxiliary features output by the song transcription model can be one or more categories of features, which can be set by actual conditions and is not limited herein. It should be noted that the song transcription model in the embodiments of the present application can refer to a specific model structure or a specific algorithm, because a part of the song auxiliary features described below can be obtained by the song transcription model, and a part of the song auxiliary features can be directly calculated by an algorithm. In view of the fact that the model itself is composed of at least one algorithm (the model and the algorithm have a strong correlation that can be mutually referred to), for the convenience of description, the model structure or algorithm used to output the song auxiliary features is collectively referred to as the song transcription model. Whether any of the following song auxiliary features is specifically processed by a model structure or an algorithm can be determined by referring to the following description or by actual conditions, and is not limited specifically.
[0099] 1) State feature: For each frame of audio, four states can be defined: singing start frame, singing end frame, singing frame and non-singing frame. These states can be represented by a three-dimensional vector processed by an activation (Sigmoid) function, where the first element and the second element represent the likelihood of the frame being a singing start frame and a singing end frame, respectively, and the third element represents the probability of the frame being a singing frame or a non-singing frame. The shape of the feature vector can be represented as [T, 3], where the size of T depends on the audio duration and the resolution of the feature in the time domain, which is the frame rate after the audio is framed.
[0100] 2) Pitch feature: The pitch feature consists of two parts: note pitch feature and frequency pitch feature. The note pitch feature is output by the song transcription model, and the shape is [T, 66]. The frequency pitch feature is obtained by the fundamental frequency extraction (pYIN, Probabilistic YIN) algorithm, and the shape is [T, 1]. In the song transcription model, the pitch recognition task can be regarded as a classification task, such as dividing the pitches from G1 to C7 into 66 pitch classes according to semitones (i.e., 100 cents). For each frame, the model will output the classification probability (processed by the Softmax function) of its pitch. The pYIN algorithm directly calculates the pitch frequency value of each frame, which can be independent of the transcription model. When using the frequency pitch feature, it needs to be mapped to the Mel scale before being input to the classifier model, because the Mel scale is more consistent with the human ear's perception of sound.
[0101] 3) Pronunciation feature: the song transcription model will output its phoneme classification probability for each frame, shaped as [Np], Np is the number of phoneme classes in the training data. To better identify the pronunciation feature, when training the song transcription model, the CTC (Connectionist Temporal Classification) loss function can be used for model tuning. In order to obtain the pronunciation feature while avoiding being limited to the phoneme classes in the training data, the values in the hidden layer before the output layer can be taken as the pronunciation feature, shaped as [Nh], Nh is the number of neurons in the hidden layer. Considering the time dimension, the shape of the pronunciation feature can be recorded as [T, Nh].
[0102] 4) Duration feature: the duration feature consists of two parts, pronunciation duration feature and phoneme duration feature. The pronunciation duration feature can be obtained by clustering the pronunciation feature, shaped as [T, 1], and the phoneme duration feature can be obtained by the model's recognition result of the phoneme, shaped as [T, 1]. When clustering the pronunciation feature, the k-means clustering algorithm can be used, the number of clustering centers k can be set by experience or actual situation, such as 50, and the class ID of each frame after clustering can be concatenated to obtain the pronunciation duration feature. For the phoneme duration feature, similar to the pronunciation duration, the class ID of the phoneme recognition result can be concatenated to obtain it.
[0103] 5) Strength feature: the root mean square (RMS) value of the audio signal can be used to represent the strength of the song, shaped as [T, 1], where each frame corresponds to a value, and the short-time energy value can also be used to represent the strength feature, which is not limited.
[0104] In some specific examples, if the song auxiliary feature includes multiple types of features, such as state feature and strength feature, etc., the training method further includes: weighting and fusing the multiple types of features to obtain an auxiliary feature fusion result; and the auxiliary feature fusion result is used at least for training the initial classifier model.
[0105] According to the importance of each feature to the song authentication effect and other considerations, the inventors of the present application have obtained the following rules: among all the features mentioned above, the song representation feature extracted by the song representation model is the most important feature, followed by the pronunciation feature and the duration feature in the auxiliary feature, and then the state feature, the pitch feature and the strength feature. Therefore, when training the classifier model, the weights of each feature participating in the training of the classifier model can be allocated according to the foregoing rules, for example, for each song auxiliary feature, the auxiliary feature fusion result = a*pronunciation feature + b*duration feature + c*state feature + d*pitch feature + e*strength feature, and the weight coefficients a, b, c, d, e can be sequentially reduced. Of course, different weight allocation rules can be set for each song auxiliary feature.
[0106] Based on the above description, as shown in FIG. 3, the classifier model can be divided into a main part and an auxiliary part. The main part is composed of several convolutional layers, fully connected layers, activation layers, and regularization layers, and finally connected with a Softmax layer to obtain the probability value of the audio belonging to the "real person singing" class and the "AI singing" class. When using the model, the "AI singing" probability value can be taken as the prediction result. The main part receives the singing feature output by the pre-trained audio representation model as input. Taking the recognition model using the MERT model for feature extraction as an example: the first layer of the classifier main part is a convolutional layer, which combines the features into 1024-dimensional features, and then connects with several fully connected layers, activation function layers and Dropout layers, finally maps to a 2-dimensional output, and performs Softmax calculation to obtain the probability value of two categories.
[0107] The auxiliary part of the classifier model receives five types of auxiliary features as input. First, the five types of auxiliary features are fused, and then put into a Bi-LSTM network to obtain a global representation, and then combined with the intermediate layer features of the main part to input into the subsequent fully connected layer for recognition. Before feature fusion, the five types of auxiliary features can be refined using a convolutional layer to ensure that the features are aligned in the time dimension. When fusing the features, each feature can be spliced according to the second dimension. In order to balance the feature dimension, the state feature, the frequency pitch feature, the duration feature and the strength feature can be copied several times (such as 10 times) before feature fusion. Normalization needs to be performed before splicing. Attention mechanism can also be added when fusing the features, that is, each feature can be adaptively weighted according to its importance.
[0108] Based on the example description of step 22, if the singing transcription model is selected, the specific operation process of "training the initial classifier model using the label information and singing features of each singing segment" in step 23 can include: training the initial classifier model using the label information, singing representation features and singing auxiliary features of each singing segment. In other words, during the model training process, the singing representation model can be frozen (regarded as a trained model that can be directly used), and only the classifier model is trained. Experiments show that this training scheme can achieve good results. Correspondingly, the training data for training the classifier model at this time can include the label information (such as label values 0 and 1), singing representation features and singing auxiliary features of each singing segment.
[0109] For example, only the training scheme of the classifier model: freeze the song representation model; for the collected audio data, sequentially perform song accompaniment separation, remove the non-song part, splicing and slicing operations (this part of the operation can be referred to as preprocessing operation). Then use the song representation model, the song transcription model and the corresponding algorithm to extract the required features. Among them, the song transcription model needs to be trained on the song transcription task before training the classifier. Use the extracted features (including song representation features) and data labels to train the classifier model to obtain the final model.
[0110] Of course, as another possible implementation, the song representation model can also be selected for training to continue to improve the overall song authentication effect of the song authentication model. For example, the parameters of the song representation model can be selected for training. It should be noted that the above frozen model parameters are to retain the existing feature extraction capability of the model, prevent the model performance from declining due to changes in training data, such as preventing overfitting, or to save computing resources. Freezing part or all of the model parameters can be set according to actual conditions, which is not limited here. Unlike the case where the song representation model is not trained, if the song representation model needs to be unfrozen and trained, the song representation features do not need to be extracted in advance, that is, the training data of the classifier model in this case can not contain song representation features.
[0111] The following describes the case where the song representation model is also trained.
[0112] If the song representation model also needs to be trained, the specific operation of step 22 can include: inputting the frequency domain signals of each song segment into the song transcription model to obtain the song auxiliary features of each song segment. Then, the specific operation for step 23 can include: using the positive sample song segments, the negative sample song segments, and their respective label information and song auxiliary features, jointly training the song representation model to be trained and the initial classifier model to obtain the trained song representation model and the target classifier model; wherein the trained song representation model meets the training target, the learning rate of the song representation model to be trained is less than the learning rate of the initial classifier model when being trained; the trained song representation model is used to output the song representation features corresponding to the song segments from the time domain signals of the song segments.
[0113] Exemplarily, the song auxiliary features of the song segments are extracted using the song transcription model and the corresponding algorithm; then, the song representation model and the classifier model are jointly trained as a whole using the song segments (which can be obtained through the above-mentioned preprocessing), the extracted song auxiliary features, and the original labels of the song segments, to obtain the finally trained models of the two. When training the song representation model, the parameters of the song representation model can be unfrozen partially or entirely. Since the purpose of training the unfrozen song representation model is to fine-tune the song representation model, a relatively small learning rate can be set for the training of the song representation model, while a relatively large learning rate can be set for the classifier model. It is to be noted that the learning rate determines the speed of the model to reach the optimal solution or convergence during the training process, and generally should not be too high or too low, and can be set according to the actual situation.
[0114] In some specific examples, the trained song authenticity verification model can be used to predict the sample song segments. For the data whose prediction result is opposite to the sample label, manual inspection can be prompted to determine whether the label needs to be modified, so as to clean the data of the sample song and its song segments.
[0115] In some specific examples, if the fusion probability obtained by fusing the prediction probabilities of the song segments in the sample song is greater than or equal to a target judgment threshold, the sample song is determined to be a machine-sung song; the determination process of the target judgment threshold includes: adjusting the initial judgment threshold initially allocated to the target classifier model based on the proportion result between the predicted number and the true number of the song samples, until the initial judgment threshold is adjusted to the proportion result that meets the preset result, and the adjustment is stopped, to obtain the target judgment threshold; wherein the predicted number refers to the number of songs in the song samples that are predicted by the target classifier model to be machine-sung based on the initial judgment threshold, and the true number refers to the number of songs in the song samples that are actually machine-sung.
[0116] Exemplarily, the above-mentioned proportion result can be used as an evaluation index to set the judgment threshold t or the prediction accuracy of the model. Generally, the fusion probability p(AI singing) obtained by fusing the prediction probabilities of the song segments in the song can be calculated, and if p(AI singing) is greater than or equal to the target judgment threshold t, the song can be determined to be AI-sung; if p(AI singing) is less than the target judgment threshold t, the song can be determined to be human-sung. The above-mentioned proportion result can be at least one of the following indicators, which can reflect the prediction effect of the model or the appropriateness of the current judgment threshold t to some extent:
[0117] a. Accuracy: Accuracy = number of correctly identified samples / total number of samples, the higher the value, the better the prediction effect of the model, or the more appropriate the current judgment threshold t;
[0118] b. Precision: Precision = Number of samples identified as AI singing and identified correctly / Number of samples identified as AI singing, the higher the value the better;
[0119] c. Recall: Recall = Number of samples identified as AI singing and identified correctly / Number of samples actually AI singing, the higher the value the better;
[0120] d. False Acceptance Rate: False Acceptance Rate = Number of samples identified as AI singing but identified incorrectly / Number of samples actually human singing, the lower the value the better;
[0121] e. False Rejection Rate: False Rejection Rate = Number of samples identified as human singing but identified incorrectly / Number of samples actually AI singing, the lower the value the better;
[0122] f. Equal Error Rate: The threshold can be adjusted so that the false acceptance rate is equal to the false rejection rate, at this time the value of the two is called the equal error rate, the lower the value the better.
[0123] In some specific examples, the song sample set can be divided into several sub-datasets, and the division basis can be data source, song style or original singer, etc. Each sub-dataset can be used to train a targeted song identification model. For example, a song identification model that is good at identifying sad songs and inspirational songs can be trained.
[0124] It is additionally explained that the following measures can be taken during model training.
[0125] Data expansion: Since AI song generation methods are constantly updated, AI song data corresponding to them can be supplemented to enable the model to better identify songs generated by new methods. After expanding AI singing data, human singing data can also be supplemented to balance positive and negative samples, so that the prediction effect of the model is not lost, that is, human singing data should also be as diverse as possible. At the same time, the proportion of data corresponding to various AI song generation methods (i.e. various data sources) can be controlled and adjusted to enable the model to have better recognition ability. In addition, when expanding data for model iteration, an incremental learning method can be used so that the model does not forget old knowledge when learning new knowledge.
[0126] Category division: Since there are various song synthesis methods, and human singing timbre and singing methods are also different, the above-mentioned binary classification (human or AI singing) problem can be converted into a multi-classification problem for solution. For example, a judgment category of AI and human chorus can be added, or a segment interval of human or AI singing can be added.
[0127] Loss function: The cross-entropy function can be used as the basic loss function of the model. At the same time, a ternary loss function can be added for assistance.
[0128] Model diversification: Since the models trained using different sub-datasets (e.g., with style differences) or different model structures (e.g., using MERT or MULE models) each have their own strengths in identifying different types of singing, multiple singing identification models can be trained to target different types of singing.
[0129] Determining the time scale: Because the length of the input segment selected at the input time will affect the results due to differences in song type and singing style, multiple models can be trained for different input lengths (e.g., 2s, 5s, 8s) when training the model. Correspondingly, when using the model, the outputs of multiple models can be averaged or weighted to obtain the final result.
[0130] In summary, the training method of the embodiments of the present application collects a large amount of real human singing and AI-generated singing to construct a dataset, uses and trains multiple network models to extract multiple audio features, can train a classifier model to have the ability to identify the authenticity of singing, and fills the gap in related datasets and detection techniques for AI-generated singing, and effectively reduces the cost of manual input.
[0131] Compared with the example illustrated in FIG. 2, the above-mentioned additional or refined examples or possible implementation manners do not necessarily have to be executed in specific implementation, such as two or more examples or possible implementation manners are added, these examples or possible implementation manners can be implemented in combination or separately, if implemented in combination, the execution order between the combined examples can be determined according to the respective processing logic, and the specific implementation can be determined according to the actual scene.
[0132] Referring to FIG. 4, the second aspect of the present application provides one specific embodiment of a singing identification method, which includes the following operation steps:
[0133] Step 41, obtaining each singing segment in the to-be-tested song.
[0134] The process of obtaining each singing segment in the to-be-tested song can refer to the preprocessing process of the sample song described in the first aspect above, which will not be described here. The length of each singing segment of the to-be-tested song can be less than or equal to the preset length of the slice required during model training.
[0135] Step 42, inputting each singing segment into a singing identification model to obtain the prediction probability of each singing segment being predicted as a machine singing segment.
[0136] The singing identification model can be trained according to the training method described in the first aspect or any specific embodiment of the first aspect, which will not be described here.
[0137] Step 43, the prediction probability of each singing segment is fused to obtain the fusion probability corresponding to the to-be-tested song.
[0138] The fusion probability is calculated from the prediction probability of each singing segment, such as average or weighted fusion to obtain the fusion probability corresponding to the whole song, so that the comparison between the fusion probability and the target judgment threshold can be made to determine whether the to-be-tested song is a machine-sung song.
[0139] Step 44, if the fusion probability is greater than or equal to the target judgment threshold, it is determined that the to-be-tested song is a machine-sung song.
[0140] In summary, when using the model for recognition, the to-be-tested song can be preprocessed to obtain its singing segments, then the singing identification model is used to output the prediction probability of each singing segment, and the fusion probability calculated from each prediction probability is used to determine whether the to-be-tested song is a machine-sung song. As shown in FIG. 3, the singing identification model can mainly consist of an audio feature pre-training model (i.e., a singing feature model), a singing transcription model (which can or can not be included), and a classifier model. Specifically, each singing segment in the to-be-tested song can be extracted by the audio feature pre-training model, the singing transcription model, and the corresponding algorithm, and then the extracted features are put into the classifier model for recognition to obtain the prediction probability of each singing segment being a machine-sung segment.
[0141] Based on the example content of FIG. 2, some specific embodiments of a singing identification method will be provided below, which examples do not necessarily have to be implemented, and these examples can be combined or implemented alone, and if combined, the execution order between the combined examples can be determined according to the respective processing logic, which can be determined according to the actual scene.
[0142] In some specific examples, if there are multiple singing identification models, and the song samples between each singing identification model have type differences, and the type differences do not include the distinction between machine singing or human singing, the singing identification method can further include the following operation (fusion of multiple model results): obtaining the fusion probability of the to-be-tested song calculated by each singing identification model; based on the fusion probability of each singing identification model, comprehensively determining whether the to-be-tested song is a machine-sung song.
[0143] Exemplarily, the type difference can refer to a genre difference, or a data source difference (e.g., from a video platform or an audio playing platform), or a model structure difference. The model structure difference can refer to that a song representation model adopts a MERT model (Music Understanding Model with Large-Scale Self-supervised Training), a song representation model adopts a MULE model (Musicset Unsupervised Large Embedding), or in another embodiment, can refer to that a classifier model combined with the song representation model contains different convolution layers or fully connected layers, such as different structures or numbers.
[0144] When fused, the output probabilities of each song identification model can be averaged or weighted summed to obtain a fusion prediction result, and then a judgment is made according to a fusion threshold corresponding to the fusion prediction result. For example, if the fusion probability is greater than or equal to the fusion threshold, it can be determined that the to-be-detected song is a machine-sung song. In an embodiment, a decision can be made using an auxiliary mechanism. For example, a song identification model assigned with a high threshold has a veto power. For example, if the fusion probability calculated by any model is greater than or equal to the higher threshold set in advance, it is considered that the song is AI-generated.
[0145] In some specific examples, after step 42, the song identification method of the embodiment of the application can further include: if the gap between the plurality of prediction probabilities exceeds a preset gap, determining that the to-be-detected song is a machine and human duet song. In other words, if the variance or standard deviation of the probability values calculated for the same song is large, it can be reasonably considered that the song is a human and AI duet or a duet.
[0146] In some specific examples, the song identification model can output time information of the AI-sung or human-sung segment.
[0147] The above identification result can be divided into two categories: AI singing and human singing. However, as a possible embodiment, the identification result can be divided into three categories according to a set threshold: AI singing, suspected AI singing, and human singing. Accordingly, subsequent operations can be continued according to the identification result: direct rejection (e.g., prohibition of release), prompt for manual review, or passing the review.
[0148] In summary, as shown in FIG. 5, the song voice identification method of the embodiment of the present application can separate the song voice and the accompaniment of the song to be tested to obtain the pure song voice without accompaniment. Then, each song voice segment obtained by slicing is input into the song voice identification model to output the prediction probability of the song voice being AI singing. The greater the probability value, the more likely the song voice is AI generated. Specifically, if the prediction probability of each song voice segment is greater than or equal to the target judgment threshold, it can be finally considered that the song is AI generated rather than sung by a real person.
[0149] In the embodiment of the present application, the operations performed by the song voice identification method are similar to the descriptions of the first aspect or any specific training method embodiment of the first aspect, and will not be repeated here. Of course, the specific implementation process of each operation of the first aspect of the present application can also be implemented by referring to the related description of the second aspect.
[0150] Referring to FIG. 6, the electronic device 600 of the embodiment of the present application can include one or more central processing units (CPUs) 601 and a memory 605 in which one or more application programs or data are stored.
[0151] The memory 605 can be volatile storage or persistent storage. The program stored in the memory 605 can include one or more modules, each of which can include a series of instruction operations in the electronic device. In another embodiment, the central processing unit 601 can be configured to communicate with the memory 605 to execute a series of instruction operations in the memory 605 on the electronic device 600.
[0152] The electronic device 600 can also include one or more power supplies 602, one or more wired or wireless network interfaces 603, one or more input and output interfaces 604, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0153] The central processing unit 601 can perform the operations performed by the first aspect or any specific method embodiment of the first aspect, and will not be repeated here.
[0154] The present application provides a computer readable storage medium, including instructions, when the instructions run on the computer, make the computer execute the method described in the first aspect or any specific implementation manner of the first aspect.
[0155] The computer program product provided in the application comprises instructions or a computer program, and when the computer program product is run on a computer, the computer is caused to perform the method described in the first aspect or any of the specific implementation manners of the first aspect.
[0156] It can be understood that, in various embodiments of the application, the sequence number of each step does not mean the order of execution, and the execution order of each step should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application. The operation contents added or refined by each example scheme of the above method, system or device (if any) do not necessarily have to be executed in specific implementation, such as two or more operations are added, which can be combined or implemented separately, and the specific implementation can be determined according to the actual scene.
[0157] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system (if any) and device can be referred to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0158] In several embodiments provided in the application, it should be understood that the disclosed apparatus and method can be implemented by other ways. For example, the apparatus embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system or apparatus, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, apparatus or unit, and can be electrical, mechanical or other forms.
[0159] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0160] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0161] The integrated unit, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the related art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product (computer program product) stored in a storage medium includes a plurality of instructions for causing a computer device (which can be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, read-only memory), a random access memory (RAM, random access memory), a magnetic disk or an optical disk, and various program code storage media. Industrial applicability:
[0162] The song authenticity identification model training method, song authenticity identification method and related products provided by the embodiments of the present application can select positive sample song segments and negative sample song segments with different labels from machine-sung song samples and human-sung song samples, which helps to provide sufficient and reliable data support for the training of the song authenticity identification model and promotes the song authenticity identification model to more deeply identify song segments in the to-be-tested song. In addition, the song feature extraction model can be used to extract the song features of each song segment, and the target classifier model can be trained using the label information, song features corresponding to the positive sample song segments and the negative sample song segments respectively. The song feature extraction model can effectively replace manual work, learn the acoustic differences between machine singing and human singing quickly and accurately, so as to ensure the efficiency and accuracy of the target classifier model in identifying the authenticity of the song. Through the above method, the technical problem of how to identify and detect machine-generated song can be solved, which can be applied to various scenes that need to identify machine-generated song, and has a wide range of applications.
Claims
1. A training method for a singing voice authentication model, characterized in that, The singing voice authentication model includes a singing voice feature extraction model and a target classifier model, and the training method includes: Each vocal segment is extracted from a machine-sung song sample as a positive sample vocal segment, and each vocal segment is extracted from a human-sung song sample as a negative sample vocal segment; the vocal segment refers to the audio segment in the song sample after removing the accompaniment and the segment without vocals. The positive sample singing segments and the negative sample singing segments are respectively input into the singing feature extraction model to extract the singing features corresponding to each of the positive sample singing segments and the negative sample singing segments; wherein, the singing feature extraction model is a feature extraction model trained towards a training objective; the training objective includes reducing the distance between the positive sample singing segments and increasing the distance between the positive sample singing segments and the negative sample singing segments; Using the label information corresponding to the positive sample singing segments and the negative sample singing segments, and the singing features, the initial classifier model is trained until the convergence condition is met, at which point training stops to obtain the target classifier model. The label information is used to mark whether the singing segments are machine-sung, and the target classifier model is used to output the predicted probability that the singing segments in the test song are machine-sung segments, so as to determine whether the test song is machine-sung.
2. The training method according to claim 1, wherein, After extracting vocal segments as positive sample vocal segments from machine-sung song samples, the training method further includes: The positive sample singing segments are enhanced to generate more positive sample singing segments that are also machine-generated.
3. The training method according to claim 1 or 2, wherein, After extracting the vocal features corresponding to the positive sample vocal segments and the negative sample vocal segments, the training method further includes: From the vocal features corresponding to the positive sample vocal segments and the negative sample vocal segments, at least one vocal representation feature is distinguished to train the initial classifier model; the vocal representation feature is used at least to represent the singing rhythm and / or emotional changes of the human voice signal in the vocal segments.
4. The training method according to claim 1 or 2, wherein, The step of training the initial classifier model using the label information corresponding to the positive sample singing segments and the negative sample singing segments, and the singing features, includes: The features other than the singing representation features in the singing features corresponding to the positive sample singing segments and the negative sample singing segments are used as singing auxiliary features; the singing representation features are used at least to represent the singing rhythm and / or emotional changes of the human voice signal in the singing segments. Using the positive sample singing segments, the negative sample singing segments, the label information corresponding to the positive sample singing segments and the negative sample singing segments respectively, and the singing auxiliary features, the singing representation model in the initial classifier model and the singing feature extraction model is jointly trained to obtain the target classifier model and the trained singing representation model. The trained singing voice representation model conforms to the training objective, and the learning rate of the singing voice representation model during training is less than the learning rate of the initial classifier model during training. The trained singing voice representation model is used to output the singing voice representation features corresponding to the singing voice segment through the time-domain signal of the singing voice segment.
5. The training method according to claim 4, wherein, If the singing auxiliary features include multiple types of features of different types, the training method further includes: The multiple features are weighted and fused to obtain an auxiliary feature fusion result; the auxiliary feature fusion result is used at least to train the initial classifier model.
6. The training method according to claim 1, wherein, If the fusion probability is greater than or equal to the target judgment threshold, the fusion probability is obtained by fusing the prediction probabilities corresponding to the positive sample singing segment and the negative sample singing segment respectively, then the sample song is determined to be a machine-sung song sample. The process of determining the target judgment threshold includes: Based on the ratio between the predicted number and the actual number of each song sample, the initial judgment threshold initially assigned to the target classifier model is adjusted until the initial judgment threshold is adjusted to the point where the ratio meets the preset result, and the adjustment is stopped to obtain the target judgment threshold. Wherein, the predicted number refers to the number of samples in each song sample that are predicted by the target classifier model to be machine-sung based on the initial judgment threshold, and the actual number refers to the number of samples in each song sample that are actually machine-sung.
7. A method for detecting fake singing voices, characterized in that, include: Obtain vocal fragments from the song to be tested; Each singing segment is input into the singing voice authentication model to obtain the prediction probability that each singing segment is predicted to be a machine-sung segment; the singing voice authentication model is trained according to the training method of any one of claims 1 to 6. The predicted probabilities of each vocal segment are fused and calculated to obtain the fusion probability corresponding to the song to be tested; If the fusion probability is greater than or equal to the target judgment threshold, then the song to be tested is determined to be a machine-sung song.
8. The method for detecting fake singing voices according to claim 7, wherein, If there are multiple singing voice authentication models, and the song samples of each singing voice authentication model have type differences, and the type differences do not include the distinction between machine singing and human singing, then the singing voice authentication method further includes: The fusion probability of the song under test is obtained by each of the vocal fakeness detection models. Based on the fusion probability of each of the aforementioned singing voice authentication models, it is comprehensively determined whether the song to be tested is sung by a machine.
9. The method for detecting fake singing voices according to claim 7, wherein, After inputting each vocal segment into the vocal probabilities detection model to obtain the predicted probability that each vocal segment is predicted to be a machine-generated segment, the vocal probabilities detection method further includes: If the difference between multiple predicted probabilities exceeds a preset gap, then the song to be tested is determined to be a song sung by a machine and a human.
10. An electronic device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 6 or any one of claims 7 to 9.
11. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 6 or any one of claims 7 to 9.
12. A computer program product comprising instructions or a computer program, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 6 or any one of claims 7 to 9.
Citation Information
Patent Citations
Training method of singing sound detection model, singing sound detection method, device and medium
CN116564278A
Turning and singing recognition model training method, turning and singing recognition method, equipment and storage medium
CN116844531A
Song identification method, computer equipment and storage medium
CN117765976A
Training method of singing authentic identification model, singing authentic identification method and related product
CN118571231A
Determining that Audio Includes Music and then Identifying the Music as a Particular Song
US20190102458A1
Cited By
A synthetic audio detection method and apparatus, an electronic device, and a medium
CN122369511A